On generalized entropy measures and pathways
A.M. Mathai, H.J. Haubold

TL;DR
This paper explores generalized entropy measures, their mathematical properties, and their applications to pathway models, emphasizing the importance of non-additivity in statistical mechanics and connecting to measures of inaccuracy.
Contribution
It introduces and examines properties of generalized entropies, including Mathai's entropy, and links them to pathway models and differential equations.
Findings
Shannon entropy's logarithmic form leads to additivity under product probability.
Mathai's generalized entropy can produce exponential and power law behaviors.
Connections are established between Mathai's entropy and Kerridge's measure of inaccuracy.
Abstract
Product probability property, known in the literature as statistical independence, is examined first. Then generalized entropies are introduced, all of which give generalizations to Shannon entropy. It is shown that the nature of the recursivity postulate automatically determines the logarithmic functional form for Shannon entropy. Due to the logarithmic nature, Shannon entropy naturally gives rise to additivity, when applied to situations having product probability property. It is argued that the natural process is non-additivity, important, for example, in statistical mechanics, even in product probability property situations and additivity can hold due to the involvement of a recursivity postulate leading to a logarithmic function. Generalizations, including Mathai's generalized entropy are introduced and some of the properties are examined. Situations are examined where Mathai's…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
ON GENERALIZED ENTROPY MEASURES AND PATHWAYS
A.M. MATHAI
Department of Mathematics and Statistics, McGill University, Montreal, Canada H3A 2K6, and
Centre for Mathematical Sciences, Pala Campus, Arunapuram P.O., Pala-686 574, Kerala, India
and
H.J. HAUBOLD
Office for Outer Space Affairs, United Nations, Vienna International Centre, P.O. Box 500, A-1400 Vienna, Austria and
Centre for Mathematical Sciences, Pala Campus, Arunapuram P.O., Pala-686 574, Kerala, India
Abstract. Product probability property, known in the literature as statistical independence, is examined first. Then generalized entropies are introduced, all of which give generalizations to Shannon entropy. It is shown that the nature of the recursivity postulate automatically determines the logarithmic functional form for Shannon entropy. Due to the logarithmic nature, Shannon entropy naturally gives rise to additivity, when applied to situations having product probability property. It is argued that the natural process is non-additivity, important, for example, in statistical mechanics (Tsallis 2004, Cohen 2005), even in product probability property situations and additivity can hold due to the involvement of a recursivity postulate leading to a logarithmic function. Generalized entropies are introduced and some of their properties are examined. Situations are examined where a generalized entropy of order leads to pathway models, exponential and power law behavior and related differential equations. Connection of this entropy to Kerridge’s measure of “inaccuracy” is also explored.
1. Introduction
Mathai and Rathie (1975) consider various generalizations of Shannon entropy (Shannon, 1948), called entropies of order , and give various properties, including additivity property, and characterization theorems. Recently, Mathai and Haubold (2006, 2006a) explored a generalized entropy of order , which is connected to a measure of uncertainty in a probability scheme, Kerridge’s (Kerridge, 1961) concept of inaccuracy in a scheme, and pathway models that are considered in this paper.
As defined in Mathai and Haubold (2006, 2006a) the entropy is a non-additive entropy and his measure is an additive entropy. It is also shown that maximization of the continuous analogue of , denoted by , gives rise to various functional forms for , depending upon the types of constraints on .
Occasionally, emphasis is placed on the fact that Shannon entropy satisfies the additivity property, leading to extensivity. It will be shown that when the product probability property (PPP) holds then a logarithmic function can give a sum and a logarithmic function enters into Shannon entropy due to the assumption introduced through a certain type of recursivity postulate. The concept of statistical independence will be examined in Section 1 to illustrate that simply because of PPP one need not expect additivity to hold or that one should not expect this PPP should lead to extensivity. The types of non-extensivity, associated with a number of generalized entropies, are pointed out even when PPP holds. The nature of non-extensivity that can be expected from a multivariate distribution, when PPP holds or when there is statistical independence of the random variables, is illustrated by taking a trivariate case.
Maximum entropy principle is examined in Section 2. It is shown that optimization of measures of entropies, in the continuous populations, under selected constraints, leads to various types of models. It is shown that the generalized entropy of order is a convenient one to obtain various probability models.
Section 3 examines the types of differential equations satisfied by the various special cases of the pathway model.
1.1. Product probability property (PPP) or statistical independence of events
Let denote the probability of the event . If the definition is taken as the definition of independence of the events and then any event , and the sure event are independent. But is contained in and then the definition of independence becomes inconsistent with the common man’s vision of independence. Even if the trivial cases of the sure event and the impossible event are deleted, still this definition becomes a resultant of some properties of positive numbers. Consider a sample space of distinct elementary events. If symmetry in the outcomes is assumed then we will assign equal probabilities each to the elementary events. Let . If and are independent then . Let
[TABLE]
Then
[TABLE]
deleting and . There is no solution for for a large number of , for example, . This means that there are no independent events in such cases and it sounds strange from a common man’s point of view.
The term “independence” of events is a misnomer. This property should have been called product probability property or PPP of events. There is no reason to expect the information or entropy in a joint distribution to be the sum of the information contents of the marginal distributions when the PPP holds for the distributions, that is when the joint density or probability function is a product of the marginal densities or probability functions. We may expect a term due to the product probability to enter into the expression for the entropy in the joint distribution in such cases. But if the information or entropy is defined in terms of a logarithm, then naturally, logarithm of a product being the sum of logarithms, we can expect a sum coming in such situations. This is not due to independence or due to the PPP of the densities but due to the fact that a functional involving logarithm is taken thereby a product has become a sum. Hence not too much importance should be put on whether or not the entropy on the joint distribution becomes sum of the entropies on marginal distributions or additivity property when PPP holds.
1.2. How is logarithm coming in Shannon’s entropy?
Several characterization theorems for Shannon entropy and its various generalizations are given in Mathai and Rathie (1975. Modified and refined versions of Shannon’s own postulates are given as postulates for the first theorem characterizing Shannon entropy in Mathai and Rathie (1975). Apart from continuity, symmetry, zero-indifference and normalization postulates the main postulate in the theorem is a recursivity postulate, which in essence says that when the PPP holds then the entropy will be a weighted sum of the entropies, thus in effect, assuming a logarithmic functional form. The crucial postulate is stated here. Consider a multinomial population , , , that is, , , . If any can take a zero value also then zero-indifferent postulate, namely that the entropy remains the same when an impossible event is incorporated into the scheme, is to be added. Let denote the entropy to be defined. Then the crucial recursivity postulate says that
[TABLE]
. This says that if the -th event is partitioned into independent events so that then the entropy becomes a weighted sum. Naturally, the result will be a logarithmic function for the measure of entropy.
There are several modifications to this crucial recursivity postulate. One suggested by Tverberg is that and and is assumed to be Lebesgue integrable in . Again a characterization of Shannon entropy is obtained. In all the characterization theorems for Shannon entropy this recursivity property enters in one form or the other as a postulate, which in effect implies a logarithmic form for the entropy measure. Shannon entropy has the following form:
[TABLE]
where is a constant. If any is assumed to be zero then is to be interpreted as zero. Since the constant is present, logarithm can be taken to any base. Usually the logarithm is taken to the base for ready application to binary systems. We will take logarithm to the base .
1.3. Generalization of Shannon entropy
Consider again a multinomial population . The following are some of the generalizations of Shannon entropy .
[TABLE]
When all the entropies of order described above in (4) to (7) go to Shannon entropy .
[TABLE]
Hence all the above measures are called generalized entropies of order .
Let us examine to see what happens to the above entropies in the case of a joint distribution. Let such that . This is a bivariate situation of a discrete distribution. Then the entropy in the joint distribution, for example,
[TABLE]
If the PPP holds and if , , , , and if then
[TABLE]
Therefore
[TABLE]
If any one of the above mentioned generalized entropies in (4) to (8) is written as then we have the relation
[TABLE]
where
[TABLE]
When the entropy is called additive and when the entropy is called non-additive. As can be expected, when a logarithmic function is involved, as in the cases of , the entropy is additive and .
1.4. Extensions to higher dimensional joint distributions
Consider a trivariate population or a trivariate discrete distribution ,,, such that . If the PPP holds mutually, that is, pair-wise as well as jointly, which then will imply that
[TABLE]
Then proceeding as before, we have for any of the measures described above in (4) to (8), calling it ,
[TABLE]
where is the same as in (13). The same procedure can be extended to any multivariable situation. If we may call the entropy additive and if then the entropy is non-additive.
1.5. Crucial recursivity postulate
Consider the multinomial population . Let the entropy measure to be determined through appropriate postulates be denoted by . For let
[TABLE]
If another parameter is to be involved in then we will denote by . From (5) to (7) it can be seen that the generalized entropies of order of Havrda-Charvát (1967), Tsallis (1988, 2004) and Shannon (1948) entropy satisfy the functional equation
[TABLE]
for with , with the boundary condition
[TABLE]
where
[TABLE]
Observe that the normalizing constant at is equal to for and it is different for other entropies. Thus equations (6),(7),(8), with the appropriate normalizing constants , can give characterization theorems for the various entropy measures. The form of is coming from the crucial recursivity postulate, assumed as a desirable property for the measures.
1.6. Continuous analogues
In the continuous case let be the density function of a real random variable . Then the various entropy measures, corresponding to the ones in (4) to (8) are the following:
[TABLE]
As expected, Shannon entropy in this case is given by
[TABLE]
where is a constant.
Note that when PPP (product probability property) or statistical independence holds then in the continuous case also we have the property in (12) and (14) and then non-additivity holds for the measures analogous to the ones in (3),(5),(6),(7) with remaining the same. Since the steps are parallel a separate derivation is not given here.
2. Maximum Entropy Principle
If we have a multinomial population or the scheme , , , then we know that the maximum uncertainty in the scheme or the minimum information from the scheme is obtained when we cannot give any preference to the occurrence of any particular event or when the events are equally likely or when . In this case, Shannon entropy becomes,
[TABLE]
and this is the maximum uncertainty or maximum Shannon entropy in this scheme. If the arbitrary functional is to be fixed by maximizing the entropy then in (19) to (21) we have to optimize for fixed , over all functional , subject to the condition and for all . For applying calculus of variation procedure we consider the functional
[TABLE]
where is a Lagrangian multiplier. Then the Euler equation is the following:
[TABLE]
Hence is the uniform density in this case, analogous to the equally likely situation in the multinomial case. If the first moment is assumed to be a given quantity for all functional then will become the following for (19) to (21).
[TABLE]
and the Euler equation leads to the power law. That is,
[TABLE]
By selecting appropriately we can create a density out of (27). For and the right side in (27) increases exponentially. If and then we have Tsallis’ -exponential function from the right side of (27). If and then (27) can produce a density in the category of a type-1 beta. From (27) it is seen that the form of the entropies of Havrda-Charvát and Tsallis need special attention to produce densities (Ferri et al. 2005). However, Tsallis has considered a different constraint on . If the density is replaced by its escort density, namely, where and if the expected value of in this escort density is assumed to be fixed for all functional then the of (26) becomes
[TABLE]
where is a constant and is the normalizing constant. If is taken as then
[TABLE]
Then (28) for is Tsallis statistics (Tsallis 2004, Cohen 2005). Then for also by writing one gets the case of Tsallis statistics for (Ferri et al. 2005). These modifications and the consideration of escort distribution are not necessary if we take the generalized entropy of order . Thus if we consider and if we assume that the first moment in itself is fixed for all functional then the Euler equation gives
[TABLE]
and for we have Tsallis statistics (Tsallis 2004, Cohen 2005)
[TABLE]
coming directly, where is the normalizing constant.
Let us start with of (20) under the assumptions that for all , , is fixed for all functional and for a specified , is the same for all functional , is the same for all functional , for some limits and , then the Euler equation becomes
[TABLE]
If is written as then we have, writing for ,
[TABLE]
where . For or the right side of (31) remains as a generalized type-1 beta model with the corresponding normalizing constant . For , writing the model in (31) goes to a generalized type-2 beta form, namely,
[TABLE]
When in (31) or in (32) we have an extended or stretched exponential form,
[TABLE]
If in (30) is taken as positive then (30) for will be increasing exponentially. Hence all possible forms are available from (30). The model in (31) is a special case of the distributional pathway model and for a discussion of the matrix-variate pathway model see Mathai (2005). Special cases of (31) and (32) for are Tsallis statistics (Gell-Mann and Tsallis, 2004; Ferri et al. 2005).
Instead of optimizing of (22) under the conditions that for all , and is fixed, let us optimize under the following conditions: for all , and the following two moment-like expressions are fixed quantities for all functional ,
[TABLE]
Then the Euler equation becomes
[TABLE]
and for , , we have the distributional pathway model for the real scalar case, namely
[TABLE]
where is the normalizing constant. For , (34) gives a generalized type-1 beta form, for it gives a generalized type-2 beta form and for we have a generalized gamma form. For , (34) gives the superstatistics of Beck (2006) and Beck and Cohen (2003). For , (34) gives Tsallis statistics (Tsallis 2004, Cohen 2005). Densities appearing in a number of physical problems are seen to be special cases of (34), a discussion of which may be seen from Mathai and Haubold (2006a). For example, (34) for is the Maxwell-Boltzmann density; for is the Gaussian density; for is the Weibull density. For we have the Wigner function giving the atomic moment distribution in the framework of Fokker-Planck equation, see Douglas, Bergamini, and Renzoni (2006) where
[TABLE]
Before closing this section we may observe one more property for . As an expected value
[TABLE]
But Kerridge’s (Kerridge, 1961) measure of “inaccuracy” in assigning for the true density , in the generalized form is
[TABLE]
which is also connected to the measure of directed divergence between and . In (37) the normalizing constant is , the same factor appearing in Havrda-Charvt́ entropy. With different normalizing constants, as seen before, (36) and (37) have the same forms as an expected value with replaced by in (36). Hence can also be looked upon as a type of directed divergence or “inaccuracy” measure.
3. Differential Equations
The functional part in (34), for a more general exponent, namely
[TABLE]
is seen to satisfy the following differential equation for which defines the differential pathway.
[TABLE]
Then for we have
[TABLE]
For in (38) we have
[TABLE]
Here (43) is the power law coming from Tsallis statistics (Gell-Mann and Tsallis, 2004).
Acknowledgement The authors would like to thank the Department of Science and Technology, Government of India, New Delhi, for the financial assistance for this work under project No. SR/S4/MS:287/05 which enabled this collaboration possible.
4. References
Beck, C. (2006). Stretched exponentials from superstatistics. Physica A, 365, 96-101.
Beck, C. and Cohen, E.G.D. (2003). Superstatistics. Physica A, 322, 267-275.
Cohen, E.G.D. (2005). Boltzmann and Einstein: Statistics and dynamics - An unsolved problem. Pramana, 64, 635-643.
Douglas, P., Bergamini, S., and Renzoni, F. (2006). Tunable Tsallis distribution in dissipative optical lattices. Physical Review Letters, 96, 110601-1-4.
Ferri, G.L., Martinez, S., and Plastino, A. (2005). Equivalence of the four versions of Tsallis’s statistics. Journal of Statistical Mechanics: Theory and Experiment, PO4009.
Gell-Mann, M. and Tsallis, C. (Eds.) (2004). Nonextensive Statistical Mechanics: Interdisciplinary Applications. Oxford University Press, Oxford.
Havrda, J. and Charvát, F. (1967). Quantification method of classification procedures: Concept of structural -entropy. Kybernetika, 3, 30-35.
Kerridge, D.F. (1961). Inaccuracy and inference. Journal of the Royal Statistical Society Series B, 23, 184-194.
Mathai, A.M. (2005). A pathway to matrix-variate gamma and normal densities. Linear Algebra and Its Applications, 396, 317-328.
Mathai, A.M. and Haubold, H.J. (2006). Pathway model, Tsallis statistics, superstatistics and a generalized measure of entropy. *Physica A *, 375), 110-122.
Mathai,A.M. and Haubold, H.J. (2006a). On generalized distributions and pathways. arXiv:cond-mat/0609526v2.
Mathai, A.M. and Rathie, P.N. (1975). Basic Concepts in Information Theory and Statistics: Axiomatic Foundations and Applications, Wiley Halstead, New York and Wiley Eastern, New Delhi.
Rényi, A. (1961). On measure of entropy and information. Proceedings of the Fourth Berkeley Symposium on Mathematical Statistics and Probability, 1960, University of California Press, 1961, Vol. 1, 547-561.
Shannon, C.E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27, 379-423, 547-561.
Tsallis, C. (1988). Possible generalization of Boltzmann-Gibbs statistics. Journal of Statistical Physics, 52, 479-487.
Tsallis, C. (2004). What should a statistical mechanics satisfy to reflect nature?, Physica D, 193, 3-34.
