A new approach to mutual information
Fumio Hiai1
Graduate School of Information Sciences,
Tohoku University, Aoba-ku, Sendai 980-8579, Japan
and
Dénes Petz2
Alfréd Rényi Institute of Mathematics, Hungarian
Academy of Sciences, H-1053 Budapest, Reáltanoda u. 13-15, Hungary
Abstract.
A new expression as a certain asymptotic limit via “discrete micro-states” of
permutations is provided to the mutual information of both continuous and discrete random variables.
1Supported in part by Grant-in-Aid for Scientific Research
(B)17340043.
2Supported in part by the Hungarian Research Grant OTKA T068258.
AMS subject classification: Primary: 62B10, 94A17.
Introduction
One of the important quantities in information theory is the mutual
information of two random variables X and Y which is expressed
in terms of the Boltzmann-Gibbs entropy H(⋅) as follows:
[TABLE]
when X,Y are continuous variables. For the expression of I(X∧Y) of discrete
variables X,Y, the above H(⋅) is replaced by the Shannon entropy. A more
practical and rigorous definition via the relative entropy is
[TABLE]
where μ(X,Y) denotes the joint distribution measure of (X,Y) and
μX⊗μY the product of the respective distribution measures
of X,Y.
The aim of this paper is to show that the mutual information I(X∧Y)
is gained as a certain asymptotic limit of the volume of “discrete
micro-states” consisting of permutations approximating joint moments of
(X,Y) in some way. In Section 1, more generally we consider an n-tuple of
real bounded random variables (X1,…,Xn). Denote by
Δ(X1,…,Xn;N,m,δ) the set of (x1,…,xn) of
xi∈RN whose joint moments (on the uniform distributed N-point set)
of order up to m approximate those of (X1,…,Xn) up to an error
δ. Furthermore, denote by Δsym(X1,…,Xn;N,m,δ)
the set of (σ1,…,σn) of permutations σi∈SN
such that
(σ1(x1),…,σn(xn))∈Δ(X1,…,Xn;N,m,δ)
for some x1,…,xn∈R≤N, where R≤N is the RN-vectors
arranged in increasing order. Then, the asymptotic volume
[TABLE]
under the uniform probability measure γSN on SN is shown to
converge as limsupN→∞ (also liminfN→∞) and then
limm→∞,δ↘0 to
[TABLE]
as long as H(Xi)>−∞ for 1≤i≤n. Thus, we obtain a kind of
discretization of the mutual information via symmetric group (or permutations).
The approach can be applied to an n-tuple of discrete random variables
(X1,…,Xn) as well. But the definition of the Δsym-set of
micro-states for discrete variables is somewhat different from the continuous
variable case mentioned above, and we discuss the discrete variable case in
Section 2 separately.
The idea comes from the paper [3]. Motivated by theory of mutual free
information in [6], a similar approach to Voiculescu’s free entropy
is provided there. The free entropy is the free probability counterpart of
the Boltzmann-Gibbs entropy, and RN-vectors and the symmetric group
SN here are replaced by Hermitian N×N matrices and the unitary
group U(N), respectively. In this way, the “discretization
approach” here is in some sense a classical analog of the “orbital approach”
in [3].
1. The continuous case
For N∈N let R≤N be the convex cone of the N-dimensional Euclidean
space RN consisting of x=(x1,…,xN) such that
x1≤x2≤⋯≤xN. The space RN is naturally regarded as the real
function algebra on the N-point set. Let SN be the symmetric group of order
N (i.e., the permutations on {1,2,…,n}). Throughout this section let
(X1,…,Xn) be an n-tuple of real random variables on a probability space
(Ω,P), and assume that the Xi’s are bounded (i.e.,
Xi∈L∞(Ω;P)). The Boltzmann-Gibbs entropy of
(X1,…,Xn) is defined to be
[TABLE]
if the joint density p(x1,…,xn) of (X1,…,Xn) exists; otherwise
H(X1,…,Xn)=−∞. Note that the above integral is well defined in
[−∞,∞) since the density p is compactly supported.
Definition 1.1**.**
The mean value of x=(x1,…,xN) in RN is given by
[TABLE]
For each N,m∈N and δ>0 we define Δ(X1,…,Xn;N,m,δ) to
be the set of all n-tuples (x1,…,xn) of
xi=(xi1,…,xiN)∈RN, 1≤i≤n, such that
[TABLE]
for all 1≤i1,…,ik≤n with 1≤k≤m, where
xi1⋯xik means the pointwise product, i.e.,
[TABLE]
and E(⋅) denotes the expectation on (Ω,P). For each R>0, define
ΔR(X1,…,Xn;N,m,δ) to be the set of all
(x1,…,xn)∈Δ(X1,…,Xn;N,m,δ) such that
xi∈[−R,R]N for all 1≤i≤n.
Heuristically, Δ(X1,…,Xn;N,m,δ) is the set of “micro-states”
consisting of n-tuples of discrete random variables on the N-point set with
the uniform probability such that all joint moments of order up to m give the
corresponding joint moments of X1,…,Xn up to an error δ.
For x∈RN write ∥x∥p:=(N−1∑j=1N∣xj∣p)1/p for
1≤p<∞ and ∥x∥∞:=max1≤j≤N∣xj∣ while ∥X∥p denotes
the Lp-norm of a real random variable X on (Ω,P).
The next lemma is seen from [4, 5.1.1] based on the Sanov large deviation
theorem, which says that the Boltzmann-Gibbs entropy is gained as an asymptotic limit
of the volume of the approximating micro-states.
Lemma 1.2**.**
For every m∈N and δ>0 and for any choice of
R≥max1≤i≤n∥Xi∥∞, the limit
[TABLE]
exists, where λN is the Lebesgue measure on RN. Furthermore, one has
[TABLE]
independently of the choice of R≥max1≤i≤n∥Xi∥∞.
In the following let us introduce some kinds of mutual information in the
discretization approach using micro-states of permutations.
Definition 1.3**.**
The action of SN on RN is given by
[TABLE]
for σ∈SN and x=(x1,…,xN)∈RN. For each N,m∈N,
δ>0 and R>0 we denote by Δsym,R(X1,…,Xn;N,m,δ) the
set of all (σ1,…,σn)∈SNn such that
[TABLE]
for some (x1,…,xn)∈(R≤N)n. For each R>0 define
[TABLE]
where γSN is the uniform probability measure on SN. Define also
Isym,R(X1,…,Xn) by replacing limsup by liminf. Obviously,
[TABLE]
Moreover, Δsym,∞(X1,…,Xn;N,m,δ) is defined by replacing
ΔR(X1,…,Xn;N,m,δ) in the above by
Δ(X1,…,Xn;N,m,δ) without cut-off by the parameter R. Then
Isym,∞(X1,…,Xn) and Isym,∞(X1,…,Xn) are
also defined as above.
Definition 1.4**.**
For each 1≤i≤n we choose and fix a sequence ξi={ξi(N)} of
ξi(N)∈R≤N, N∈N, such that κN(ξi(N)k)→E(Xik) as
N→∞ for all k∈N, i.e., ξi(N)→Xi in moments. For each
N,m∈N and δ>0 we define
Δsym(X1,…,Xn:ξ1(N),…,ξn(N);N,m,δ) to be the set of all
(σ1,…,σn)∈SNn such that
[TABLE]
Define
[TABLE]
and Isym(X1,…,Xn:ξ1,…ξn) by replacing limsup by
liminf.
The next proposition asserts that the quantities in Definitions 1.3 and
1.4 are all equivalent.
Lemma 1.5**.**
For any choice of R≥max1≤i≤n∥Xi∥∞ and for any choices of
approximating sequences ξ1,…,ξn one has
[TABLE]
Proof.
It is obvious that Δsym(X1,…,Xn:ξ1(N),…,ξn(N);N,m,δ)
is included in Δsym,∞(X1,…,Xn;N,m,δ) for any approximating
sequences ξi. Moreover, for each 1≤i≤n an approximating sequence ξi
can be chosen so that ∥ξi(N)∥∞≤∥Xi∥∞ for all N; then
Δsym(X1,…,Xn:ξ1(N),…,ξn(N);N,m,δ)⊂Δsym,R(X1,…,Xn;N,m,δ) for any
R≥R0:=max1≤i≤n∥Xi∥∞. Hence it suffices to prove that for any
approximating sequences ξi and for every m∈N and δ>0, there are an
m′∈N, a δ′>0 and an N0∈N so that
[TABLE]
for all N≥N0. Choose a ρ∈(0,1) with m(R0+1)m−1ρ<δ/2.
By [5, Lemma 4.3] (also [4, 4.3.4]) there exist an m′∈N with
m′≥2m, a δ′>0 with δ′≤min{1,δ/2} and an N0∈N
such that for every 1≤i≤n and every x∈R≤N with N≥N0, if
∣κN(xk)−E(Xik)∣<δ′ for all 1≤k≤m′, then
∥x−ξi(N)∥m<ρ. Suppose N≥N0 and
(σ1,…,σn)∈Δsym,∞(X1,…,Xn;N,m′,δ′);
then (σ1(x1),…,σn(xn))∈Δ(X1,…,Xn;N,m′,δ′)
for some (x1,…,xn)∈(R≤N)n. Since
∣κN(xik)−E(Xik)∣<δ′ for all 1≤k≤m′, we get
∥xi−ξi(N)∥m≤ρ and
[TABLE]
Therefore,
[TABLE]
for all 1≤i1,…,ik≤n with 1≤k≤m. The above latter inequality
follows from the Hölder inequality. Hence (σ1,…,σn)∈Δsym(X1,…,Xn:ξ1(N),…,ξn(N);N,m,δ), and the result
follows.
∎
Consequently, we denote all the quantities in (1.1) by the same
Isym(X1,…,Xn) and those in (1.2) by Isym(X1,…,Xn).
We call Isym(X1,…,Xn) and Isym(X1,…,Xn) the
mutual information and upper mutual information of (X1,…,Xn),
respectively. The terminology “mutual information” will be justified after the next
theorem.
In the continuous variable case, our main result is the following exact relation of
Isym and Isym with the Boltzmann-Gibbs entropy H(⋅), which
says that Isym(X1,…,Xn) is formally the sum of the separate entropies
H(Xi)’s minus the compound H(X1,…,Xn). Thus, a naive meaning of
Isym(X1,…,Xn) is the entropy (or information) overlapping among the Xi’s.
Theorem 1.6**.**
[TABLE]
Proof.
If the coordinates si of s∈RN are all distinct, then s is uniquely
written as s=σ(x) with x∈R≤N and σ∈SN. Note that
the set of s∈RN with si=sj for some i=j is a closed subset of
λN-measure zero. Under the correspondence
[TABLE]
(well defined on a co-negligible subset of RN), the measure λN is
transformed into the product of λN∣R≤N and the counting measure on
SN.
In the following proof we adopt, due to Lemma 1.5, the description of
Isym and Isym as Isym,R(X1,…,Xn) and
Isym,R(X1,…,Xn) with R:=max1≤i≤n∥Xi∥∞.
For each N,m∈N and δ>0, suppose
(s1,…,sn)∈ΔR(X1,…,Xn;N,m,δ) and write
si=σi(xi) with xi∈R≤N and σi∈SN. Then it is
obvious that
[TABLE]
By Lemma 1.2 and the fact stated at the beginning of the proof, we obtain
[TABLE]
This implies that
[TABLE]
Conversely, for each m∈N and δ>0, by [5, Lemma 4.3] (also
[4, 4.3.4]) there are an m′∈N with m′≥m, a δ′>0 with
δ′≤δ/2 and an N0∈N such that for every N∈N and
for every x,y∈R≤N, if ∥x∥∞≤R and
∣κN(xk)−κN(yk)∣<2δ′ for all 1≤k≤m′, then
∥x−y∥1<δ/2m(R+1)m−1. Suppose N≥N0 and
[TABLE]
so that
(σ1(y1),…,σn(yn))∈ΔR(X1,…,Xn;N,m′,δ′)
for some (y1,…,yn)∈(R≤N)n. Since
[TABLE]
for all 1≤k≤m′, we get ∥xi−yi∥1<δ/2m(R+1)m−1 for
1≤i≤n. Therefore,
[TABLE]
for all 1≤i1,…,ik≤n with 1≤k≤m. This implies that
(σ1(x1),…,σn(xn))∈ΔR(X1,…,Xn;N,m,δ). By
Lemma 1.2 we obtain
[TABLE]
This implies by Lemma 1.2 once again that
[TABLE]
The result follows from (1.3) and (1.4).
∎
Let μ(X1,…,Xn) be the joint distribution measure on Rn of
(X1,…,Xn) while μXi is that of Xi for 1≤i≤n. Let
S(μ(X1,…,Xn),μX1⊗⋯⊗μXn) denote the
relative entropy (or the Kullback-Leibler divergence) of
μ(X1,…,Xn) with respect to the product measure
μX1⊗⋯⊗μXn, i.e.,
[TABLE]
if μ(X1,…,Xn) is absolutely continuous with respect to
μX1⊗⋯⊗μXn; otherwise
S(μ(X1,…,Xn),μX1⊗⋯⊗μXn):=+∞.
When H(Xi)>−∞ for all 1≤i≤n, one can easily verify that
[TABLE]
Thus, the above theorem yields the following:
Corollary 1.7**.**
If H(Xi)>−∞ for all 1≤i≤n, then
[TABLE]
Corollary 1.8**.**
Under the same assumption as the above corollary, Isym(X1,…,Xn)=0 if and
only if X1,…,Xn are independent.
In particular, the original mutual information I(X1∧X2) of two real random
variables X1,X2 is normally defined as
[TABLE]
Hence we have
[TABLE]
as long as H(X1)>−∞ and H(X2)>−∞ (and X1,X2 are bounded). For this
reason, we gave the term “mutual information” to Isym.
Finally, some open problems are in order:
Without the assumption H(Xi)>−∞ for 1≤i≤n, does
Isym(X1,…,Xn)=Isym(X1,…,Xn) hold true?
More strongly, does the limit such as
[TABLE]
or
[TABLE]
exist as in Lemma 1.2?
Without the assumption H(Xi)>−∞ for 1≤i≤n, does
Isym(X1,…,Xn)=S(μ(X1,…,Xn),μX1⊗⋯⊗μXn) hold true?
Also, is Isym(X1,…,Xn)=0 equivalent to the independence of X1,…,Xn?
Although the boundedness assumption for X1,…,Xn is rather essential
in the above discussions, it is desirable to extend the results in this section to
X1,…,Xn not necessarily bounded but having all moments.
2. The discrete case
Let Y be a finite set with a probability measure p. The Shannon entropy
of p is
[TABLE]
For each sequence y=(y1,…,yN)∈YN, the type of y is a
probability measure on Y given by
[TABLE]
The number of possible types is smaller than (N+1)#Y.
If ν is a type and TN(ν) denotes the set of all sequences of type ν
from YN, then the cardinality of TN(ν) is estimated as follows:
[TABLE]
(see [1, 12.1.3] and [2, Lemma 2.2]).
Let p be a probability meausre on Y. For each N∈N and
δ>0 we define Δ(p;N,δ) to be the set of all sequences
y∈YN such that ∣νy(t)−p(t)∣<δ for all t∈Y.
In other words, Δ(p;N,δ) is the set of all δ-typical
sequeces (with respect to the measure p). Then the next lemma is well known.
Lemma 2.1**.**
[TABLE]
In fact, this easily follows from (2.1). Let PN,δ be the
maximizer of the Shannon entropy on the set of all types νy,
y∈YN, such that ∣νy(t)−p(t)∣<δ for all t∈Y. We can use the Shannon entropy of the type class corresponding to
PN,δ to estimate the cardinality of Δ(p;N,δ):
[TABLE]
It follows that
[TABLE]
and the lemma follows.
We consider the case where p is the joint distribution of an n-tuple
(X1,…,Xn) of discrete random variables on (Ω,P). Throughout this
section we assume that the random variables X1,…,Xn have their values in
a finite set X={t1,…,td}.
Definition 2.2**.**
Let p(X1,…,Xn) denote the joint distribution of (X1,…,Xn),
which is a measure on Xn while the distribution pXi of Xi is a
measure on X, 1≤i≤n. We write Δ(Xi;N,δ) for
Δ(pXi;N,δ) and Δ(X1,…,Xn;N,δ) for
Δ(p(X1,…,Xn);N,δ).
Next, we introduce the counterparts of Definitions 1.3 and 1.4 in
the discrete variable case.
Definition 2.3**.**
The action of SN on XN is similar to that on RN given in Defintion
1.3. For N∈N let X≤N denote the set of all sequences of length
N of the form
[TABLE]
Oviously, such a sequence x is uniquely determined by
(Nx(t1),…,Nx(td)) or the type of x. That is, X≤N is
regarded as the set of all types from XN. For each N∈N and δ>0 we
denote by Δsym(X1,…,Xn;N,δ) the set of all
(σ1,…,σn)∈SNn such that
[TABLE]
for some (x1,…,xn)∈(X≤N)n. Define
[TABLE]
and Isym(X1,…,Xn) by replacing limsup by liminf. Moreover,
for each 1≤i≤n, choose a sequence ξi={ξi(N)} of
ξi(N)=(ξi(N)1,…,ξi(N)N)∈X≤N such that
νξi(N)→pXi as N→∞. We then
define Δsym(X1,…,Xn:ξ1(N),…,ξn(N);N,δ),
Isym(X1,…,Xn:ξ1,…,ξn) and
Isym(X1,…,Xn:ξ1,…,ξn) as in Definition 1.4.
Lemma 2.4**.**
For any choices of approximating sequences ξ1,…,ξn one has
[TABLE]
Proof.
It suffices to show that for each δ>0 there are a δ′>0 and an
N0∈N such that
[TABLE]
for all N≥N0. Choose δ′>0 so that 3ndn+1δ′≤δ, where
d=#X. Suppose (σ1,…,σn) is in the left-hand side of
(2.2) so that
(σ1(x1),…,σn(xn))∈Δ(X1,…,Xn;N,δ′) for some (x1,…,xn),
xi=(xi1,…,xiN)∈X≤N. Since
[TABLE]
[TABLE]
[TABLE]
it follows that
[TABLE]
for any 1≤i≤n and t∈X. Now, choose an N0∈N so that
∣νξi(N)(t)−pXi(t)∣<δ′ and hence
[TABLE]
for any 1≤i≤n and t∈X and for all N≥N0. Since
[TABLE]
for every 1≤l≤d thanks to (2.5), it is easily seen that
[TABLE]
for any 1≤i≤n. Hence we get
[TABLE]
so that thanks to (2.3)
[TABLE]
for every (z1,…,zn)∈Xn. Therefore, (σ1,…,σn) is in
the right-hand side of (2.2), as required.
∎
The next theorem is the discrete variable version of Theorem 1.6.
Theorem 2.5**.**
[TABLE]
Proof.
For each sequence (N1,…,Nd) of integers Nl≥0 with ∑l=1dNl=N,
let S(N1,…,Nd) denote the subgroup of SN consisting of products of
permutations of {1,…,N1}, {N1+1,…,N1+N2}, …,
{N1+⋯+Nd−1+1,…,N}, and let
[TABLE]
be the set of left cosets of S(N1,…,Nd). For each x∈X≤N and
σ∈SN we write [σ]x for the left coset of
S(Nx(t1),…,Nx(td)) containing σ. Then it is clear that every
s∈XN is represented as s=σ(x) with a unique pair
(x,[σ]x) of x∈X≤N and
[σ]x∈SN/S(Nx(t1),…,Nx(td)).
For any ε>0 one can choose a δ>0 such that for every 1≤i≤n and
every probability measure p on X, if ∣p(t)−pXi(t)∣<δ for all
t∈X, then ∣S(p)−S(pXi)∣<ε. This implies that for each N∈N and
1≤i≤n, one has ∣S(νx)−S(pXi)∣<ε whenever
x∈Δ(Xi;N,δ). Notice that
Δsym(X1,…,Xn;N,δ/dn−1) is the union of
[σ1]x1×⋯×[σn]xn for all
(x1,…,xn;[σ1]x1,…,[σn]xn) of
xi∈X≤N and
[σi]xi∈SN/S(Nxi(t1),…,Nxi(td)) such that
(σ1(x1),…,σn(xn))∈Δ(X1,…,Xn;N,δ/dn−1).
Now, suppose (x1,…,xn)∈(X≤N)n,
(σ1,…,σn)∈SNn and
(σ1(x1),…,σn(xn))∈Δ(X1,…,Xn;N,δ/dn−1).
Then, for each 1≤i≤n we get xi∈Δ(Xi;N,δ), i.e.,
∣νxi(t)−pXi(t)∣<δ for all t∈X as (2.4).
Hence we have
[TABLE]
so that
[TABLE]
Therefore,
[TABLE]
For each 1≤i≤n and for any x∈Δ(Xi;N,,δ), the Stirling
formula yields
[TABLE]
thanks to the above choice of δ>0. Here, note that the o(1) in the above
estimate is uniform for x∈Δ(Xi;N,δ).
Hence, by (2), (2) and by Lemma 2.1 applied to
p(X1,…,Xn) on Xn, we obtain
[TABLE]
and hence
[TABLE]
Next, we prove the converse direction. For any ε>0 choose a δ>0 as above.
For N∈N let Ξ(N,δ/dn−1) be the set of all
(x1,…,xn)∈(X≤N)n such that
[TABLE]
for some (σ1,…,σn)∈SNn. Furthermore, for each
(x1,…,xn)∈Ξ(N,δ/dn−1), let
Σ(x1,…,xn;N,δ/dn−1) be the set of all
[TABLE]
such that
(σ1(x1),…,σn(xn))∈Δ(X1,…,Xn;N,δ/dn−1).
Then it is obvious that
[TABLE]
When (x1,…,xn)∈Ξ(N,δ/dn−1), we get
xi∈Δ(Xi;N,δ) as (2.4) for 1≤i≤n. Hence it is seen
that
[TABLE]
For any fixed (x1,…,xn)∈Ξ(N,δ/dn−1), suppose
([σ1]x1,…,[σn]xn)∈Σ(x1,…,xn;N,δ/dn−1); then we get
[TABLE]
similarly to (2.6). Therefore,
[TABLE]
By (2.10)–(2) we obtain
[TABLE]
so that
[TABLE]
Since it follows similarly to (2) that
[TABLE]
with uniform o(1) for all x∈Δ(Xi;N,δ), we obtain
[TABLE]
by Lemma 2.1 again, and hence
[TABLE]
The conclusion follows from (2.9) and (2.13).
∎
In particular, the mutual information I(X1∧X2) of X1 and X2 is equivalently
expressed as
[TABLE]
Similarly to the problem (2) mentioned in the last of Section 1, it is unknown
whether the limit
[TABLE]
exists or not.