Non-monotone convergence in the quadratic Wasserstein distance
Walter Schachermayer, Uwe Schmock, Josef Teichmann

TL;DR
This paper provides a simple counter-example demonstrating that the quadratic Wasserstein distance does not always decrease monotonically when considering n-fold normalized convolutions of two measures, challenging an existing assumption.
Contribution
It presents the first explicit counter-example to the presumed monotonicity of quadratic Wasserstein distance under convolution.
Findings
Quadratic Wasserstein distance can increase under convolution.
Counter-example disproves the monotonicity assumption.
Challenges previous beliefs in mass transport theory.
Abstract
We give an easy counter-example to Problem 7.20 from C. Villani's book on mass transport: in general, the quadratic Wasserstein distance between -fold normalized convolutions of two given measures fails to decrease monotonically.
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsMathematical Analysis and Transform Methods · Advanced Banach Space Theory · Approximation Theory and Sequence Spaces
Non-monotone convergence in the quadratic Wasserstein distance
Walter Schachermayer, Uwe Schmock, and Josef Teichmann
Financial and Actuarial Mathematics, Technical University Vienna, Wiedner Hauptstrasse 8–10, A-1040 Vienna, Austria.
(Date: October 4, 2006)
Abstract.
We give an easy counter-example to Problem 7.20 from C. Villani’s book on mass transport: in general, the quadratic Wasserstein distance between -fold normalized convolutions of two given measures fails to decrease monotonically.
Financial support from the Austrian Science Fund (FWF) under grant P 15889, from the Vienna Science and Technology Fund (WWTF) under grant MA13, from the European Union under grant HPRN-CT-2002-00281 is gratefully acknowledged. Furthermore this work was financially supported by the Christian Doppler Research Association (CDG) via PRisMa Lab. The authors gratefully acknowledge a fruitful collaboration and continued support by Bank Austria Creditanstalt (BA-CA) and the Austrian Federal Financing Agency (ÖBFA) through CDG
We use the terminology and notation from [5]. For Borel measures , on we define the quadratic Wasserstein distance
[TABLE]
where is the Euclidean distance on and the pairs run through all random vectors defined on some common probabilistic space , such that has distribution and has distribution . By a slight abuse of notation we define for two random vectors , , such that has distribution and has distribution . The following theorem (see [5, Proposition 7.17]) is due to Tanaka [4].
Theorem 1**.**
For and square integrable random vectors , , , such that is independent of , and is independent of , and or , we have
[TABLE]
For a sequence of i.i.d. random vectors we define the normalized partial sums
[TABLE]
If denotes the law of , we write for the law of . Clearly equals, up to the scaling factor , the -fold convolution of .
We shall always deal with measures , with vanishing barycenter. Given two measures and on with finite second moments, we let and be i.i.d. sequences with law and , respectively, and denote by and the corresponding normalized partial sums. From Theorem 1 we obtain
[TABLE]
from which one may quickly deduce a proof of the Central Limit Theorem (compare [5, Ch. 7.4] and the references given there).
However, we can not deduce from Theorem 1 that the inequality
[TABLE]
holds true for all . Specializing to the case , an estimate, which we can obtain from Tanaka’s Theorem, is
[TABLE]
This contains some valid information, but does not imply (1). It was posed as Problem 7.20 of [5], whether inequality (1) holds true for all probability measures , on and all .
The subsequent easy example shows that the answer is no, even for and symmetric measures. We can choose and for sufficiently large , as the proposition (see also Remark 1) shows.
Proposition 1**.**
Denote by n the distribution of , and by the distribution of with i.i.d. and . Then
[TABLE]
while for all .
Remark 1**.**
If one only wants to find a counter-example to Problem 7.20 of [5], one does not really need the full strength of Proposition 1, i.e. the estimate that . In fact, it is sufficient to consider the case in order to contradict the monotonicity of inequality (1). Indeed, a direct calculation reveals that
[TABLE]
Proof of Proposition 1.
We start with the final assertion, which is easy to show. The -fold convolutions of the measures and , respectively, are supported on odd and even numbers, respectively. Hence they have disjoint supports with distance and so the quadratic transportation costs are bounded from below by .
For the proof of (2), fix , define and , and note that and are supported by the even numbers. For we denote by the probability of the point under , i.e.
[TABLE]
We define for . We have , where is the distribution giving probability , , to , [math], , respectively. We deduce that for ,
[TABLE]
Notice that for . The term in the first parentheses is therefore non-negative. It can easily be calculated and estimated via
[TABLE]
for .
Following [5] we know that the quadratic Wasserstein distance can be given by a cyclically monotone transport plan . We define the transport plan via an intuitive transport map . It is sufficient to define for , since it acts symmetrically on the negative side. moves mass from the point to for . At the transport moves to every side, which is possible, since there is enough mass concentrated at [math].
By equation (3) we see that the transport moves to , since, for , the first terms corresponds to the mass, which arrives from the left and is added to , and the second term to the mass, which is transported away: summing up one obtains . For , mass only arrives from the left. At mass is only transported away. By the symmetry of the problem around [math] and by the quadratic nature of the cost function (the distance of the transport is , hence cost ), we finally have
[TABLE]
By the Central Limit Theorem and uniform integrability of the function with respect to the binomial approximations, we obtain
[TABLE]
Hence
[TABLE]
In order to obtain equality we start from the local monotonicity of the respective transport maps on non-positive and non-negative numbers. It easily follows that the given transport plan is cyclically monotone and hence optimal (see [5, Ch. 2]). The subsequent equality allows also to consider estimates from below. Rewriting (3) yields
[TABLE]
for , and
[TABLE]
for . Furthermore,
[TABLE]
for . This yields by a reasoning similar to the above that
[TABLE]
hence
[TABLE]
∎
Remark 2**.**
Let be an integer. By slight modifications of the proof of Proposition 1 we can construct sequences of measures and , such that the quadratic Wasserstein distances of -fold convolutions are bounded from below by for all which are not multiples of , while
[TABLE]
Remark 3**.**
Assume the notations of [5]. In the previous considerations we can replace the quadratic cost function by any other lower semi-continuous cost function , which is bounded on parallels to the diagonal and vanishes on the diagonal. For example, if we choose for , then we obtain the same asymptotics as in Proposition 1 (with a different constant).
Remark 4**.**
We have used in the above proof that is obtained from by convolving with the measure . In fact, this theme goes back (at least) as far as L. Bachelier’s famous thesis from 1900 on option pricing [2, p. 45]. Strictly speaking, L. Bachelier deals with the measure assigning mass to , and considers consecutive convolutions, instead of the above . Hence convolutions with correspond to Bachelier’s result after two time steps. Bachelier makes the crucial observation that this convolution leads to a radiation of probabilities: Each stock price radiates during a time unit to its neighboring price a quantity of probability proportional to the difference of their probabilities. This was essentially the argument which allowed us to prove (1). Let us mention that Bachelier uses this argument to derive the fundamental relation between Brownian motion (which he was the first to define and analyse in his thesis) and the heat equation (compare e.g. [3] for more on this topic).
Remark 5**.**
Having established the above counterexample, it becomes clear how to modify Problem 7.20 from [5] to give it a chance to hold true. This possible modification was also pointed out to us by C. Villani.
Problem 1**.**
Let be a probability measure on with finite second moment and vanishing barycenter, and the Gaussian measure with same first and second moments. Does decrease monotonically to zero?
When entropy is considered instead of the quadratic Wasserstein distance the corresponding question on monotonicity was answered affirmatively in the recent paper [1].
One may also formulate a variant of Problem 7.20 as given in (1) by replacing the measure through a log-concave probability distribution. This would again generalize problem 1.
The reference list from the paper itself. Each links out to its DOI / PubMed record.
- 1[1] S. Artstein, K. M. Ball, F. Barthe and A. Naor, Solution of Shannon’s Problem on the Monotonicity of Entropy , Journal of the AMS 17 (4), 2004, pp. 975–982.
- 2[2] L. Bachelier, Theorie de la Speculation , Paris, 1900, see also: http://www.numdam.org/en/ .
- 3[3] W. Schachermayer, Introduction to the Mathematics of Financial Markets , LNM 1816 - Lectures on Probability Theory and Statistics, Saint-Flour summer school 2000 (Pierre Bernard, editor), Springer Verlag, Heidelberg (2003), pp. 111–177.
- 4[4] H. Tanaka, An inequality for a functional of probability distributions and its applications to Kac’s one-dimensional model of a Maxwell gas , Zeitschrift f r Wahrscheinlichkeitstheorie und verwandte Gebiete 27 , 47–52, 1973
- 5[5] C. Villani, Topics in Optimal Transportation , Graduate Studies in Mathematics 58 , American Mathematical Society, Providence Rhode Island, 2003.
