α-Geodesical Skew Divergence

Kimura, Masanari; Hino, Hideitsu

doi:10.3390/e23050528

Open AccessEditor’s ChoiceArticle

α-Geodesical Skew Divergence

by

Masanari Kimura

^1,*

and

Hideitsu Hino

²

¹

Department of Statistical Science, School of Multidisciplinary Sciences, The Graduate University for Advanced Studies (SOKENDAI), Kanagawa 240-0193, Japan

²

The Institute of Statistical Mathematics, Tokyo 190-0014, Japan

^*

Author to whom correspondence should be addressed.

Entropy 2021, 23(5), 528; https://doi.org/10.3390/e23050528

Submission received: 1 April 2021 / Revised: 22 April 2021 / Accepted: 24 April 2021 / Published: 25 April 2021

Download

Browse Figures

Review Reports Versions Notes

Abstract

:

The asymmetric skew divergence smooths one of the distributions by mixing it, to a degree determined by the parameter

λ

, with the other distribution. Such divergence is an approximation of the KL divergence that does not require the target distribution to be absolutely continuous with respect to the source distribution. In this paper, an information geometric generalization of the skew divergence called the

α

-geodesical skew divergence is proposed, and its properties are studied.

Keywords:

KL-divergence; JS-divergence; skew divergence; information geometry

1. Introduction

Let

(X, F, μ)

be a measure space where

X

denotes the sample space,

F

the

σ

-algebra of measurable events, and

μ

a positive measure. The set of the strictly positive probability measure

P

is defined as

P \{f (x) > 0 (\forall x \in X), and \int_{X} f (x) d μ (x) = 1\},

(1)

and the set of nonnegative probability measure

P_{+}

is defined as

P_{+} \{f (x) \geq 0 (\forall x \in X), and \int_{X} f (x) d μ (x) = 1\} .

(2)

Then a number of divergences that appear in statistics and information theory [1,2] are introduced.

Definition 1

(Kullback–Leibler divergence [3]). The Kullback–Leibler divergence or KL-divergence

D_{K L} : P_{+} \times P \to [0, \infty]

is defined between two Radon–Nikodym densities p and q of μ-absolutely continuous probability measures by

D_{K L} [p ∥ q] \int_{X} p ln \frac{p}{q} d μ .

(3)

KL-divergence is a measure of the difference between two probability distributions in statistics and information theory [4,5,6,7]. This is also called the relative entropy and is known not to satisfy the axiom of distance. Because the KL-divergence is asymmetric, several symmetrizations have been proposed in the literature [8,9,10].

Definition 2

(Jensen–Shannon divergence [8]). The Jensen–Shannon divergence or JS-divergence

D_{J S} : P \times P \to [0, \infty)

is defined between two Radon–Nikodym densities p and q of μ-absolutely continuous probability measures by

\begin{matrix} D_{J S} [p ∥ q] & : = \frac{1}{2} (D_{K L} [p ∥ \frac{p + q}{2}] + D_{K L} [q ∥ \frac{p + q}{2}]) \\ = \frac{1}{2} \int_{X} (p ln \frac{2 p}{p + q} + q ln \frac{2 q}{p + q}) d μ \\ = D_{J S} [q ∥ p] . \end{matrix}

(4)

The JS-divergence is a symmetrized and smoothed version of the KL-divergence, and it is bounded as

0 \leq D_{J S} [p ∥ q] \leq ln 2 .

(5)

This property contrasts with the fact that KL-divergence is unbounded.

Definition 3

(Jeffreys divergence [11]). The Jeffreys divergence

D_{J} [p ∥ q] : P \times P \to [0, \infty]

is defined between two Radon–Nikodym densities p and q of μ-absolutely continuous probability measures by

D_{J} [p ∥ q] : = D_{K L} [p ∥ q] + D_{K L} [q ∥ p] .

(6)

Such symmetrized KL-divergences have appeared in various pieces of literature [12,13,14,15,16,17,18].

For continuous distributions, the KL-divergence is known to have computational difficulty. To be more specific, if q takes a small value relative to p, the value of

D_{K L} [p ∥ q]

may diverge to infinity. The simplest idea to avoid this is to use very small

ϵ > 0

and modify

D_{K L} [p ∥ q]

as follows:

D_{K L}^{+} [p ∥ q] : = \int_{X} p ln \frac{p}{q + ϵ} d μ .

However, such an extension is unnatural in the sense that

q + ϵ

no longer satisfies the condition for a probability measure:

\int_{X} ϵ + q (x) d μ (x) \neq 1

. As a more natural way to stabilize KL-divergence, the following skew divergences have been proposed:

Definition 4

(Skew divergence [8,19]). The skew divergence

D_{S}^{(λ)} [p ∥ q] : P \times P \to [0, \infty]

is defined between two Radon–Nikodym densities p and q of μ-absolutely continuous probability measures by

\begin{matrix} D_{S}^{(λ)} [p ∥ q] & : = D_{K L} [p ∥ (1 - λ) p + λ q] \\ = \int_{X} p ln \frac{p}{(1 - λ) p + λ q} d μ, \end{matrix}

(7)

where

λ \in [0, 1]

.

Skew divergences have been experimentally shown to perform better in applications such as natural language processing [20,21], image recognition [22,23] and graph analysis [24,25]. In addition, there is research on quantum generalization of skew divergence [26].

The main contributions of this paper are summarized as follows:

Several symmetrized divergences or skew divergences are generalized from an information geometry perspective.
It is proved that the natural skew divergence for the exponential family is equivalent to the scaled KL-divergence.
Several properties of geometrically generalized skew divergence are proved. Specifically, the functional space associated with the proposed divergence is shown to be a Banach space.

Implementation of the proposed divergence is available on GitHub (https://github.com/nocotan/geodesical_skew_divergence (accessed on 3 April 2021)).

2. α-Geodesical Skew Divergence

The skew divergence is generalized based on the following function.

Definition 5

(f-interpolation). For any

a, b, \in R

,

λ \in [0, 1]

and

α \in R

, f-interpolation is defined as

m_{f}^{(λ, α)} (a, b) = f_{α}^{- 1} ((1 - λ) f_{α} (a) + λ f_{α} (b)),

(8)

where

f_{α} (x) = \{\begin{matrix} x^{\frac{1 - α}{2}} & (α \neq 1) \\ ln x & (α = 1) \end{matrix}

(9)

is the function that defines the f-mean [27].

The f-mean function satisfies

\begin{matrix} lim_{α \to \infty} f_{α} (x) & = \{\begin{matrix} \infty & (| x | < 1), \\ 1 & (| x | = 1), \\ 0 & (| x | > 1), \end{matrix} \\ lim_{α \to - \infty} f_{α} (x) & = \{\begin{matrix} 0 & (| x | < 1), \\ 1 & (| x | = 1), \\ \infty & (| x | > 1) . \end{matrix} \end{matrix}

It is easy to see that this family includes various known weighted means including the e-mixture and m-mixture for

α = \pm 1

in the literature of information geometry [28]:

\begin{matrix} (α = 1) & m_{f}^{(λ, 1)} (a, b) = exp {(1 - λ) ln a + λ ln b} \\ (α = - 1) & m_{f}^{(λ, - 1)} (a, b) = (1 - λ) a + λ b \\ (α = 0) & m_{f}^{(λ, 0)} (a, b) = {((1 - λ) \sqrt{a} + λ \sqrt{b})}^{2} \\ (α = 3) & m_{f}^{(λ, 3)} (a, b) = \frac{1}{(1 - λ) \frac{1}{a} + λ \frac{1}{b}} \\ (α = \infty) & m_{f}^{(λ, \infty)} (a, b) = min {a, b} \\ (α = - \infty) & m_{f}^{(λ, - \infty)} (a, b) = max {a, b} \end{matrix}

The inverse function

f_{α}^{- 1}

is convex when

α \in [- 1, 1]

, and concave when

α \in (- \infty, - 1] \cup (1, \infty)

. It is worth noting that the f-interpolation is a special case of the Kolmogorov–Nagumo average [29,30,31] when

α

is restricted in the interval

[- 1, 1]

.

In order to consider the geometric meaning of this function, the notion of the statistical manifold is introduced.

2.1. Statistical Manifold

Let

S = {p_{ξ} = p (x; ξ) \in P | ξ = (ξ^{1}, \dots, ξ^{n}) \in Ξ}

(10)

be a family of probability distribution on

X

, where each element

p_{ξ}

is parameterized by n real-valued variables

ξ = (ξ^{1}, \dots, ξ^{n}) \in Ξ \subset R^{n}

. The set

S

is called a statistical model and is a subset of

P

. We also denote

(S, g_{i j})

as a statistical model equipped with the Riemannian metric

g_{i j}

. In particular, let

g_{i j}

be the Fisher–Rao metric, which is the Riemannian metric induced from the Fisher information matrix [32].

In the rest of this paper, the abbreviations

\begin{matrix} \partial_{i} & = \partial_{ξ^{i}} = \frac{\partial}{\partial ξ^{i}}, \\ ℓ & = ℓ_{x} (ξ) = ln p_{ξ} (x) \end{matrix}

are used.

Definition 6

(Christoffel symbols). Let

g_{i j}

be a Riemannian metric, particularly the Fisher information matrix, then the Christoffel symbols are given by

Γ_{i j, k} = \frac{1}{2} (\partial_{i} g_{j k} + \partial_{j} g_{i k} - \partial_{k} g_{i j}), i, j, k = 1, \dots, n .

(11)

Definition 7

(Levi-Civita connection). Let g be a Fisher–Riemannian metric on

S

which is a 2-covariant tensor defined locally by

g (X_{ξ}, Y_{ξ}) = \sum_{i, j = 1}^{n} g_{i j} (ξ) a^{i} (ξ) b^{j} (ξ),

where

X_{ξ} = \sum_{i = 1}^{n} a^{i} (ξ) \partial_{i} p_{ξ}

and

Y_{ξ} = \sum_{i = 1}^{n} b^{i} (ξ) \partial_{i} p_{ξ}

are vector fields in the 0-representation on

S

. Then, its associated Levi-Civita connection

\nabla^{(0)}

is defined by

g (\nabla_{\partial_{i}}^{(0)} \partial_{j}, \partial_{k}) = Γ_{i j, k} .

(12)

The fact that

\nabla^{(0)}

is a metrical connection can be written locally as

\partial_{k} g_{i j} = Γ_{k i, j} + Γ_{k j, i} .

(13)

It is worth noting that the superscript

α

of

\nabla^{(α)}

corresponds to a parameter of the connection. Based on the above definitions, several connections parameterized by the parameter

α

are introduced. The case

α = 0

corresponds to the Levi-Civita connection induced by the Fisher metric.

Definition 8

(

\nabla^{(1)}

-connection). Let g be the Fisher-Riemannian metric on

S

, which is a 2-covariant tensor. Then, the

\nabla^{(1)}

-connection is defined by

g (\nabla_{\partial_{i}}^{(1)} \partial_{j}, \partial_{k}) = E_{ξ} [\partial_{i} \partial_{j} ℓ \partial_{k} ℓ] .

(14)

It can also be expressed equivalently by explicitly writing as the Christoffel coefficients

Γ_{i j, k}^{(1)} (ξ) = E_{ξ} [\partial_{i} \partial_{j} ℓ \partial_{k} ℓ] .

(15)

Definition 9

(

\nabla^{(- 1)}

-connection). Let g be the Fisher–Riemannian metric on

S

, which is a 2-covariant tensor. Then, the

\nabla^{(- 1)}

-connection is defined by

g (\nabla_{\partial_{i}}^{(- 1)} \partial_{j}, \partial_{k}) = Γ_{i j, k}^{(- 1)} (ξ) = E_{ξ} [(\partial_{i} \partial_{j} ℓ + \partial_{i} ℓ \partial_{j} ℓ) \partial_{k} ℓ] .

(16)

In the following, the ∇-flatness is considered with respect to the corresponding coordinates system. More details can be found in [28].

Proposition 1.

The exponential family is

\nabla^{(1)}

-flat.

Proposition 2.

The exponential family is

\nabla^{(- 1)}

-flat if and only if it is

\nabla^{(0)}

-flat.

Proposition 3.

The mixture family is

\nabla^{(- 1)}

-flat.

Proposition 4.

The mixture family is

\nabla^{(1)}

-flat if and only if it is

\nabla^{(0)}

-flat.

Proposition 5.

The relation between the foregoing three connections is given by

\nabla^{(0)} = \frac{1}{2} (\nabla^{(- 1)} + \nabla^{(1)}) .

(17)

Proof.

It suffices to show

Γ_{i j, k}^{(0)} = \frac{1}{2} (Γ_{i j, k}^{(- 1)} + Γ_{i j, k}^{(1)}) .

From the definitions of

Γ^{(- 1)}

and

Γ^{(1)}

,

\begin{matrix} Γ_{i j, k}^{(- 1)} + Γ_{i j, k}^{(1)} & = E_{ξ} [(\partial_{i} \partial_{j} ℓ + \partial_{i} ℓ \partial_{j} ℓ) \partial_{k} ℓ] + E_{ξ} [\partial_{i} \partial_{j} ℓ \partial_{k} ℓ] \\ = E_{ξ} [(2 \partial_{i} \partial_{j} ℓ + \partial_{i} ℓ \partial_{j} ℓ) \partial_{k} ℓ] \\ = 2 E_{ξ} [(\partial_{i} \partial_{j} ℓ + \frac{1}{2} \partial_{i} ℓ \partial_{j} ℓ) \partial_{k} ℓ] \\ = 2 Γ_{i j, k}^{(0)}, \end{matrix}

which proves the proposition. □

The connections

\nabla^{(- 1)}

and

\nabla^{(1)}

are two special connections on

S

with respect to the mixture family and the exponential family, respectively. Moreover, they are related by the duality condition, and the following 1-parameter family of connections is defined.

Definition 10

(

\nabla^{(α)}

-connection). For

α \in R

, the

\nabla^{(α)}

-connection on the statistical model

S

is defined as

\nabla^{(α)} = \frac{1 + α}{2} \nabla^{(1)} + \frac{1 - α}{2} \nabla^{(- 1)} .

(18)

Proposition 6.

The components

Γ_{i j, k}^{(α)}

can be written as

Γ_{i j, k}^{(α)} = E_{ξ} [(\partial_{i} \partial_{j} ℓ + \frac{1 - α}{2} \partial_{i} ℓ \partial_{j} ℓ) \partial_{k} ℓ] .

(19)

The

α

-coordinate system associated with the

\nabla^{(α)}

-connection is endowed with the

α

-geodesic, which is a straight line on the corresponding coordinates system. Then, we introduce some relevant notions.

Definition 11

(α-divergence [33]). Let α be a real parameter. The α-divergence between two probability vectors

p

and

q

is defined as

D_{α} [p ∥ q] = \frac{4}{1 - α^{2}} (1 - \sum_{i} p_{i}^{\frac{1 - α}{2}} q_{i}^{\frac{1 + α}{2}}) .

(20)

The KL-divergence, which is a special case with

α = 1

, induces the linear connection

\nabla^{(1)}

as follows.

Proposition 7.

The diagonal part of the third mixed derivatives of the KL-divergence is the negative of the Christoffel symbol:

- \partial_{ξ^{i}} \partial_{ξ^{j}} \partial_{ξ_{0}^{k}} D_{K L} [p_{ξ_{0}} ∥ p_{ξ}] |_{ξ = ξ_{0}} = Γ_{i j, k}^{(1)} (ξ_{0}) .

(21)

Proof.

The second derivative in the argument

ξ

is given by

\partial_{ξ^{i}} \partial_{ξ^{j}} D_{K L} [p_{ξ_{0}} ∥ p_{ξ}] = - \int_{X} p_{ξ_{0}} (x) \partial_{ξ^{i}} \partial_{ξ^{j}} ℓ_{x} (ξ) d x,

and differentiating it with respect to

ξ_{0}^{k}

yields

\begin{matrix} - \partial_{ξ^{i}} \partial_{ξ^{j}} \partial_{ξ_{0}^{k}} D_{K L} [p_{ξ_{0}} ∥ p_{ξ}] & = \partial_{ξ_{0}^{k}} \int_{X} p_{ξ_{0}} (x) \partial_{ξ^{i}} \partial_{ξ^{j}} ℓ_{x} (ξ) d x \\ = \int_{X} p_{ξ_{0}} (x) \partial_{ξ^{i}} \partial_{ξ^{j}} ℓ_{x} (ξ) \partial_{ξ_{0}^{k}} ℓ_{x} (ξ) d x . \end{matrix}

Then, considering the diagonal part, one yields

\begin{matrix} - \partial_{ξ^{i}} \partial_{ξ^{j}} \partial_{ξ_{0}^{k}} D_{K L} [p_{ξ_{0}} ∥ p_{ξ}] |_{ξ = ξ_{0}} & = E_{ξ_{0}} [\partial_{i} \partial_{j} ℓ (ξ) \partial_{k} ℓ (ξ)] \\ = Γ_{i j, k}^{(1)} (ξ_{0}) . \end{matrix}

□

More generally, the

α

-divergence with

α \in R

induces the

\nabla^{(α)}

-connection.

Definition 12

(α-representation [34]). For some positive measure

m_{i}^{\frac{1 - α}{2}}

, the coordinate system

θ = (θ^{i})

derived from the α-divergence is

θ^{i} = m_{i}^{\frac{1 - α}{2}} = f_{α} (m_{i})

(22)

and

θ^{i}

is called the α-representation of a positive measure

m_{i}^{\frac{1 - α}{2}}

.

Definition 13

(α-geodesic [28]). The α-geodesic connecting two probability vectors

p (x)

and

q (x)

is defined as

r_{i} (t) = c (t) f_{α}^{- 1} \{(1 - t) f_{α} (p (x_{i})) + t f_{α} (q (x_{i}))\}, t \in [0, 1]

(23)

where

c (t)

is determined as

c (t) = \frac{1}{\sum_{i = 1}^{n} r_{i} (t)} .

(24)

It is known that the appropriate reparameterizations for the parameter t are necessary for a rigorous discussion in the space of probability measures [35,36]. However, as mentioned in the literature [35], an explicit expression for the reparametrizations

τ_{p, a}

and

τ_{p, q}

is unknown. A similar discussion has been made in the derivation of the

ϕ_{β}

-path [37], where it is mentioned that the normalizing factor is unknown in general. Furthermore, the f-mean is not convex depending on the

α

. For these reasons, it is generally difficult to discuss

α

-geodesics in probability measures by normalization or reparameterization, and to avoid unnecessary complexity, the parameter t is assumed to be appropriately reparameterized.

Let

ψ_{α} (θ) = \frac{1 - α}{2} \sum_{i = 1}^{n} m_{i}

. Then, the dual coordinate system

η

is given by

η = \nabla ψ_{α} (θ)

as

η_{i} = {(θ^{i})}^{\frac{1 + α}{1 - α}} = f_{- α} (m_{i}) .

(25)

Hence, it is the (

- α

)-representation of

m_{i}

.

2.2. Generalization of Skew Divergences

From Definition 13, the f-interpoloation is considered as an unnormalized version of the

α

-geodesic. Using the notion of geodesics, skew divergence is generalized in terms of information geometry as follows.

Definition 14

(α-Geodesical Skew Divergence). The α-geodesical skew divergence

D_{G S}^{(α, λ)} : P \times P \to [0, \infty]

is defined between two Radon–Nikodym densities p and q of μ-absolutely continuous probability measures by:

\begin{matrix} D_{G S}^{(α, λ)} [p ∥ q] & : = D_{K L} [p ∥ m_{f}^{(λ, α)} (p, q)] \\ = \int_{X} p ln \frac{p}{m_{f}^{(λ, α)} (p, q)} d μ, \end{matrix}

(26)

where

α \in R

and

λ \in [0, 1]

.

Some special cases of

α

-geodesical skew divergence are listed below:

\begin{matrix} (\forall α \in R, λ = 1) & D_{G S}^{(α, 1)} [p ∥ q] = D_{K L} [p ∥ q] \\ (\forall α \in R, λ = 0) & D_{G S}^{(α, 0)} [p ∥ q] = D_{K L} [p ∥ p] = 0 \\ (α = 1, \forall λ \in [0, 1]) & D_{G S}^{(1, λ)} [p ∥ q] = λ D_{K L} [p ∥ q] (scaled KL-divergence) \\ (α = - 1, \forall λ \in [0, 1]) & D_{G S}^{(- 1, λ)} [p ∥ q] = D_{S}^{(λ)} [p ∥ q] (skew divergence) \\ (α = 0, \forall λ \in [0, 1]) & D_{G S}^{(0, λ)} [p ∥ q] = \int_{X} p ln \frac{p}{{(1 - λ) \sqrt{p} + λ \sqrt{q}}^{2}} d μ \\ (α = 3, \forall λ \in [0, 1]) & D_{G S}^{(3, λ)} [p ∥ q] = D_{S}^{(λ)} [p ∥ q] + H (p) + H (q) \\ (α = \infty, \forall λ \in [0, 1]) & D_{G S}^{(\infty, λ)} [p ∥ q] = \int_{X} p ln \frac{p}{min {p, q}} d μ \\ (α = - \infty, \forall λ \in [0, 1]) & D_{G S}^{(- \infty, λ)} [p ∥ q] = \int_{X} p ln \frac{p}{max {p, q}} d μ \end{matrix}

Furthermore,

α

-geodesical skew divergence is a special form of the generalized skew K-divergence [10,38], which is a family of abstract means-based divergences. In this paper, the skew K-divergence touched upon in [10] is characterized in terms of

α

-geodesic on positive measures, and its geometric and functional analytic properties are investigated. When the Kolmogorov–Nagumo average (i.e., when the function

f^{- 1}

in Equation (8) is a strictly monotone convex function) the geodesic has been shown to be well-defined [37].

2.3. Symmetrization of $α$ -Geodesical Skew Divergence

It is easy to symmetrize the

α

-geodesical skew divergence as follows.

Definition 15

(Symmetrized α-Geodesical Skew Divergence). The symmetrized α-geodesical skew divergence

{\bar{D}}_{G S}^{(α, λ)} : P \times P \to [0, \infty]

is defined between two Radon–Nikodym densities p and q of μ-absolutely continuous probability measures by:

\begin{matrix} {\bar{D}}_{G S}^{(α, λ)} [p ∥ q] & : = \frac{1}{2} (D_{G S}^{(α, λ)} [p ∥ q] + D_{G S}^{(α, λ)} [q ∥ p]), \end{matrix}

(27)

where

α \in R

and

λ \in [0, 1]

.

It is seen that

{\bar{D}}_{G S}^{(α, λ)} [p ∥ q]

includes several symmetrized divergences.

\begin{matrix} {\bar{D}}_{G S}^{(α, 1)} [p ∥ q] & = \frac{1}{2} (D_{K L} [p ∥ q] + D_{K L} [q ∥ p]), (half of Jeffreys divergence) \\ {\bar{D}}_{G S}^{(- 1, \frac{1}{2})} [p ∥ q] & = \frac{1}{2} (D_{K L} [p ∥ \frac{p + q}{2}] + D_{K L} [q ∥ \frac{p + q}{2}]), (JS-divergence) \\ {\bar{D}}_{G S}^{(- 1, λ)} [p ∥ q] & = \frac{1}{2} (D_{K L} [p ∥ (1 - λ) p + λ q] + D_{K L} [q ∥ (1 - λ) q + λ p]) . \end{matrix}

The last one is the

λ

-JS-divergence [39], which is a generalization of the JS-divergence.

3. Properties of α-Geodesical Skew Divergence

In this section, the properties of the

α

-geodesical skew divergence are studied.

Proposition 8

(Non-negativity of the α-geodesical skew divergence). For

α \geq - 1

and

λ \in [0, 1]

, the α-geodesical skew divergence

D_{G S}^{(α, λ)} [p ∥ q]

satisfies the following inequality:

D_{G S}^{(α, λ)} [p ∥ q] \geq 0 .

(28)

Proof.

When λ is fixed, the f-interpolation has the following inverse monotonicity with respect to α:

m_{f}^{(λ, α)} (p, q) \geq m_{f}^{(λ, α^{'})} (p, q), (α \leq α^{'}) .

(29)

From Gibbs’ inequality [40] and Equation (29), one obtains

\begin{matrix} D_{G S}^{(α, λ)} [p ∥ q] & = \int_{X} p ln \frac{p}{m_{f}^{(α, λ)} (p, q)} d μ \\ \geq (\int_{X} p d μ) ln \frac{p}{m_{f}^{(α, λ)} (p, q)} \\ \geq 1 \cdot ln 1 = 0 . \end{matrix}

□

Proposition 9

(Asymmetry of the α-geodesical skew divergence). α-Geodesical skew divergence is not symmetric in general:

D_{G S}^{(α, λ)} [p | q] \neq D_{G S}^{(α, λ)} [q ∥ p] .

(30)

Proof.

For example, if

λ = 1

, then

\forall α \in R

, it holds that

\begin{matrix} D_{G S}^{(α, 1)} [p ∥ q] - D_{G S}^{(α, 1)} [q ∥ p] & = D_{K L} [p ∥ q] - D_{K L} [q ∥ p], \end{matrix}

and the asymmetry of the KL-divergence results in an asymmetry of the geodesic skew divergence. □

When a function

f (x)

of

x \in [0, 1]

satisfies

f (x) = f (1 - x)

, it is referred to as centrosymmetric.

Proposition 10

(Non-centrosymmetricity of the α-geodesical skew divergence with respect to λ). α-Geodesical skew divergence is not centrosymmetric in general with respect to the parameter

λ \in [0, 1]

:

D_{G S}^{(α, λ)} [p ∥ q] \neq D_{G S}^{(α, 1 - λ)} [p ∥ q] .

(31)

Proof.

For example, if

λ = 1

, then

\forall α \in R

, we have

\begin{matrix} D_{G S}^{(α, λ)} [p ∥ q] - D_{G S}^{(α, 1 - λ)} [p ∥ q] & = D_{G S}^{(α, 1)} [p ∥ q] - D_{G S}^{(α, 0)} [p ∥ q] \\ = \int_{X} p ln \frac{p}{q} - \int_{X} p ln \frac{p}{p} \\ = \int_{X} p ln \frac{p}{q} \geq 0 . \end{matrix}

(32)

□

Proposition 11

(Monotonicity of the α-geodesical skew divergence with respect to α). α-Geodesical skew divergence satisfies the following inequality for all

α \in R, λ \in [0, 1]

.

D_{G S}^{(α, λ)} [p ∥ q] \geq D_{G S}^{(α^{'}, λ)} [p ∥ q], (α \geq α^{'}) .

Proof.

Obvious from the inverse monotonicity of the f-interpolation (29) and the monotonicity of the logarithmic function. □

Figure 1 shows the monotonicity of the geodesic skew divergence with respect to

α

. In this figure, divergence is calculated between two binomial distributions.

Proposition 12

(Subadditivity of the α-geodesical skew divergence with respect to α). α-Geodesical skew divergence satisfies the following inequality for all

α, β \in R, λ \in [0, 1]

D_{G S}^{(α + β, λ)} [p ∥ q] \leq D_{G S}^{(α, λ)} [p ∥ q] + D_{G S}^{(β, λ)} [p ∥ q] .

Proof.

For some

α

and

λ

,

m_{f}^{(λ, α)}

takes the form of the Kolmogorov mean [29], which is obvious from its continuity, monotonicity and self-distributivity. □

Proposition 13

(Continuity of the α-geodesical skew divergence with respect to α and λ). α-Geodesical skew divergence has the continuity property.

Proof.

We can prove from the continuity of the KL-divergence and the Kolmogorov mean. □

Figure 2 shows the continuity of the geodesic skew divergence with respect to

α

and

λ

. Both source and target distributions are binomial distributions. From this figure, it can be seen that the divergence changes smoothly as the parameters change.

Lemma 1.

Suppose

α \to \infty

. Then,

lim_{α \to \infty} D_{G S}^{(α, λ)} [p ∥ q] = \int_{X} p ln \frac{p}{min {p, q}} d μ

(33)

holds for all

λ \in [0, 1]

.

Proof.

Let

u = \frac{1 - α}{2}

. Then

{lim}_{α \to \infty} u = - \infty

. Assuming

p_{0} \leq p_{1}

, it holds that

\begin{matrix} lim_{α \to \infty} m_{f}^{(λ, α)} (p_{0}, p_{1}) & = lim_{u \to - \infty} {((1 - λ) p_{0}^{u} + λ p_{1}^{u})}^{\frac{1}{u}} \\ = p_{0} lim_{u \to - \infty} {((1 - λ) + λ {(\frac{p_{1}}{p_{0}})}^{u})}^{\frac{1}{u}} \\ = p_{0} = min {p_{0}, p_{1}} . \end{matrix}

Then, the following equality

\begin{matrix} lim_{α \to \infty} D_{G S}^{(α, λ)} [p ∥ q] & = \int_{X} p ln \frac{p}{{lim}_{α \to \infty} m_{f}^{(λ, α)} (p_{0}, p_{1})} d μ \\ = \int_{X} p ln \frac{p}{min {p, q}} d μ \end{matrix}

holds. □

Lemma 2.

Suppose

α \to - \infty

. Then,

lim_{α \to \infty} D_{G S}^{(α, λ)} [p ∥ q] = \int_{X} p ln \frac{p}{max {p, q}} d μ

(34)

holds for all

λ \in [0, 1]

.

Proof.

Let

u = \frac{1 - α}{2}

. Then

{lim}_{α \to - \infty} u = \infty

. Assuming

p_{0} \leq p_{1}

, it holds that

\begin{matrix} lim_{α \to \infty} m_{f}^{(λ, α)} (p_{0}, p_{1}) & = lim_{u \to - \infty} {((1 - λ) p_{0}^{u} + λ p_{1}^{u})}^{\frac{1}{u}} \\ = p_{1} lim_{u \to - \infty} {((1 - λ) {(\frac{p_{0}}{p_{1}})}^{u} + λ)}^{\frac{1}{u}} \\ = p_{1} = max {p_{0}, p_{1}} . \end{matrix}

Then, the following equality

\begin{matrix} lim_{α \to - \infty} D_{G S}^{(α, λ)} [p ∥ q] & = \int_{X} p ln \frac{p}{{lim}_{α \to - \infty} m_{f}^{(λ, α)} (p_{0}, p_{1})} d μ \\ = \int_{X} p ln \frac{p}{max {p, q}} d μ \end{matrix}

holds. □

Proposition 14

(Lower bound of the α-geodesical skew divergence). α-Geodesical skew divergence satisfies the following inequality for all

α \in R, λ \in [0, 1]

.

D_{G S}^{(α, λ)} [p ∥ q] \geq \int_{X} p ln \frac{p}{max {p, q}} d μ .

(35)

Proof.

It follows from the definition of the inverse monotonicity of f-interpolation (29) and Lemma 2. □

Proposition 15

(Upper bound of the α-geodesical skew divergence). α-Geodesical skew divergence satisfies the following inequality for all

α \in R, λ \in [0, 1]

.

D_{G S}^{(α, λ)} [p ∥ q] \leq \int_{X} p ln \frac{p}{min {p, q}} d μ .

(36)

Proof.

It follows from the definition of the f-interpolation (29) and Lemma 1. □

Theorem 1

(Strong convexity of the α-geodesical skew divergence). α-Geodesical skew divergence

D_{G S}^{(α, λ)} [p ∥ q]

is strongly convex in p with respect to the total variation norm.

Proof.

Let

r : = m_{f}^{(α, λ)} (p, q)

and

f_{j} : = \frac{p_{j}}{r} (j = 0, 1)

, so that

f_{t} = \frac{p_{t}}{r} (t \in (0, 1))

. From Taylor’s theorem, for

g (x) : = x ln x

and

j = 0, 1

, it holds that

\begin{matrix} g (f_{j}) = g (f_{t}) + g^{'} (f_{t}) (f_{j} - f_{t}) + {(f_{j} - f_{t})}^{2} \int_{0}^{1} g^{″} ((1 - s) f_{t} + s f_{j}) (1 - s) d s . \end{matrix}

Let

\begin{matrix} δ & : = (1 - t) g (f_{0}) + t g (f_{1}) - g (f_{t}) \\ = (1 - t) t {(f_{1} - f_{0})}^{2} \int_{0}^{1} (\frac{t}{(1 - s) f_{t} + s f_{0}} + \frac{1 - t}{(1 - s) f_{t} + s f_{1}}) (1 - s) d s \\ = (1 - t) t {(f_{1} - f_{0})}^{2} \int_{0}^{1} (\frac{t}{f_{u_{0}} (t, s)} + \frac{1 - t}{f_{u_{1}} (t, s)}) (1 - s) d s, \end{matrix}

where

\begin{matrix} u_{j} (t, s) : = & (1 - s) t + j t, \\ f_{μ_{j}} (t, s) : = & (1 - s) f_{t} + s f_{j} . \end{matrix}

Then,

\begin{matrix} Δ & : = (1 - t) H (p_{0}) + t H (p_{1}) - H (p_{t}) \\ = \int δ d r \\ = (1 - t) t \int_{0}^{1} (1 - s) d s [t I (u_{0} (t, s)) + (1 - t) I (u_{1} (t, s))], \end{matrix}

where

\begin{matrix} ∥ p_{1} - p_{0} ∥ : = & \int | d p_{1} - d p_{0} | d μ, \\ H (p) : = & D_{G S}^{(α, λ)} [p ∥ r] = \int p ln \frac{p}{r} d μ, \\ I (u) : = & \int \frac{{(f_{1} - f_{0})}^{2}}{f_{u}} d r . \end{matrix}

Now, it is suffice to prove that

Δ \geq \frac{t (1 - t)}{2} {∥ p_{1} - p_{0} ∥}^{2}

. For all

u \in (0, 1)

, it is seen that

p_{1}

is absolutely continuous with respect to

p_{u}

. Let

g_{u} : = \frac{p_{1}}{p_{u}} = \frac{f_{1}}{f_{u}}

. One obtains

\begin{matrix} I (u) & = \frac{1}{{(1 - u)}^{2}} \int \frac{{(f_{1} - f_{u})}^{2}}{f_{u}} d r \\ = \frac{1}{{(1 - u)}^{2}} \int {(g_{u} - 1)}^{2} d p_{u} \\ \geq \frac{1}{{(1 - u)}^{2}} {(\int | g_{u} - 1 | d p_{u})}^{2} \\ = \frac{1}{{(1 - u)}^{2}} ∥ p_{1} - p_{u} ∥^{2} = {∥ p_{1} - p_{0} ∥}^{2}, \end{matrix}

and hence, for

j = 0, 1

,

Δ \geq \frac{t (1 - t)}{2} {∥ p_{1} - p_{0} ∥}^{2} .

□

4. Natural α-Geodesical Skew Divergence for Exponential Family

In this section, the exponential family is considered in which the probability density function is given by

p (x; θ) = exp \{θ \cdot x + k (x) - ψ (θ)\},

(37)

where

x

is a random variable. In the above equation,

θ = (θ^{1}, \dots, θ^{n})

is an n-dimensional vector parameter to specify distribution,

k (x)

is a function of

x

and

ψ

corresponds to the normalization factor.

In skew divergence, the probability distribution of the target is a weighted average of the two distributions. This implicitly assumes that interpolation of the two probability distributions is properly given by linear interpolation. Here, in the exponential family, the interpolation between natural parameters rather than interpolation between probability distributions themselves is considered. Namely, the geodesic connecting two distributions

p (x; θ_{p})

and

q (x; θ_{q})

on the

θ

-coordinate system is considered:

θ (λ) = (1 - λ) θ_{p} + λ θ_{q},

(38)

where

λ \in [0, 1]

is the parameter. The probability distributions on the geodesic

θ (λ)

are

\begin{matrix} p (x; λ) & = p (x; θ (λ)) \\ = exp \{λ (θ_{q} - θ_{p}) \cdot x + θ_{p} \cdot x - ψ (λ)\} . \end{matrix}

(39)

Hence, a geodesic itself is a one-dimensional exponential family, where

λ

is the natural parameter. A geodesic consists of a linear interpolation of the two distributions in the logarithmic scale because

ln p (x; λ) = (1 - λ) ln p (x; θ_{p}) + λ ln p (x; θ_{q}) - ψ (λ) .

(40)

This corresponds to the case

α = 1

on the f-interpolation with normalization factor

c (λ) = exp {- ψ (λ)}

,

p (x; θ (λ)) = m_{f}^{(λ, 1)} (p (x; θ_{p}), p (x; θ_{q})) .

(41)

This induces the natural geodesic skew divergence with

α = 1

as

\begin{matrix} D_{G S}^{(1, λ)} [p ∥ q] & = \int_{X} p ln (\frac{p}{m_{f}^{(λ, 1)} (p, q)}) d μ \\ = \int_{X} p ln p - p ln (m_{f}^{(λ, 1)} (p, q)) d μ \\ = \int_{X} p ln p - p ln (exp {(1 - λ) ln p + λ ln q}) d μ \\ = \int_{X} (p ln p - (1 - λ) p ln p - λ p ln q) d μ \\ = \int_{X} (λ p ln p - λ p ln q) d μ \\ = λ \int_{X} p ln \frac{p}{q} d μ \\ = λ D_{K L} [p ∥ q], \end{matrix}

and this is equal to the scaled KL divergence.

More generally, let

θ_{P}^{(α)}

and

θ_{Q}^{(α)}

be the parameter representations on the

α

-coordinate system of probability distributions P and Q. Then, the geodesics between them are represented as in Figure 3, and it induces the

α

-geodesical skew divergence.

5. Function Space Associated with the α-Geodesical Skew Divergence

To discuss the functional nature of the

α

-geodesical skew divergence in more depth, the function space it constitutes is considered. For an

α

-geodesical skew divergence

f_{q}^{(α, λ)} (p) = D_{G S}^{(α, λ)} [p ∥ q]

with one side of the distribution fixed, let the entire set be

F_{q} = \{f_{q}^{(α, λ)} ∣ α \in R, λ \in [0, 1]\} .

(42)

For

f_{q}^{(α, λ)} \in F_{q}

, its semi-norm is defined by

∥ f_{q}^{(α, λ)} ∥_{p} : = \int_{X} {(| f_{q}^{(α, λ)} |^{p} d μ)}^{\frac{1}{p}} .

(43)

By defining addition and scalar multiplication for

f_{q}^{(α, λ)}, g_{q}^{(α, λ)} \in F_{q}

,

c \in R

as follows,

F_{q}

becomes a semi-norm vector space:

\begin{matrix} (f_{q}^{(α, λ)} + g_{q}^{(α, λ)}) (u) : = & f_{q}^{(α, λ)} (u) + g_{q}^{(α, λ)} (u) = D_{G S}^{(α, λ)} [u ∥ q] + D_{G S}^{(α^{'}, λ^{'})} [u ∥ q], \end{matrix}

(44)

\begin{matrix} (c f) (u) : = & c f_{q}^{(α, λ)} (u) = c \cdot D_{G S}^{(α, λ)} [u ∥ q] . \end{matrix}

(45)

Theorem 2.

Let

N

be the kernel of

{∥ \cdot ∥}_{p}

as follows:

{N k e r (∥ \cdot ∥}_{p}) = \{f_{q}^{(α, λ)} ∣ f_{q}^{(α, λ)} = 0\} .

(46)

Then the quotient space

V (F_{q}, ∥ \cdot ∥_{p}) / N

is a Banach space.

Proof.

It is sufficient to prove that

f_{q}^{(α, λ)}

is integrable to the power of p and that

V

is complete. From Proposition 15, the

α

-geodesical skew divergence is bounded from above for all

α \in R

and

λ \in [0, 1]

. Since

f_{q}^{(α, λ)}

is continuous, we know that it is p-power integrable.

Let

{f_{n}}

be a Cauchy sequence of

V

:

lim_{n, m \to \infty} {∥ f_{n} - f_{m} ∥}_{p} = 0 .

As

n (k), k = 1, 2, \dots,

can be taken to be monotonically increasing and

∥ f_{n} - f_{n (k)} ∥_{p} < 2^{- k}

with respect to

n > n (k)

, let

∥ f_{n (k + 1)} - f_{n (k)} ∥_{p} < 2^{- k} .

If

g_{n} = | f_{n (1)} | + \sum_{j = 1}^{n - 1} | f_{n (j + 1)} - f_{n (j)} | \in V

, it is non-negatively monotonically increasing at each point, and from the subadditivity of the norm,

∥ g_{n} ∥_{p} \leq {∥ f_{n (1)} ∥}_{p} + \sum_{j = 1}^{n - 1} 2^{- j}

. From the monotonic convergence theorem, we have

∥ lim_{n \to \infty} g_{n} ∥_{p} = lim_{n \to \infty} ∥ g_{n} ∥_{p} \leq {∥ f_{n (1)} ∥}_{p} + 1 < \infty .

That is,

{lim}_{n \to \infty} g_{n}

exists almost everywhere, and

{lim}_{n \to \infty} g_{n} \in V

. From

{lim}_{n \to \infty} g_{n} < \infty

, we have

f_{n (1)} + \sum_{j = 1}^{n - 1} (f_{n (j + 1)} - f_{n (j)}) = lim_{n \to \infty} f_{n (1)}

converges absolutely almost everywhere to

| {lim}_{n \to \infty} f_{n (n)} | \leq {lim}_{n \to \infty} g_{n}, a . e .

That is,

{lim}_{n \to \infty} f_{n (n)} \in V

. Then

| lim_{n \to \infty} f_{n} - f_{n (n)} | \leq lim_{n \to \infty} g_{n}

and from the superior convergence theorem, we can obtain

lim_{n \to \infty} ∥ lim_{n \to \infty} f_{n} - f_{n (n)} ∥_{p} = 0

We have now confirmed the completeness of

V

. □

Corollary 1.

Let

F_{+} = \{f_{q}^{(α, λ)} ∣ α \in R, λ \in (0, 1], q \in P\} .

(47)

Then the space

V_{+} : = (F_{+} {, ∥ \cdot ∥}_{p})

is a Banach space.

Proof.

If we restrict

λ \in (0, 1]

,

D_{G S}^{(α, λ)} [u ∥ q] = 0

if and only if

u = q

. Then,

V_{+}

has the unique identity element, and then

V_{+}

is a complete norm space. □

Consider the second argument Q of

D_{G S}^{(α, λ)} (P | | Q)

is fixed, which is referred to as the reference distribution. Figure 4 shows values of the

α

-geodesical skew divergence for a fixed reference Q, where both P and Q are restricted to be Gaussian. In this figure, the reference distribution is

N (0, 0.5)

and the parameters of input distributions are varied in

μ \in [0, 4.5]

and

σ^{2} \in [0.5, 2.3]

. From this figure, one can see that a larger value of

α

emphasizes the discrepancy between distributions P and Q. Figure 5 illustrates a coordinate system associated with the

α

-geodesical skew divergence for different

α

. As seen from the figure, for the same pair of distributions P and Q, the value of divergence with

α = 3

is larger than that with

α = - 1

.

6. Conclusions and Discussion

In this paper, a new family of divergence is proposed to address the computational difficulty of KL-divergence. The proposed

α

-geodesical skew divergence is a natural derivation from the concept of

α

-geodesics in information geometry and generalizes many existing divergences.

Furthermore,

α

-geodesical skew divergence leads to several applications. For example, the new divergence can be applied to the annealed importance sampling by the same analogy as in previous studies using q-paths [41]. It could also be applied to linguistics, a field in which skew divergence was originally used [19].

Author Contributions

Formal analysis, M.K. and H.H.; Investigation, M.K.; Methodology, M.K. and H.H.; Software, M.K.; Supervision, H.H.; Validation, H.H.; Visualization, M.K.; Writing—original draft, M.K.; Writing—review & editing, H.H. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by JPSJ (KAKENHI) grant number JP17H01793, JST CREST Grant No. JPMJCR2015 and NEDO JPNP18002.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

Not applicable.

Acknowledgments

The authors express special thanks to the editor and reviewers, whose comments led to valuable improvements to the manuscript.

Conflicts of Interest

The authors declare no conflict of interest.

References

Deza, M.M.; Deza, E. Encyclopedia of distances. In Encyclopedia of Distances; Springer: Berlin/Heidelberg, Germany, 2009; pp. 1–583. [Google Scholar]
Basseville, M. Divergence measures for statistical data processing—An annotated bibliography. Signal Process. 2013, 93, 621–633. [Google Scholar] [CrossRef]
Kullback, S.; Leibler, R.A. On information and sufficiency. Ann. Math. Stat. 1951, 22, 79–86. [Google Scholar] [CrossRef]
Sakamoto, Y.; Ishiguro, M.; Kitagawa, G. Akaike Information Criterion Statistics; D. Reidel: Dordrecht, The Netherlands, 1986; Volume 81, p. 26853. [Google Scholar]
Goldberger, J.; Gordon, S.; Greenspan, H. An Efficient Image Similarity Measure Based on Approximations of KL-Divergence Between Two Gaussian Mixtures. ICCV 2003, 3, 487–493. [Google Scholar]
Yu, D.; Yao, K.; Su, H.; Li, G.; Seide, F. KL-divergence regularized deep neural network adaptation for improved large vocabulary speech recognition. In Proceedings of the 2013 IEEE International Conference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada, 16–31 May 2013. [Google Scholar]
Solanki, K.; Sullivan, K.; Madhow, U.; Manjunath, B.; Chandrasekaran, S. Provably secure steganography: Achieving zero KL divergence using statistical restoration. In Proceedings of the 2006 International Conference on Image Processing, Atlanta, GA, USA, 8–11 October 2006. [Google Scholar]
Lin, J. Divergence measures based on the Shannon entropy. IEEE Trans. Inf. Theory 1991, 37, 145–151. [Google Scholar] [CrossRef] [Green Version]
Menéndez, M.; Pardo, J.; Pardo, L.; Pardo, M. The jensen-shannon divergence. J. Frankl. Inst. 1997, 334, 307–318. [Google Scholar] [CrossRef]
Nielsen, F. On the Jensen–Shannon symmetrization of distances relying on abstract means. Entropy 2019, 21, 485. [Google Scholar] [CrossRef] [Green Version]
Jeffreys, H. An Invariant Form for the Prior Probability in Estimation Problems. Available online: https://royalsocietypublishing.org/doi/10.1098/rspa.1946.0056 (accessed on 24 April 2021).
Chatzisavvas, K.C.; Moustakidis, C.C.; Panos, C. Information entropy, information distances, and complexity in atoms. J. Chem. Phys. 2005, 123, 174111. [Google Scholar] [CrossRef] [Green Version]
Bigi, B. Using Kullback-Leibler distance for text categorization. In European Conference on Information Retrieval; Springer: Berlin/Heidelberg, Germany, 2003; pp. 305–319. [Google Scholar]
Wang, F.; Vemuri, B.C.; Rangarajan, A. Groupwise point pattern registration using a novel CDF-based Jensen-Shannon divergence. In Proceedings of the 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06), New York, NY, USA, 17–22 June 2006. [Google Scholar]
Nishii, R.; Eguchi, S. Image classification based on Markov random field models with Jeffreys divergence. J. Multivar. Anal. 2006, 97, 1997–2008. [Google Scholar] [CrossRef] [Green Version]
Bayarri, M.; García-Donato, G. Generalization of Jeffreys divergence-based priors for Bayesian hypothesis testing. J. R. Stat. Soc. Ser. B (Stat. Methodol.) 2008, 70, 981–1003. [Google Scholar] [CrossRef]
Nielsen, F. Jeffreys centroids: A closed-form expression for positive histograms and a guaranteed tight approximation for frequency histograms. IEEE Signal Process. Lett. 2013, 20, 657–660. [Google Scholar] [CrossRef] [Green Version]
Nielsen, F. On a generalization of the Jensen–Shannon divergence and the Jensen–Shannon centroid. Entropy 2020, 22, 221. [Google Scholar] [CrossRef] [Green Version]
Lee, L. Measures of distributional similarity. In Proceedings of the 37th Annual Meeting of the Association for Computational Linguistics on Computational Linguistics, College Park, MD, USA, 20–26 June 1999. [Google Scholar]
Lee, L. On the Effectiveness of the Skew Divergence for Statistical Language Analysis. In Proceedings of the Eighth International Workshop on Artificial Intelligence and Statistics, Key West, FL, USA, 4–7 January 2001; pp. 176–783. [Google Scholar]
Xiao, F.; Wu, Y.; Zhao, H.; Wang, R.; Jiang, S. Dual skew divergence loss for neural machine translation. arXiv 2019, arXiv:1908.08399. [Google Scholar]
Carvalho, B.M.; Garduño, E.; Santos, I.O. Skew divergence-based fuzzy segmentation of rock samples. J. Phys. Conf. Ser. 2014, 490, 012010. [Google Scholar] [CrossRef]
Revathi, P.; Hemalatha, M. Cotton leaf spot diseases detection utilizing feature selection with skew divergence method. Int. J. Sci. Eng. Technol. 2014, 3, 22–30. [Google Scholar]
Ahmed, N.; Neville, J.; Kompella, R.R. Network Sampling via Edge-Based Node Selection with Graph Induction. Available online: https://docs.lib.purdue.edu/cgi/viewcontent.cgi?article=2743&context=cstech (accessed on 24 April 2021).
Hughes, T.; Ramage, D. Lexical semantic relatedness with random graph walks. In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL), Prague, Czech Republic, 28–30 June 2007. [Google Scholar]
Audenaert, K.M. Quantum skew divergence. J. Math. Phys. 2014, 55, 112202. [Google Scholar] [CrossRef] [Green Version]
Hardy, G.H.; Littlewood, J.E.; Pólya, G. Inequalities; Cambridge University Press: Cambridge, UK, 1952. [Google Scholar]
Amari, S.I. Information Geometry and Its Applications; Springer: Berlin/Heidelberg, Germany, 2016. [Google Scholar]
Kolmogorov, A.N.; Castelnuovo, G. Sur la Notion de la Moyenne; Atti Accad. Naz: Lincei, French, 1930. [Google Scholar]
Nagumo, M. Über eine klasse der mittelwerte. Jpn. J. Math. 1930, 7, 71–79. [Google Scholar] [CrossRef] [Green Version]
Nielsen, F. Generalized Bhattacharyya and Chernoff upper bounds on Bayes error using quasi-arithmetic means. Pattern Recognit. Lett. 2014, 42, 25–34. [Google Scholar] [CrossRef] [Green Version]
Amari, S.I. Differential-Geometrical Methods in Statistics; Springer Science & Business Media: Berlin/Heidelberg, Germany, 2012; Volume 28. [Google Scholar]
Amari, S. Differential-geometrical methods in statistics. Lect. Notes Stat. 1985, 28, 1. [Google Scholar]
Amari, S. α-Divergence Is Unique, Belonging to Both f-Divergence and Bregman Divergence Classes. IEEE Trans. Inf. Theory 2009, 55, 4925–4931. [Google Scholar] [CrossRef]
Ay, N.; Jost, J.; Lê, H.V.; Schwachhöfer, L. Information Geometry; Springer International Publishing: Berlin/Heidelberg, Germany, 2017. [Google Scholar] [CrossRef]
Morozova, E.A.; Chentsov, N.N. Markov invariant geometry on manifolds of states. J. Sov. Math. 1991, 56, 2648–2669. [Google Scholar] [CrossRef]
Eguchi, S.; Komori, O. Path Connectedness on a Space of Probability Density Functions. In Lecture Notes in Computer Science; Springer International Publishing: Berlin/Heidelberg, Germany, 2015; pp. 615–624. [Google Scholar] [CrossRef]
Nielsen, F. On a Variational Definition for the Jensen-Shannon Symmetrization of Distances Based on the Information Radius. Entropy 2021, 23, 464. [Google Scholar] [CrossRef]
Nielsen, F. A family of statistical symmetric divergences based on Jensen’s inequality. arXiv 2010, arXiv:1009.4004. [Google Scholar]
Cover, T.M. Elements of Information Theory; John Wiley & Sons: Hoboken, NJ, USA, 1999. [Google Scholar]
Brekelmans, R.; Masrani, V.; Bui, T.D.; Wood, F.D.; Galstyan, A.; Steeg, G.V.; Nielsen, F. Annealed Importance Sampling with q-Paths. arXiv 2020, arXiv:2012.07823. [Google Scholar]

Figure 1. Monotonicity of the

α

-geodesical skew divergence with respect to

α

. The

α

-geodesical skew divergence between the binomial distributions

p = B (10, 0.3)

and

q = B (10, 0.7)

has been calculated.

Figure 1. Monotonicity of the

α

-geodesical skew divergence with respect to

α

. The

α

-geodesical skew divergence between the binomial distributions

p = B (10, 0.3)

and

q = B (10, 0.7)

has been calculated.

Figure 2. Continuity of the

α

-geodesmcal skew divergence with respect to

α

and

λ

. The

α

-geodesical skew divergence between the binomial distributions

p = B (10, 0.3)

and

q = B (10, 0.7)

has been calculated.

Figure 2. Continuity of the

α

-geodesmcal skew divergence with respect to

α

and

λ

. The

α

-geodesical skew divergence between the binomial distributions

p = B (10, 0.3)

and

q = B (10, 0.7)

has been calculated.

Figure 3. The geodesic between two probability distributions on the

α

-coordinate system.

Figure 3. The geodesic between two probability distributions on the

α

-coordinate system.

Figure 4.

α

-geodesical skew divergence between two normal distributions. The reference distribution is

Q = N (0, 0.5)

. For

P_{1}, P_{2}, \dots, P_{j}, (j = 1, 2, \dots, 10)

, let their mean and variance be

μ_{j}

and

σ_{j}^{2}

, respectively, where

μ_{j + 1} - μ_{j} = 0.5

and

σ_{j + 1}^{2} - σ_{j}^{2} = 0.2

.

Figure 4.

α

-geodesical skew divergence between two normal distributions. The reference distribution is

Q = N (0, 0.5)

. For

P_{1}, P_{2}, \dots, P_{j}, (j = 1, 2, \dots, 10)

, let their mean and variance be

μ_{j}

and

σ_{j}^{2}

, respectively, where

μ_{j + 1} - μ_{j} = 0.5

and

σ_{j + 1}^{2} - σ_{j}^{2} = 0.2

.

Figure 5. Coordinate system of

F_{q}

or

F_{+}

. Such a coordinate system is not Euclidean.

Figure 5. Coordinate system of

F_{q}

or

F_{+}

. Such a coordinate system is not Euclidean.

Publisher’s Note: MDPI stays neutral with regard to jurisdictional claims in published maps and institutional affiliations.

© 2021 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).

Share and Cite

MDPI and ACS Style

Kimura, M.; Hino, H. α-Geodesical Skew Divergence. Entropy 2021, 23, 528. https://doi.org/10.3390/e23050528

AMA Style

Kimura M, Hino H. α-Geodesical Skew Divergence. Entropy. 2021; 23(5):528. https://doi.org/10.3390/e23050528

Chicago/Turabian Style

Kimura, Masanari, and Hideitsu Hino. 2021. "α-Geodesical Skew Divergence" Entropy 23, no. 5: 528. https://doi.org/10.3390/e23050528

APA Style

Kimura, M., & Hino, H. (2021). α-Geodesical Skew Divergence. Entropy, 23(5), 528. https://doi.org/10.3390/e23050528

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Menu

α-Geodesical Skew Divergence

Abstract

1. Introduction

2. α-Geodesical Skew Divergence

2.1. Statistical Manifold

2.2. Generalization of Skew Divergences

2.3. Symmetrization of $α$ -Geodesical Skew Divergence

3. Properties of α-Geodesical Skew Divergence

4. Natural α-Geodesical Skew Divergence for Exponential Family

5. Function Space Associated with the α-Geodesical Skew Divergence

6. Conclusions and Discussion

Author Contributions

Funding

Institutional Review Board Statement

Informed Consent Statement

Data Availability Statement

Acknowledgments

Conflicts of Interest

References

Share and Cite

Article Metrics

Article Access Statistics

Further Information

Guidelines

MDPI Initiatives

Follow MDPI

Article Menu

α-Geodesical Skew Divergence

Abstract

1. Introduction

2. α-Geodesical Skew Divergence

2.1. Statistical Manifold

2.2. Generalization of Skew Divergences

2.3. Symmetrization of α -Geodesical Skew Divergence

3. Properties of α-Geodesical Skew Divergence

4. Natural α-Geodesical Skew Divergence for Exponential Family

5. Function Space Associated with the α-Geodesical Skew Divergence

6. Conclusions and Discussion

Author Contributions

Funding

Institutional Review Board Statement

Informed Consent Statement

Data Availability Statement

Acknowledgments

Conflicts of Interest

References

Share and Cite

Article Metrics

Article Access Statistics

Further Information

Guidelines

MDPI Initiatives

Follow MDPI

2.3. Symmetrization of $α$ -Geodesical Skew Divergence