4 Information and Asymptotics

Orthogonality

We now discuss parameter orthogonality, which has important consequences for likelihood characteristics and inference.

Let θ→=(ϕ→,λ→) of dimension d1 and d2 respectively (d1+d2=d), i.e. ϕ→=(ϕ1,…,ϕd1) and λ→=(λ1,…,λd2). The expected information matrix, IE⁢(θ→), can be partitioned as follows

IE⁢(θ→)=[Iϕ→ϕ→⁢(θ→)Iϕ→λ→⁢(θ→)Iϕ→λ→⁢(θ→)Iλ→λ→⁢(θ→)],

where Iϕ→ϕ→⁢(θ→) has (i,j)th element

iϕi⁢ϕj=E⁢{-∂2⁡ℓ⁢(θ→)∂⁡ϕi⁢∂⁡ϕj},

similarly Iλ→λ→⁢(θ→) has (i,j)th element

iλi⁢λj=E⁢{-∂2⁡ℓ⁢(θ→)∂⁡λi⁢∂⁡λj},

and Iϕ→λ→⁢(θ→) has (i,j)th element

iϕi⁢λj=E⁢{-∂2⁡ℓ⁢(θ→)∂⁡ϕi⁢∂⁡λj}.

Definition of Orthogonality
We define ϕ→ to be orthogonal to λ→ if Iϕ→λ→⁢(θ→)=𝟎, i.e. the elements of this matrix satisfy the property

iϕi⁢λj=E⁢(-∂2⁡ℓ⁢(θ→)∂⁡ϕi⁢∂⁡λj)=0

for i=1,…,d1 and j=1,…,d2 whatever value of θ→ in Ω.

Closely linked to this is the orthogonality of the observed information matrix at θ→^.

Orthogonality of (ϕ→,λ→) means that the maximum likelihood estimates ϕ→^ and λ→^ are asymptotically independent. This has a number of practical advantages:

  1. 1.

    Parameter interpretation;

  2. 2.

    The asymptotic standard error for estimating ϕ→ is the same whether λ→ is treated as known or unknown;

  3. 3.

    Likelihood based confidence intervals are simpler to summarize and interpret.