On the Statistical Properties and Computational Inference of the Generalized Kumaraswamy Distribution Family
José Evandeilton Lopes
2026-09-27
Source:vignettes/theory-gkwdist.Rmd
theory-gkwdist.RmdAbstract. We present a comprehensive mathematical
treatment of the Generalized Kumaraswamy (GKw) distribution, a
five-parameter family for modeling continuous random variables on the
unit interval introduced by Carrasco et al. (2010). We establish the
hierarchical structure connecting GKw to several nested submodels,
including the package-parameterized Beta subfamily and the Kumaraswamy
distribution, derive the log-likelihood, score vector, and observed
information matrix, and state likelihood asymptotics under explicit
regularity and identifiability conditions. Particular attention is given
to non-identifiable parameter manifolds, boundary cases, and numerically
stable evaluation of the cascade transformations. The analytical
derivatives are written in a form suitable for direct implementation and
validation. These results provide the mathematical foundation for the
numerical routines in the R package gkwdist.
Keywords: Bounded distributions, Beta distribution, Kumaraswamy distribution, Maximum likelihood estimation, Fisher information, Numerical stability
1. Introduction and Preliminaries
1.1 Motivation and Background
The analysis of continuous random variables constrained to the unit interval arises naturally in numerous statistical applications, including proportions, rates, percentages, and index measurements. The classical Beta distribution (Johnson et al., 1995) has long served as the canonical model for such data, offering analytical tractability and well-understood properties. However, its cumulative distribution function (CDF) involves the incomplete beta function, requiring numerical evaluation of special functions for quantile computation and simulation.
Kumaraswamy (1980) introduced an alternative two-parameter family with closed-form CDF and quantile function, facilitating computational efficiency while maintaining comparable flexibility to the Beta distribution. Jones (2009) demonstrated that the Kumaraswamy distribution exhibits similar shape characteristics to the Beta family while offering superior computational advantages.
Building upon these foundations, Carrasco, Ferrari and Cordeiro (2010) introduced the five-parameter Generalized Kumaraswamy (GKw) distribution by combining the Kumaraswamy CDF with a generalized-beta generator. Cordeiro and de Castro (2011) developed the related Kumaraswamy-generated (Kw-G) family, which is conceptually connected but is not the five-parameter GKw model studied here. The GKw distribution encompasses a rich hierarchy of submodels and provides substantial flexibility for modeling diverse patterns in bounded data. Identifiability issues for this class were revisited explicitly by Carrasco and Cordeiro (2017), so likelihood asymptotics must be stated conditionally on a regular identifiable parameter region.
Despite its theoretical appeal, a fully explicit and internally
consistent analytical treatment of the GKw family—particularly for
likelihood-based inference—has remained incomplete in the literature.
This vignette fills this gap by providing a rigorous development,
including validated expressions for all first and second derivatives of
the log-likelihood function, written in a form convenient for
implementation in the gkwdist R package.
1.2 Mathematical Preliminaries
We establish notation and fundamental results required for subsequent development.
Notation 1.1. Throughout, we denote
- : the gamma function
- : the beta function
- : the regularized incomplete beta function
- : the digamma function
- : the trigamma function
- : the indicator function of a set
We recall basic derivatives of the beta function.
Lemma 1.1 (Derivatives of the beta function). For ,
Proof. Since the identities follow immediately from the definitions of and and the chain rule.
We will also repeatedly use the following cascade of transformations.
Lemma 1.2 (Cascade transformations). Define, for ,
Then, for ,
Proof. Direct differentiation and repeated application of the chain rule.
For brevity we will often write , and when the dependence on is clear from the context.
2. The Generalized Kumaraswamy Distribution and Its Subfamilies
2.1 Definition and Fundamental Properties
We start from the five-parameter Generalized Kumaraswamy family.
Definition 2.1 (Generalized Kumaraswamy
distribution).
A random variable
has a Generalized Kumaraswamy distribution with parameter vector
denoted
,
if its probability density function (pdf) is
where
and the extended parameter space used
by the package is
Its interior with respect to the
five-dimensional parameterization is
Although
is well-defined for
,
the package and the submodel hierarchy considered here use
.
Boundary points and certain interior manifolds require special care
because the full five-parameter representation is not everywhere
identifiable.
Proposition 2.1 (Exact non-identifiability
manifolds).
For the full five-parameter GKw representation, at least the following
parameter manifolds are non-identifiable:
- If , then the density depends on only through :
- If , then the density depends on only through :
- If and , then the density depends on only through : Thus the full five-parameter embedding is not globally identifiable, even though reduced submodels can be identifiable when expressed in their own minimal parameterizations.
Proof. If , then and , so (2.1) reduces immediately to (2.2a). Hence any two pairs with the same product generate the same density. If , then , and (2.1) becomes (2.2b), which depends on and only through . Finally, if , then , , and . Therefore which is (2.2c). Hence distinct parameter vectors can induce exactly the same distribution on each of these manifolds. Whenever the usual differentiability conditions hold, the Fisher information for the full parameterization is necessarily singular there.
We now verify that (2.1) defines a proper density.
Theorem 2.1 (Validity of the pdf).
For any
,
the function
in (2.1) is a valid probability density on
.
Proof. Non-negativity is immediate from the definition. To prove normalization, consider the change of variable From Lemma 1.2 and the chain rule, Hence
Substituting into the integral of , because and .
The same change of variable yields the CDF.
Theorem 2.2 (Cumulative distribution
function).
If
,
then for
,
Proof. From the same substitution as in (2.3), At the endpoints, and , so and .
Corollary 2.1 (Quantile function and random
generation).
Let
and let
denote the inverse, with respect to its
argument, of the regularized incomplete beta function. Then
Consequently, if
,
then
has the GKw distribution.
Proof. Set in (2.4). Then , hence , , and inversion of yields (2.4a). Equation (2.4b) follows by the probability integral transform.
2.2 The Hierarchical Structure
The GKw family exhibits a rich nested structure. Several well-known bounded distributions, or package-restricted portions of their usual parameter spaces, arise as particular choices (and mild reparameterizations) of .
For a data point we will often write
2.2.1 Beta–Kumaraswamy distribution
Theorem 2.3 (Beta–Kumaraswamy distribution).
Setting
in (2.1) yields the four-parameter Beta–Kumaraswamy (BKw) distribution
with pdf
and CDF
Proof. For , we have , so from (2.1), which is (2.5). The CDF follows from Theorem 2.2 with .
2.2.2 Kumaraswamy–Kumaraswamy distribution
Theorem 2.4 (Kumaraswamy–Kumaraswamy
distribution).
Setting
in (2.1) yields the four-parameter Kumaraswamy–Kumaraswamy (KKw)
submodel, parameterized consistently with gkwdist as
,
with pdf
CDF
and quantile function
Proof. With , , and (2.1) gives (2.7). From (2.4), which is (2.8). Solving successively for , , , and yields (2.9).
The standard Kumaraswamy distribution is recovered from KKw at the boundary value together with .
2.2.3 Exponentiated Kumaraswamy distribution
Theorem 2.5 (Exponentiated Kumaraswamy
distribution).
Setting
and
in (2.1) yields the three-parameter exponentiated Kumaraswamy (EKw)
distribution
with CDF
and quantile function
Proof. With and , we have , , and . Thus (2.1) reduces to (2.10). From (2.4), yielding (2.11). Inverting gives and then as in (2.12).
Note that the standard Kumaraswamy distribution appears as the special case of EKw.
2.2.4 McDonald distribution
Theorem 2.6 (McDonald/GB1 subfamily under the package
parameterization).
Setting
in (2.1) yields the package-parameterized McDonald (GB1) subfamily
with CDF
Proof. For we have , , and . Substituting into (2.1) yields (2.13); the CDF follows from (2.4).
Because , the second GB1 shape is . Thus, as with the Beta embedding in Theorem 2.8, this is the portion of the classical McDonald/GB1 family compatible with the package’s shifted parameterization.
2.2.5 Kumaraswamy distribution
Theorem 2.7 (Kumaraswamy distribution).
The standard two-parameter Kumaraswamy distribution is obtained from GKw
by taking
equivalently as the submodel
EKw()
with
.
Its pdf is
with CDF
quantile function
and
-th
moment, for
,
Proof. With , and , (2.1) reduces to (2.15) because , and . Equations (2.16) and (2.17) follow from (2.11) with . For the moment, which yields (2.18).
2.2.6 Beta distribution
Theorem 2.8 (Beta subfamily under the package
parameterization).
Setting
in (2.1) yields
with CDF
and
-th
moment, for
,
Proof. For , we have , , and . Substituting into (2.1) gives (2.19); the CDF and moment follow from standard Beta distribution theory with shape parameters .
Remark 2.2 (Range of the embedded Beta
family).
Because the gkwdist parameterization uses
,
this embedding has Beta shape parameters
and
.
It is therefore a genuine Beta subfamily, but it does
not span the entire classical two-shape Beta family
.
The latter would require extending the mathematical parameter range to
,
which is outside the package parameterization used in this vignette.
3. Likelihood-Based Inference
3.1 The Log-Likelihood Function
Let be an i.i.d. sample from , with observed values . For each , define
Definition 3.1 (Log-likelihood function).
The log-likelihood is
Theorem 3.1 (Decomposition of the
log-likelihood).
The log-likelihood can be written as
where
Equivalently, where
Proof. Take logarithms of (2.1) and sum over .
3.2 Maximum Likelihood Estimation
Definition 3.2 (Maximum likelihood estimator).
Whenever the likelihood maximum is attained, an MLE is any maximizer
The set of maximizers need not be a
singleton on non-identifiable manifolds. In a regular identifiable
parameterization, the local maximizer considered below is assumed to be
well defined with probability tending to one.
We will refer to the true parameter value as .
Theorem 3.2 (Consistency and asymptotic normality on a
regular identifiable region).
Let
belong to an open identifiable parameter region
.
Assume:
- is the unique maximizer of on the parameter region considered, and the normalized log-likelihood obeys a uniform law of large numbers sufficient for M-estimator consistency;
- is twice continuously differentiable in a neighborhood of , with domination conditions allowing differentiation under the integral sign and a law of large numbers for the Hessian;
- the per-observation score has finite second moment and is positive definite.
Then any sequence of MLEs in this regular region satisfying the usual approximate-maximization condition is consistent, and
Proof. The first assertion follows from the standard consistency theorem for M-estimators because converges uniformly to , whose maximizer is unique at . For asymptotic normality, expand the score around : where lies between and . By the multivariate central limit theorem, , while consistency and the Hessian law of large numbers give . Slutsky’s theorem yields (3.15).
Remark 3.1 (Why the qualification is
essential).
Theorem 3.2 does not apply to the full five-parameter representation on
the non-identifiable manifolds in Proposition 2.1. In particular,
makes
confounded through
;
makes
confounded through
;
and
makes
confounded through
.
At such points the information matrix for the full embedding is singular
and the usual five-dimensional Gaussian limit cannot hold. A reduced
model can nevertheless possess regular asymptotics when it is written in
an identifiable minimal parameterization.
3.3 Likelihood Ratio Tests
For nested models, likelihood ratio tests follow Wilks’ theorem.
Theorem 3.3 (Wilks’ theorem on a regular interior
null).
Consider nested hypotheses
versus
,
where locally
is a smooth identifiable submanifold of codimension
,
the true parameter is an interior point relative to the regular
parameterization, and the Fisher information is nonsingular. Let
Then, under
,
Proof. This is the standard Wilks expansion of the log-likelihood under local asymptotic normality and a regular interior constraint. The assumptions exclude boundary points and loss of identifiability.
Corollary 3.1 (LR tests within the GKw
hierarchy).
The following regular reductions have the usual chi-squared limit
provided the null point lies in an identifiable region with nonsingular
information:
For the first two comparisons, this
excludes, in particular, null points at which the full GKw
representation loses identifiability.
The comparison KKw versus Kw is different. In distribution space, the Kw null inside KKw is the singular slice because, when is already fixed by KKw and , Proposition 2.1(3) gives which is exactly a Kw density for every . The commonly used representative restriction selects one parameterization of the same Kw null but is not its full preimage in the redundant KKw coordinates. Consequently, under , the split of into and is unidentified. The null is therefore a singular loss-of-identifiability problem (and the representative also lies on the package boundary). Wilks’ ordinary chi-squared theorem does not apply, and one should not assume a standard chi-bar-square mixture without a model-specific derivation of the local experiment. In practice, a parametric bootstrap under the fitted two-parameter Kw null is a defensible calibration strategy, but its finite-sample performance should itself be checked by simulation for the intended sample size and parameter region. Self and Liang (1987) provide a general boundary-likelihood framework; the unidentified nuisance direction present here requires additional care.
4. Analytical Derivatives and Information Matrix
We now derive explicit expressions for the score vector and the observed information matrix in terms of the cascade transformations and their derivatives.
4.1 The Score Vector
Definition 4.1 (Score function).
The score vector is
For the derivatives of with respect to the parameters we will use
Theorem 4.1 (Score components).
The components of
are
Proof. We differentiate the decomposition (3.4)–(3.12) term by term.
(i) Derivative with respect to .
From (3.6) and (3.9), Using , Next, , so Similarly, so Collecting terms gives (4.2).
(ii) Derivative with respect to .
From (3.7) and (3.10), since does not depend on . Using , Furthermore, so Combining terms yields (4.3).
(iii) Derivative with respect to .
Only and depend on . From Lemma 1.1, giving (4.4).
(iv) Derivative with respect to .
Similarly, giving (4.5).
(v) Derivative with respect to .
We have and Together these yield (4.6).
Implementation sign convention.
Equations (4.2)–(4.6) are derivatives of the
log-likelihood
.
The optimization helpers in gkwdist are written for the
negative log-likelihood; therefore their gradient and Hessian should
satisfy
This distinction is essential when
comparing the formulas in this vignette with gr* and
hs* implementation routines.
4.2 The Hessian and Observed Information Matrix
We now consider second-order derivatives. Let denote the Hessian matrix of the log-likelihood, where .
Definition 4.2 (Observed information).
The observed information matrix is defined as
To keep the formulas compact, for each observation and each transformation we define, for parameters ,
In particular,
For implementation, all raw second derivatives entering can be written explicitly. Let Then For , These formulas, together with the definition of , give every Hessian entry below without numerical differentiation.
4.2.1 Diagonal elements
Theorem 4.2 (Diagonal elements of the
Hessian).
The second derivatives of
with respect to each parameter are
Equivalently, one may write so that (4.8)–(4.9) can be expressed in the same notation as in the original decomposition.
Proof. Differentiate the score components (4.2)–(4.6) with respect to the same parameter. The term contributes to , and similarly for and . The contributions from are precisely the second derivatives of , and , which yield the terms in or .
For and , only depends on these parameters through ; the formulas (4.10)–(4.11) follow from Lemma 1.1. Finally, does not depend on beyond the linear factor , so , and the only contribution to (4.12) besides comes from , whose second derivative w.r.t. is .
4.2.2 Off-diagonal elements
Theorem 4.3 (Off-diagonal elements of the
Hessian).
For
,
the mixed second derivatives of
are:
Proof.
Equation (4.18) follows by differentiating (4.4) with respect to ; only the term involving contributes a non-zero derivative, giving .
Equations (4.19)–(4.22) follow from differentiating (4.4) and (4.5) with respect to or . For example,
Equation (4.17) is obtained by differentiating (4.2) with respect to ; only the factor contributes the term , while the dependence of on is captured by and .
For , note from (4.6) that Differentiating with respect to or yields (4.23)–(4.24); differentiating with respect to and gives (4.25)–(4.26).
Mixed derivatives commute by Schwarz’s theorem, so the Hessian is symmetric.
In practice, the resulting expressions for are evaluated numerically by plugging in the analytic first and second derivatives of , which follow recursively from the definitions of these transformations.
4.3 Asymptotic Variance–Covariance Matrix
On a regular identifiable parameter region satisfying Theorem 3.2, the asymptotic variance–covariance matrix of the MLE is governed by the Fisher information.
Let denote the per-observation Fisher information,
For the full sample of size , the expected information is , while the observed information is .
Theorem 4.4 (Variance–covariance matrix of the
MLE).
Under the regularity assumptions of Theorem 3.2,
and the asymptotic standard error of
is approximated by
Proof. The convergence follows from a law of large numbers for the Hessian (Cox and Hinkley, 1974). Combining this with the asymptotic normality (3.15) yields (4.27)–(4.28).
5. Computational Aspects and Discussion
5.1 Numerical Stability
Direct evaluation of , , and can suffer catastrophic cancellation when the subtracted quantity is close to one. The implementation should therefore propagate the cascade on the log scale.
Algorithm 5.1 (Stable computation of for ).
For , define The first branch avoids loss of relative accuracy when is small, while the second avoids cancellation when is close to zero. Thus This is the stable branch structure recommended by Mächler (2012).
The same principle should be used for ratios appearing in the score and Hessian. For example, Products such as should likewise be assembled from sums and differences of logged factors whenever direct powers or divisions would overflow, underflow, or magnify cancellation.
5.2 Optimization
For the regular five-parameter model it is convenient to optimize an unconstrained negative log-likelihood. Define so that . Boundary submodels such as EKw or Kw should be fitted as their own reduced parameterizations rather than approximated by driving numerically to zero in the non-identifiable five-parameter representation. Let
Algorithm 5.2 (Maximum likelihood via BFGS applied to ).
- Initialization. Obtain a finite starting vector , preferably from a simpler identifiable submodel.
- BFGS iteration. At iteration , evaluate and its exact gradient. With a positive-definite approximation to the inverse Hessian of , use the descent direction choose a step length satisfying standard line-search conditions (e.g. Wolfe conditions), and update Apply the standard BFGS inverse-Hessian update to .
- Convergence checks. Require a small gradient norm together with stable parameter and objective changes; inspect the Hessian conditioning and profile likelihood when parameters are weakly identified.
- Observed information. Transform the Hessian back to the original parameterization if standard errors are reported on the -scale, and compute the observed information only at a regular local maximum.
This formulation avoids the sign inconsistency that arises if a positive-definite BFGS matrix is described as an approximation to the Hessian of the maximized log-likelihood, whose Hessian is negative definite at a strict local maximum.
Theorem 5.1 (Local superlinear convergence of
BFGS).
Suppose
is twice continuously differentiable in a neighborhood of a local
minimizer
,
is positive definite, the Hessian is locally Lipschitz continuous, the
BFGS iterates converge to
,
exact gradients are used, and the line search satisfies the standard
conditions required by BFGS convergence theory. Then the BFGS iterates
are q-superlinearly convergent:
This is a local convergence statement;
exact gradients alone do not guarantee global convergence for an
arbitrary non-concave likelihood.
Justification. Under the stated smoothness, curvature, convergence and line-search hypotheses, the standard BFGS theory yields the Dennis–Moré condition for the inverse-Hessian approximations, which implies q-superlinear convergence of the iterates. This is a theorem about the local minimization problem for , not a claim that an arbitrary starting value reaches the global MLE; see Nocedal and Wright (2006), Chapter 6.
5.3 Gradient Accuracy
Numerical differentiation can be used to validate the analytic derivatives but is less efficient for routine computation.
Lemma 5.1 (Finite-difference error).
Consider the central finite-difference approximation to
with step size
:
Then
where
is machine precision (approximately
in double precision). Balancing truncation and roundoff errors gives the
familiar cube-root scaling
after the objective and parameter have been appropriately scaled. There
is no universal numerical value of
,
because the constants depend on the local third derivative and on the
scale of the likelihood evaluation.
Proof. Taylor expansion about gives so subtraction and division by leave a truncation error . If each floating-point objective evaluation has absolute perturbation of order on the working scale, subtracting two such values and dividing by contributes . Balancing these two leading orders gives , hence after suitable scaling. See Nocedal and Wright (2006), Chapter 8.
In contrast, the analytical gradients of Theorem 4.1 can be evaluated with accuracy limited essentially only by floating-point roundoff and can be accumulated in one pass through the data, whereas a central-difference approximation to the full -dimensional gradient requires likelihood evaluations (here ).
5.4 Practical Recommendations
Guideline 5.1 (Model selection within the GKw hierarchy).
- Begin with simpler identifiable models such as the embedded Beta subfamily or the Kumaraswamy model when they are scientifically plausible.
- Consider EKw or McDonald-type extensions when one additional shape mechanism is needed.
- Use BKw or KKw when the additional flexibility is supported by the data and by stable likelihood geometry.
- Fit the full five-parameter GKw model only when the likelihood is well identified: inspect profile likelihoods, the observed-information spectrum, sensitivity to starting values, and parametric-bootstrap behavior. There is no distribution-free sample-size cutoff such as that guarantees identifiability or numerical stability.
- Compare candidate models with likelihood criteria such as while remembering that ordinary likelihood-ratio reference distributions can fail for boundary or non-identifiable reductions.
Guideline 5.2 (Diagnostics).
- Q-Q plot. Use the quantile representation (2.4a), with the inverse beta function evaluated numerically.
- Probability integral transform. The fitted values should resemble a Uniform sample, but they are not exactly i.i.d. Uniform in finite samples because the same data were used to estimate . Formal calibration should therefore use a parametric bootstrap when needed.
- Information conditioning. Examine the eigenvalues or a suitably scaled condition number of . The raw condition number is parameterization-dependent, so fixed thresholds should be treated only as numerical warnings rather than universal inferential rules.
- Positive definiteness. At a strict regular local maximum, the Hessian of is negative definite and the observed information is positive definite. Failure of this property can indicate a saddle point, a flat/non-identifiable direction, a boundary optimum, or numerical error.
- Profile likelihood. For a five-parameter model, profile plots are especially useful for detecting ridges involving , , or near the non-identifiable manifolds in Proposition 2.1.
5.5 Discussion
We have developed a rigorous mathematical framework for the Generalized Kumaraswamy (GKw) family, including:
Hierarchical embedding.
The GKw family contains the McDonald, Kumaraswamy, exponentiated Kumaraswamy, Beta–Kumaraswamy and Kumaraswamy–Kumaraswamy distributions as submodels, together with the shifted Beta subfamily described in Theorem 2.8, with explicit parameter mappings.Likelihood theory.
We derived explicit expressions for the log-likelihood, the score vector and the full observed information matrix in terms of the cascade transformations , in a form suitable for stable numerical implementation.Likelihood asymptotics and identifiability.
Standard MLE asymptotics hold only on regular identifiable regions with nonsingular Fisher information. The full five-parameter representation is non-identifiable on important manifolds including , , and , so boundary and reduced models must be handled separately.Computational considerations.
Log-scale evaluations and carefully structured derivatives provide numerical stability and efficiency. In our C++ implementation via RcppArmadillo, analytical gradients and Hessians yield substantial speedups over finite-difference approximations, together with better numerical accuracy.
Natural extensions and computational developments include:
- Compact and numerically efficient representations for the full GKw moments. Carrasco et al. (2010) provide series representations; simple elementary closed forms are generally unavailable.
- Higher-order likelihood corrections and profile-likelihood inference in weakly identified regions.
- Boundary-aware likelihood-ratio calibration and parametric-bootstrap procedures for comparisons such as KKw versus Kw.
- Multivariate generalizations using copulas constructed from GKw marginals.
- Fully Bayesian treatments with priors and parameterizations designed to avoid or regularize non-identifiable directions.
The full GKw quantile is already available computationally through the inverse regularized beta function followed by the explicit cascade inversion in (2.4a); no elementary inverse-beta formula is required for simulation or Q-Q diagnostics.
The gkwdist R package provides the distributional and
likelihood-computation routines underlying this vignette. The
identifiability, singular-null and numerical-stability qualifications
established here should be treated as part of the mathematical contract
for interpreting those routines and for future package-level inference
helpers.
References
Carrasco, J. M. F., Ferrari, S. L. P., & Cordeiro, G. M. (2010). A new generalized Kumaraswamy distribution. arXiv:1004.0911. arxiv.org/abs/1004.0911
Carrasco, J. M. F. and Cordeiro, G. M. (2017). An extension of the Kumaraswamy distribution. International Journal of Statistics and Probability, 6(3), 61. doi:10.5539/ijsp.v6n3p61
Casella, G. and Berger, R. L. (2002). Statistical Inference, 2nd ed. Duxbury Press, Pacific Grove, CA.
Cordeiro, G. M. and de Castro, M. (2011). A new family of generalized distributions. J. Stat. Comput. Simul. 81, 883–898.
Cox, D. R. and Hinkley, D. V. (1974). Theoretical Statistics. Chapman and Hall, London.
Johnson, N. L., Kotz, S. and Balakrishnan, N. (1995). Continuous Univariate Distributions, Volume 2, 2nd ed. Wiley, New York.
Jones, M. C. (2009). Kumaraswamy’s distribution: A beta-type distribution with some tractability advantages. Statist. Methodol. 6, 70–81.
Kumaraswamy, P. (1980). A generalized probability density function for double-bounded random processes. J. Hydrol. 46, 79–88.
Lehmann, E. L. and Casella, G. (1998). Theory of Point Estimation, 2nd ed. Springer, New York.
Mächler, M. (2012). Accurately computing . R package vignette, https://CRAN.R-project.org/package=Rmpfr.
Nocedal, J. and Wright, S. J. (2006). Numerical Optimization, 2nd ed. Springer, New York.
Self, S. G. and Liang, K.-Y. (1987). Asymptotic properties of maximum likelihood estimators and likelihood ratio tests under nonstandard conditions. Journal of the American Statistical Association, 82(398), 605–610. doi:10.1080/01621459.1987.10478472
van der Vaart, A. W. (1998). Asymptotic Statistics. Cambridge University Press, Cambridge.
Author’s address:
J. E. Lopes
Laboratory of Statistics and Geoinformation (LEG)
Graduate Program in Numerical Methods in Engineering (PPGMNE)
Federal University of Paraná (UFPR)
Curitiba, PR, Brazil
E-mail: evandeilton@gmail.com