Yield Statistics#
A simulated event catalog is one realization of a transient population and its observation by a survey. Its detected event count is only a realized yield and not a formal estimate of the expected yield.
This note defines the yield estimate: the analytically expected number of intrinsic events multiplied by the detection efficiency measured from the catalog. It establishes confidence bounds for this estimate, including uncertainty in the astrophysical rate normalization. All intervals are computed from distribution quantiles or confidence-set transformations; no symmetric-error approximation is used.
Conventions and notation#
We consider one transient population, one fixed survey schedule, and one specified detection criterion. The cosmology, redshift evolution, and distribution of transient properties are held fixed. Only the rate normalization is uncertain in the population model considered here.
Symbol |
Name |
Definition |
|---|---|---|
\(R(z;A)\) |
Volumetric event rate |
Events per comoving volume per source-frame time. |
\(A\) |
Rate normalization |
Unknown physical normalization of the volumetric rate. |
\(R_0\) |
Fiducial rate normalization |
Adopted central estimate of \(A\), used to generate the catalog. |
\(f(z)\) |
Rate evolution |
Fixed, nonnegative, dimensionless redshift dependence. |
\(\mathcal R(A)\) |
All-sky event rate |
Expected events per observer-frame time in the redshift domain. |
\(\mathcal V\) |
Rate-weighted comoving volume |
Full-sky volume integral including evolution and time dilation. |
\(\mathcal E\) |
Population exposure |
\(T\mathcal V\), with units of volume times time. |
\(\mu(A)\) |
Expected intrinsic count |
\(A\mathcal E\) in the population sampling domain. |
\(\epsilon\) |
Detection efficiency |
Probability that an event sampled from that domain is detected. |
\(\lambda(A,\epsilon)\) |
Expected detected yield |
\(A\mathcal E\epsilon\). |
\(n\) |
Simulated event count |
Number of intrinsic events in the realized catalog. |
\(k\) |
Detected event count |
Number of unique catalog events satisfying the detection criterion. |
\(\widehat\epsilon\) |
Estimated detection efficiency |
\(k/n\), when \(n>0\). |
\(\widehat\lambda\) |
Yield estimate |
\(R_0\mathcal E\widehat\epsilon\). |
Lowercase \(n\) and \(k\) denote realized integer counts. Greek \(\mu\) and \(\lambda\) denote expected counts, which need not be integers. Hats denote estimators, and subscripts \(\mathrm L\) and \(\mathrm U\) denote lower and upper bounds. A subscript \(0\) denotes evaluation at the fiducial normalization \(R_0\).
Every confidence interval carries a stated confidence level \(C=1-\alpha\). A frequentist confidence level describes the coverage of the interval-construction procedure under repeated sampling; it is not a probability distribution assigned to a fixed unknown parameter.
The population rate#
We write the intrinsic volumetric rate (events / comoving volume / rest-frame time) as
Here, \(A\) is the physical rate normalization, with units \(\mathrm{Mpc}^{-3}\,\mathrm{yr}^{-1}\), and \(f(z)\) is the redshift dependence of the rate.
Tip
The convention \(f(0)=1\) makes \(A\) the local volumetric rate, but another fixed normalization is equally valid.
In general, these intrinsic rates are taken from the literature and are reported with a given uncertainty,
If a publication reports \(R_0 {}^{+\Delta R_+}_{-\Delta R_-}\), the corresponding endpoints are
Convention
Wherever relevant in the code-base, uncertainties in the intrinsic rate are taken to be the 90% confidence interval.
All-sky rate and intrinsic count#
Cosmological time dilation gives
Consequently, the differential observer-frame event rate is
For an isotropic population in \(0\leq z\leq z_{\max}\), we define the rate-weighted comoving volume as
where \(D_M\) is the transverse comoving distance.
The all-sky event rate and its rate-only confidence bounds are
For an observer-frame sampling window of duration \(T\). The expected all-sky yield and its rate-only bounds are then
Detection efficiency and catalog generation#
Let \(x\) denote an event’s redshift, sky position, reference time, and intrinsic properties. Let \(p(x)\) be their normalized joint density in the sampling domain. For a fixed survey, define \(s(x)\in[0,1]\) as the probability that the event satisfies the detection criterion. A deterministic selection is the special case \(s(x)\in\{0,1\}\).
The population-averaged detection efficiency is
and the expected detected yield is
This efficiency includes the footprint, cadence, observing depth, extinction treatment, light-curve model, and detection criterion. Do not apply an additional sky-coverage factor if it is already included in \(\epsilon\).
At the fiducial normalization, the simulation draws
Conditional on \(n\), independently sampled events and independent event-level detection outcomes give
Individual events may have different detectabilities. The binomial model holds because their properties are independent draws from the same population and those properties are averaged over in \(\epsilon\). Conditional on a fixed list of unequal detection probabilities, the count instead follows a Poisson-binomial distribution.
Marginalizing over the Poisson-distributed intrinsic count gives
This is Poisson thinning. It assumes that detecting one event does not alter the selection of another. Shared random survey conditions, population-dependent follow-up competition, and correlations between events require additional modeling. The present bounds are conditional on the specified survey and population shape.
Estimating the expected yield#
For \(n>0\), the conditional binomial maximum-likelihood estimate is
The yield estimate is
At fixed \(R_0\), this estimator is conditionally unbiased:
The raw detected count \(k\) is also an unbiased estimator of \(\lambda_0\) under Poisson catalog generation. The efficiency-based estimator uses the known analytic normalization \(\mu_0\) and the observed denominator \(n\), rather than retaining fluctuations in the total simulated population count.
For example, if \(\mu_0=1000\) and \(n=k=917\), then \(\widehat\lambda=1000\). This does not establish that \(\epsilon=1\) exactly: a finite catalog in which all events are detected still permits efficiencies below unity.
The same estimator works when \(n\) is fixed deliberately, or when more events are simulated to improve precision. In those cases, \(\mu_0\) remains the target population’s expected intrinsic count; it is not replaced by the larger simulation sample size. For independent catalogs sampling the same \(p(x)\) and selection, pool their counts:
Confidence bounds from the simulated catalog#
For a requested simulation-only confidence level \(C_\epsilon=1-\alpha_\epsilon\), use the central Clopper–Pearson binomial interval. Write \(Q_{\rm B}(q;a,b)\) for the quantile at probability \(q\) of a Beta distribution with shape parameters \(a,b\). Then
These bounds invert binomial tail tests. Their coverage is at least \(1-\alpha_\epsilon\) for every fixed \(\epsilon\); discreteness generally makes the interval conservative. The Beta quantiles are a computational identity, not an assumed Bayesian posterior. See the SciPy Clopper–Pearson documentation.
With the normalization held at \(R_0\), the corresponding simulation-only yield bounds are
These bounds quantify uncertainty in the expected yield due to a finite simulation sample. They are not prediction bounds for the count in a future survey. Coverage also holds for a Poisson-distributed \(n\), because the binomial construction covers conditionally for each sample size, including the empty-catalog convention below.
Boundary cases#
For no detections in a nonempty catalog, \(k=0<n\),
The estimated yield is zero, but its upper confidence bound is positive. Zero simulated detections do not establish a zero expected yield.
For detections of every simulated event, \(k=n>0\),
For an empty catalog, \(n=k=0\), the detection efficiency is unidentified. Return no point estimate and use
Do not set \(\widehat\epsilon=0\) when \(n=0\).
The bounds above are two-sided central intervals. A deliberately one-sided upper bound at confidence \(1-\alpha_\epsilon\) instead uses
with upper bound one when \(k=n\). In particular, for \(k=0<n\) it is \(1-\alpha_\epsilon^{1/n}\). The upper endpoint of a central interval and a one-sided upper bound at the same stated confidence level are different quantities.
Including rate uncertainty#
Rate-only bounds#
Holding the estimated efficiency fixed gives the rate-only yield bounds
For \(R_0>0\), equivalently,
This displays the rate contribution conditional on the fitted efficiency. It does not account for uncertainty in that efficiency. In particular, when \(k=0\), these rate-only plug-in bounds collapse to zero even though the combined upper bound need not do so.
Combined bounds with guaranteed coverage#
A joint frequentist interval can be constructed from separate confidence sets without imposing a probability distribution on \(A\). Choose a target combined confidence level \(C=1-\alpha\) and allocate the failure probability between the two inputs:
A simple default is \(\alpha_\epsilon=\alpha_R=\alpha/2\). Compute the efficiency interval at \(1-\alpha_\epsilon\) and obtain the external rate interval at \(1-\alpha_R\). Both levels must be available; do not reinterpret a published interval at a different level.
By the union bound, the rectangle
covers the true pair \((\epsilon,A)\) with probability at least \(1-\alpha_\epsilon-\alpha_R\). Since \(\lambda=\mathcal E A\epsilon\) is nondecreasing in both nonnegative parameters, project the rectangle onto the expected yield:
Whenever the rectangle covers the true pair, this interval covers the true expected yield. Thus its coverage is at least \(C\). This is a conservative confidence construction, not an exact equal-tail interval for the product. It requires no Gaussian approximation and remains valid at \(k=0\) or \(k=n\).
For example, a combined 95% interval can use a 97.5% efficiency interval and a 97.5% rate interval. Multiplying the endpoints of two 95% intervals does not by itself establish 95% combined coverage: the union-bound guarantee is only 90%. Actual coverage can be higher.
The union-bound construction does not require independence. If the two interval-coverage events are independent, the stronger lower bound is
Under that additional assumption, equal component levels \(1-\alpha_\epsilon=1-\alpha_R=\sqrt C\) suffice. Merely using independent random seeds does not establish independence from every external modeling uncertainty. The union-bound allocation is the default here.
If a rate interval is available only at one confidence level, retain that level and state the coverage guaranteed by the available component intervals. Asymmetric endpoints alone cannot determine bounds at a new confidence level. They also do not determine a unique likelihood or sampling distribution. In that situation, report the simulation-only and rate-only bounds separately, or report the combined interval with its supported coverage guarantee.
For an empty catalog, the combined interval reduces to \([0,\mathcal E R_{\mathrm U}]\); there is still no efficiency-based point estimate. All statements of coverage are conditional on the fixed cosmology, evolution shape, and survey model.