Sitelet https://github.com/statsmodels/statsmodels/pull/10406
Skip to content

BUG: validate interval, alpha, and p_alt in binomial TOST helpers - #10406

Merged
bashtage merged 1 commit into
statsmodels:mainfrom
simpleqt:sq/binom-tost-helpers-validation
Oct 4, 2026
Merged

bashtage merged 1 commit into
statsmodels:mainfrom
simpleqt:sq/binom-tost-helpers-validation

Conversation

@simpleqt

@simpleqt simpleqt commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

Problem

The binomial TOST helpers silently accept impossible inputs:

>>> binom_tost_reject_interval(0.6, 0.3, 100)
(69.0, 22.0)     # reversed rejection region, silent
>>> binom_tost_reject_interval(0.1, 0.3, 100, alpha=2)
(nan, nan)       # silent
>>> power_binom_tost(0.6, 0.3, 100, p_alt=0.5)
-0.9999084204    # negative power, silent
>>> power_binom_tost(0.1, 0.3, 100, p_alt=1.5)
nan              # silent

This is the exact-binomial counterpart of the interval validation already added to `binom_tost`, `proportions_ztost`, and the weightstats TOST tests.

Fix

Raise `ValueError`:

  • `binom_tost_reject_interval`: `low < upp` and `alpha` in (0, 1)
  • `power_binom_tost` (wrapper): `low < upp`, `alpha` in (0, 1), `p_alt` in [0, 1]
  • `power_ztost_prop` (shared implementation): same three checks

Array-valued `p_alt` (used by the existing vectorized tests) is handled via `np.any`.

Tests

Added `test_binom_tost_helpers_invalid_inputs_raises` in `statsmodels/stats/tests/test_proportion.py` covering both functions and all three parameters. Verified red without the fix and green with it; the full test module passes (576 passed, 5 skipped) and the diff introduces no new `ruff check` findings.

AI Disclosure

Contributions comply with the statsmodels AI Policy. Completed option:

  • No AI tools were used to develop this pull request.
  • AI tools were used.
    Tool(s): `ZCode (GLM-based coding agent)`.
    Used for: `drafting the repro script, the implementation, the tests, and the PR text; I executed all code locally, verified the red/green test cycle, and reviewed every line of the diff`.
    I have personally read, understood, and can explain every line of this diff, and
    have verified the statistical/numerical correctness of the change.

binom_tost_reject_interval silently returned a reversed rejection region
(69, 22) for an inverted interval and (nan, nan) for alpha outside (0, 1);
power_binom_tost returned a negative power (-0.9999) for an inverted
interval and nan for p_alt outside [0, 1]. Raise ValueError in both, plus
the shared implementation power_ztost_prop. Array-valued p_alt inputs are
handled via np.any.
Copilot AI balanced review requested due to automatic review settings October 4, 2026 12:46

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@bashtage
bashtage merged commit 8b4743a into statsmodels:main Oct 4, 2026
10 of 11 checks passed
@bashtage

bashtage commented Oct 4, 2026

Copy link
Copy Markdown
Member

Thanks.

bashtage added a commit that referenced this pull request Oct 5, 2026
The checks added in GH-10379, GH-10386, GH-10388, GH-10392, GH-10393,
GH-10399, GH-10405, GH-10406, GH-10411, GH-10420, GH-10421, GH-10423,
GH-10425, GH-10426 and GH-10431 compare the arguments with scalars,
for example `if not 0 < alpha < 1`, `if low > upp` or `if dispersion < 0`.
These raise "The truth value of an array with more than one element is
ambiguous" for arrays and Series. The functions broadcast over these
arguments and return the same values as the calls for the elements in
0.15.0, for example a confidence interval for several alpha, the power
for several equivalence margins or a table of sample sizes.

Test with np.all and the comparison ufuncs instead, which work for
scalars, lists, arrays and Series, and which also reject NaN. The
error messages are unchanged. A new test checks for 50 arguments that an
array of two valid values gives the same result as the two calls with the
scalars, and that one invalid element (including NaN) is rejected.

Follow-up to GH-10379, GH-10386, GH-10388, GH-10392, GH-10393, GH-10399,
GH-10405, GH-10406, GH-10411, GH-10420, GH-10421, GH-10423, GH-10425,
GH-10426, GH-10431

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
bashtage added a commit that referenced this pull request Oct 5, 2026
The arguments that broadcast are documented as float or int, but arrays
work, and the previous commit keeps them working. Document them as
float or array_like in the functions of the previous commit and in
chisquare_power, which says in its Notes that it works vectorized.

Follow-up to GH-10379, GH-10386, GH-10388, GH-10392, GH-10393, GH-10399,
GH-10405, GH-10406, GH-10411, GH-10420, GH-10421, GH-10423, GH-10425,
GH-10426, GH-10431

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
bashtage added a commit that referenced this pull request Oct 5, 2026
binom_tost_reject_interval, power_binom_tost and power_ztost_prop did not
check nobs. power_binom_tost(0.4, 0.6, nobs=0) returned a power of -1.0 and
a negative nobs returned nan, or "nobs must be positive" in the other
functions of the module. Reject nobs <= 0 and NaN, for arrays too.

Follow-up to GH-10406
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants