Yet another Python survival analysis tool.
This is another pure python survival analysis tool so why was it needed? The intent of this package was to closely mimic the scipy API as close as possible with a simple .fit() method for any type of distribution (parametric or non-parametric); other survival analysis packages don't completely mimic that API. Further, there is currently (at the time of writing) no pacakage that can take an arbitrary comination of observed, censored, and truncated data. Finally, surpyval is unique in that it can be used with multiple parametric estimation methods. This allows for an analyst to determine a distribution for the parameters if another method fails. The parametric methods available are Maximum Likelihood Estimation (MLE), Probability Plotting (MPP), Mean Square Error (MSE), Method of Moments (MOM), and Maximum Product of Spacing (MPS). Surpyval can, for each type of estimator, take the following types of input data:
| Method | Para/Non-Para | Observed | Censored | Truncated |
|---|---|---|---|---|
| MLE | Parametric | Yes | Yes | Yes |
| MPP | Parametric | Yes | Yes | Limited |
| MSE | Parametric | Yes | Yes | Limited |
| MOM | Parametric | Yes | No | No |
| MPS | Parametric | Yes | Yes | No |
| Kaplan-Meier | Non-Parametric | Yes | Right only | Left only |
| Nelson-Aalen | Non-Parametric | Yes | Right only | Left only |
| Fleming-Harrington | Non-Parametric | Yes | Right only | Left only |
| Turnbull (EM, or EM-ICM without truncation) | Non-Parametric | Yes | Yes | Yes |
SurPyval also offers many different distributions for users, and because of the flexible implementation adding new distributions is easy. Further, the power of SurPyval lay in the robust parameter estimation, as such, some distributions, those that are supported on the half real line, can be offset to make a three- or four-parameter version. The currently available distributions are:
| Distribution | Offsetable |
|---|---|
| Weibull | Yes |
| Normal | No |
| LogNormal | Yes |
| Gamma | Yes |
| Beta | No |
| Beta (4 parameter) | No |
| Uniform | No |
| Exponential | Yes |
| Exponentiated Weibull | Yes |
| Rayleigh | Yes |
| Gumbel (and GumbelLEV) | No |
| Logistic | No |
| LogLogistic | Yes |
Discrete distributions: Poisson, Geometric, NegativeBinomial,
DiscreteWeibull, BetaGeometric, Bernoulli and Binomial (with
a number of trials per row); any
continuous distribution can be discretised with Discretize. Any of them can
be combined in a MixtureModel, and custom distributions are supported.
This project spawned from a Reliaility Engineering project; due to the history of reliability engineers estimating parameters from a probability plot. SurPyval has continued this tradition to ensure that any parametric distribution can have the estimate plotted on a probability plot. These visualisations enable an analyst to get a sense of the goodness of fit of the parametric distribution with the non-parametric distribution.
SurPyval's models can be placed on a set of orthogonal axes. The table below
cross-tabulates four of those axes — the time scale (continuous-time
durations vs discrete-time trials), event recurrence, competing
events, and covariates — against the estimation axis, and fills
each cell with what can be used to implement it. The discrete-time models are
single-event, not recurrent: Bernoulli is a single trial and Binomial is
the sum of n such trials — neither is a recurrent-event process.
A — marks a combination that is
either not applicable (e.g. semiparametric estimation requires covariates) or
not yet built.
| Time | Recurrence | Events | Covariates | Parametric | Semi-/Nonparametric |
|---|---|---|---|---|---|
| Continuous time | Single event | Single | Without | Weibull, Exponential, LogNormal, Gamma, … |
KaplanMeier, NelsonAalen, FlemingHarrington, Turnbull |
| Continuous time | Single event | Single | With | WeibullPH, WeibullAFT, WeibullPO, WeibullAH (every distribution), AcceleratedLife with surpyval.life_models, RoystonParmar, WeibullFrailty |
CoxPH, ProportionalOdds, AdditiveHazards, BuckleyJames, CoxFrailty; survival trees and forests (surpyval.beta.ml) |
| Continuous time | Single event | Competing | Without | ParametricCompetingRisks |
CompetingRisks (CIF) |
| Continuous time | Single event | Competing | With | — | FineGray, CompetingRisksProportionalHazards |
| Continuous time | Recurrent | Single | Without | HPP; NHPP: CrowAMSAA (with growth projection), Duane, CoxLewis; renewal, with each unit's next failure: GeneralizedRenewal, GeneralizedOneRenewal, ARA, ARI |
NonParametricCounting (MCF) |
| Continuous time | Recurrent | Single | With | ProportionalIntensityHPP, ProportionalIntensityNHPP |
— |
| Continuous time | Recurrent | Competing | Without | — | CauseSpecificMCF |
| Continuous time | Recurrent | Competing | With | — | — |
| Discrete time | Single event | Single | Without | Bernoulli (single trial), Binomial (n trials, or a number per row) |
— |
| Discrete time | Single event | Single | With | use logistic / binomial regression (out of scope for this package) | — |
Beyond these axes: the dependence between two lifetimes with copulas
(surpyval.multivariate: Gaussian, Student-t, Clayton, Frank, Gumbel, Joe
and AMH, with rotations, standard errors and confidence bounds), degradation
and remaining useful life (DegradationAnalysis, WienerProcess,
GammaProcess, DestructiveDegradation), and surpyval.forecast, the
expected failures of units in service with prediction intervals, from a
univariate, regression or repairable-system model. The parametric regression
models have qf, cs and quantile_cb, with Wald, likelihood-ratio and
bootstrap bounds.
SurPyval can be installed via pip using the PyPI repository
pip install surpyvalIf you're familiar with survival analysis, and Weibull plotting, the following is a quick start.
from surpyval import Weibull
from surpyval.datasets import load_bofors_steel
# Fetch some data that comes with SurPyval
data = load_bofors_steel()
x = data['x']
n = data['n']
model = Weibull.fit(x=x, n=n, offset=True)
model.plot();SurPyval is well documented, and improving, at the main documentation.
Every model in SurPyval keeps the same rules, so what you learn about one holds for the others:
- Inputs. One data format everywhere (
x,c,n,t,tl,tr). Invalid input raises aValueErrorthat says how to fix it. Missing values gonanin,nanout. Row order, time units and counts-versus-repeated-rows never change an answer. - Outputs. Results keep the shape of the query. The functions of a model agree with each other (
sf + ff = 1,Hf = -log(sf), ...). Probabilities stay in [0, 1] and are monotone in time. Behaviour outside the data is defined and documented. - Estimation. A fit returns the optimum it claims, or says it could not. Every way of fitting a model (
fit,fit_from_df, formulas) gives the same answer. Defaults are the statistically best standard choice and the same everywhere. Conventions follow R'ssurvivaland the other established references. - Uncertainty. Intervals achieve their stated coverage and behave consistently across confidence levels and one- or two-sided bounds.
- Behaviour. One seed rule. Every model round-trips through JSON. Names and defaults are consistent across families. Defaults stay simple; a method for harder cases is added as an option, not a replacement. Warnings are useful and never raw numpy noise. Every public item has a runnable example.
Each principle is enforced by tests, most as properties checked against every registered model. The Design Principles page lists them in full, with the tests that check each one and the open issues where a model does not yet comply.
The bigger ideas (distributional regression, multi-state models, multivariate copulas, ...) are in ROADMAP.md; the issues are for work that can start now.
pip install -r requirements_dev.txt
Run the testing suite by simply executing:
pytestor use coverage to get a coverage report:
coverage run -m pytest # Run pytest under coverage's watch
coverage report # Print coverage report
coverage html # Make a html coverage report (really useful), open htmlcov/index.html- Pip install
pre-commit(it's inrequirements_dev.txtanyways) - Run
pre-commit installwhich sets up the git hook scripts - If you'd like, run
pre-commit run --all-filesto run the hooks on all files - When you go to commit, it will only proceed after all the hooks succeed
Email derryn if you want any features or to see how SurPyval can be used for you.

