Signature
n_j = N_in - N_out; S_j = S_prev * (1 - d_j / n_j)
| Inputs | Definition | Unit |
|---|---|---|
N_in | People whose entry time L_i is before t_j | count |
N_out | People whose exit time X_i, by event or censoring, is before t_j | count |
S_prev | Product-limit survival after the previous event time, 1 before the first | probability |
d_j | Events recorded at t_j | count |
n_j | People under observation and event free just before event time t_j | count |
|---|---|---|
S_j | Product-limit survival after the step at t_j | probability |
Function
Survival estimation from left-truncated records with entry-restricted risk sets
Maps records with an entry time, an exit time and an event indicator to survival from the time origin when people come under observation only after the origin. Each person informs survival only beyond entry, conditional on surviving to entry, so risk sets and likelihood contributions start at entry. The notation follows the Left Truncation article and its six registry patients.
Computational function
Computational function: left-truncated Kaplan-Meier curve and restricted mean from entry, exit and event records
Builds the left-truncated Kaplan-Meier curve and its restricted mean from patient records, applying HE-FM-LTR-002 at every event time. The inputs differ from the formula's: whole vectors of entry times, exit times and event indicators instead of counts at one time.
Inputs and outputs:
entry: Entry time of each record from the origin, 0 for people observed from the origin. Unit: months.;stop: Exit time at the event or censoring, after entry. Unit: months.;status: 1 for an event, 0 for censoring. Unit: indicator.;tau: Restriction time for the restricted mean, no later than the last follow-up. Unit: months.;time,n_risk,n_event,surv: Event times in increasing order, the number at risk, events and survival after each. Unit: months, count, count, probability.;rmst: Restricted mean survival time to tau, the area under the step curve (HE-FM-ADMC-003). Unit: months.Assumption: Entry and event times are independent given that the event comes after entry; risk sets near the origin may be small or empty, so the curve should be read only where the risk set has support.
Worked example (Six registry patients, article example): Entries 0, 0, 0, 12, 12 and 20 months, exits 6, 18, 30, 24, 36 and 33, deaths for A, B, D and F: risk sets 3, 4, 4 and 2, survival 0.6667, 0.5, 0.375 and 0.1875, and a restricted mean to 36 months of 20.9375, as in the article.
entry = [0, 0, 0, 12, 12, 20]; stop = [6, 18, 30, 24, 36, 33]; status = [1, 1, 0, 1, 0, 1]; tau = 36; n_risk = [3, 4, 4, 2]; surv = [0.6667, 0.5, 0.375, 0.1875]; rmst = 20.9375Worked example (Same records with entry ignored): Setting every entry to 0 gives risk sets 6, 5, 4 and 2, survival 0.8333, 0.6667, 0.5 and 0.25 and a restricted mean of 25.25 months, about 21 per cent more time alive, as in the article.
entry = [0, 0, 0, 0, 0, 0]; n_risk = [6, 5, 4, 2]; surv = [0.8333, 0.6667, 0.5, 0.25]; rmst = 25.25Excel: With records in EntryTimes, ExitTimes and Status and the distinct event times in EventCol,
=COUNTIFS(EntryTimes,"<"&EventCol,ExitTimes,">="&EventCol)gives the risk sets (RiskCol),=COUNTIFS(ExitTimes,EventCol,Status,1)the deaths (DeathCol), a running product of 1 minus DeathCol / RiskCol the survival column, and the restricted mean is the sum of each survival level times the months it lasts up to tau.R:
lt_km <- function(entry, stop, status, tau) { tj <- sort(unique(stop[status == 1])); n <- sapply(tj, function(t) sum(entry < t & stop >= t)); d <- sapply(tj, function(t) sum(stop == t & status == 1)); S <- cumprod(1-d/n); k <- tj <= tau; w <- diff(c(0, tj[k], tau)); list(time = tj, n_risk = n, n_event = d, surv = S, rmst = sum(c(1, S[k])*w)) }Base R only;lt_km(c(0, 0, 0, 12, 12, 20), c(6, 18, 30, 24, 36, 33), c(1, 1, 0, 1, 0, 1), 36)returns the first example. With the survival package,survfit(Surv(entry, stop, status) ~ 1)gives the same curve.Python:
def lt_km(entry, stop, status, tau): tj = sorted({x for x, s in zip(stop, status) if s == 1}); n = [sum(1 for l, x in zip(entry, stop) if l < t <= x) for t in tj]; d = [sum(1 for x, s in zip(stop, status) if x == t and s == 1) for t in tj]; S = list(itertools.accumulate([1-a/b for a, b in zip(d, n)], operator.mul)); k = [t for t in tj if t <= tau]; w = [b-a for a, b in zip([0]+k, k+[tau])]; return {"time": tj, "n_risk": n, "n_event": d, "surv": S, "rmst": sum(s*v for s, v in zip([1]+S[:len(k)], w))}Needsimport itertools, operator; returns the same values as the R function.Test (Risk sets follow the open-interval rule): Every risk set equals the number who entered before its event time less the number who left before it. Expected result: TRUE. FALSE shows risk sets counted from the origin. Excel check:
=SUMPRODUCT(--(RiskCol=COUNTIF(EntryTimes,"<"&EventCol)-COUNTIF(ExitTimes,"<"&EventCol)))=ROWS(EventCol)Common error (Ordinary Kaplan-Meier on delayed-entry data): Leaving out entry times raises survival at every step, 0.8333 against 0.6667 at month 6 in the example, and carries the bias into restricted means, life-years and QALYs.
Source: StataCorp. Stata Survival Analysis Reference Manual, Release 19. College Station, TX: Stata Press; 2025. Entry sts, Methods and formulas (n_j at risk just before t_j); Therneau TM. survival: Survival analysis. R package version 3.7-0. 2024. Help page Surv. (counting process intervals (start, end]).
n_j = sum_i I(L_i < t_j <= X_i); S(t) = prod_(t_j <= t) (1 - d_j / n_j); RMST_tau = sum_k S_k * w_k
Try this function
Implementations
Excel
Entry-restricted risk set and survival step from named ranges
With entry and exit times in EntryTimes and ExitTimes and EventTime, Deaths and SurvPrev named, the first formula returns the number at risk, held in AtRisk, and the second the survival after the step, held in SurvNew.
=COUNTIF(EntryTimes,"<"&EventTime)-COUNTIF(ExitTimes,"<"&EventTime); =SurvPrev*(1-Deaths/AtRisk)
Assumptions
Entry and event times independent in the risk-set correction
The corrected risk sets estimate survival from the origin only if entry time and event time are independent given that the event comes after entry; the correction fixes the denominators, not the selection.
Follow-up intervals open at entry and closed at exit
Someone entering exactly at t_j is not at risk at t_j, and someone exiting at t_j is; censoring at t_j is treated as after the event.
Worked examples
Risk set at month 6 in the six-patient registry example
At month 6 only A, B and C have entered and none has left, so 3 are at risk; one death gives survival of 2 out of 3, about 0.6667, as in the article.
N_in = 3; N_out = 0; d_j = 1; S_prev = 1; n_j = 3; S_j = 0.6667
Risk set at month 18 in the six-patient registry example
By month 18 five patients have entered (F enters at 20) and A has left, so 4 are at risk; survival falls from 0.6667 to 0.5, as in the article.
N_in = 5; N_out = 1; d_j = 1; S_prev = 0.6667; n_j = 4; S_j = 0.5
Risk set at month 33 in the six-patient registry example
All six have entered and A, B, D and C have left, so E and F are at risk; survival falls from 0.375 to 0.1875, as in the article.
N_in = 6; N_out = 4; d_j = 1; S_prev = 0.375; n_j = 2; S_j = 0.1875
Month 6 when entry times are ignored
Counting everyone from diagnosis puts all six at risk at month 6, so survival is 5 out of 6, about 0.8333 instead of 0.6667, as in the article.
N_in = 6; N_out = 0; d_j = 1; S_prev = 1; n_j = 6; S_j = 0.8333
Common errors
Starting every record at the origin
In the example the naive curve ends at 0.25 at 33 months instead of 0.1875; in Betensky and Mandel's simulation of two groups with identical survival, one entering 1.5 years after the origin, a Cox model ignoring delayed entry gave a hazard ratio of 5.01 instead of 1.
Dropping late entrants instead of using their entry times
Restricting the analysis to patients with no delay discards most of a registry cohort and, as Brown and colleagues note, introduces other selection biases.
Numbers-at-risk tables that only fall
With late entry the number at risk can rise as well as fall; a table that only decreases suggests entry times were not used.
Sources
Risk set at each event time restricted to those who have entered
Betensky RA, Mandel M. Recognizing the problem of delayed entry in time-to-event studies: better late than never for clinical neuroscientists. Annals of Neurology. 2015;78(6):839-844. doi:10.1002/ana.24538 (author manuscript read). Time-to-event analysis and delayed entry: the risk set at an event time includes only subjects who have entered the study prior to that event time and have not yet had the event or been censored; Simulation studies: with entry at 1.5 years for group 2, the unadjusted Cox model yields a biased hazard ratio of 5.01 instead of the true 1.
Counting-process intervals open on the left and closed on the right
Therneau TM. survival: Survival analysis. R package version 3.7-0. 2024. Help page Surv. Argument time2: intervals are assumed to be open on the left and closed on the right, (start, end], the counting process form used for delayed entry.
Product-limit estimator with the number at risk just before each failure time
StataCorp. Stata Survival Analysis Reference Manual, Release 19. College Station, TX: Stata Press; 2025. Entry sts, Methods and formulas: n_j is the number at risk of failure just before time t_j and d_j the number of failures at t_j; Remarks: the enter option adds a number-who-enter column.
Rising numbers at risk and selection from dropping late entrants
Brown S, Lavery JA, Shen R, et al. Implications of selection bias due to delayed study entry in clinical genomic studies. JAMA Oncology. 2022;8(2):287-291. doi:10.1001/jamaoncol.2021.5153 (author manuscript read). Discussion: due to staggered entry the number of patients at risk does not consistently decrease over time; restricting analyses to patients who are not left truncated greatly reduces available sample size and introduces other selection biases.
Canonical Identity
Stable URI · Machine-readable · Resolvable · CC BY 4.0