Kaplan-Meier step with a risk set restricted to people who have already entered

Counts the risk set at an event time as those who have entered and not yet left. With follow-up intervals open at entry and closed at exit, a person is at risk at t_j when L_i < t_j <= X_i, so the count is the number who entered before t_j less the number who exited before t_j (anyone who left before t_j had also entered before it). The product-limit step itself is HE-FM-ADMC-002; only the risk set changes.

Signature

n_j = N_in - N_out; S_j = S_prev * (1 - d_j / n_j)
Inputs
InputsDefinitionUnit
N_inPeople whose entry time L_i is before t_jcount
N_outPeople whose exit time X_i, by event or censoring, is before t_jcount
S_prevProduct-limit survival after the previous event time, 1 before the firstprobability
d_jEvents recorded at t_jcount
Output
n_jPeople under observation and event free just before event time t_jcount
S_jProduct-limit survival after the step at t_jprobability

Function

Survival estimation from left-truncated records with entry-restricted risk sets

Maps records with an entry time, an exit time and an event indicator to survival from the time origin when people come under observation only after the origin. Each person informs survival only beyond entry, conditional on surviving to entry, so risk sets and likelihood contributions start at entry. The notation follows the Left Truncation article and its six registry patients.

Computational function

  • Computational function: left-truncated Kaplan-Meier curve and restricted mean from entry, exit and event records

    Builds the left-truncated Kaplan-Meier curve and its restricted mean from patient records, applying HE-FM-LTR-002 at every event time. The inputs differ from the formula's: whole vectors of entry times, exit times and event indicators instead of counts at one time.

    Inputs and outputs: entry: Entry time of each record from the origin, 0 for people observed from the origin. Unit: months.; stop: Exit time at the event or censoring, after entry. Unit: months.; status: 1 for an event, 0 for censoring. Unit: indicator.; tau: Restriction time for the restricted mean, no later than the last follow-up. Unit: months.; time, n_risk, n_event, surv: Event times in increasing order, the number at risk, events and survival after each. Unit: months, count, count, probability.; rmst: Restricted mean survival time to tau, the area under the step curve (HE-FM-ADMC-003). Unit: months.

    Assumption: Entry and event times are independent given that the event comes after entry; risk sets near the origin may be small or empty, so the curve should be read only where the risk set has support.

    Worked example (Six registry patients, article example): Entries 0, 0, 0, 12, 12 and 20 months, exits 6, 18, 30, 24, 36 and 33, deaths for A, B, D and F: risk sets 3, 4, 4 and 2, survival 0.6667, 0.5, 0.375 and 0.1875, and a restricted mean to 36 months of 20.9375, as in the article. entry = [0, 0, 0, 12, 12, 20]; stop = [6, 18, 30, 24, 36, 33]; status = [1, 1, 0, 1, 0, 1]; tau = 36; n_risk = [3, 4, 4, 2]; surv = [0.6667, 0.5, 0.375, 0.1875]; rmst = 20.9375

    Worked example (Same records with entry ignored): Setting every entry to 0 gives risk sets 6, 5, 4 and 2, survival 0.8333, 0.6667, 0.5 and 0.25 and a restricted mean of 25.25 months, about 21 per cent more time alive, as in the article. entry = [0, 0, 0, 0, 0, 0]; n_risk = [6, 5, 4, 2]; surv = [0.8333, 0.6667, 0.5, 0.25]; rmst = 25.25

    Excel: With records in EntryTimes, ExitTimes and Status and the distinct event times in EventCol, =COUNTIFS(EntryTimes,"<"&EventCol,ExitTimes,">="&EventCol) gives the risk sets (RiskCol), =COUNTIFS(ExitTimes,EventCol,Status,1) the deaths (DeathCol), a running product of 1 minus DeathCol / RiskCol the survival column, and the restricted mean is the sum of each survival level times the months it lasts up to tau.

    R: lt_km <- function(entry, stop, status, tau) { tj <- sort(unique(stop[status == 1])); n <- sapply(tj, function(t) sum(entry < t & stop >= t)); d <- sapply(tj, function(t) sum(stop == t & status == 1)); S <- cumprod(1-d/n); k <- tj <= tau; w <- diff(c(0, tj[k], tau)); list(time = tj, n_risk = n, n_event = d, surv = S, rmst = sum(c(1, S[k])*w)) } Base R only; lt_km(c(0, 0, 0, 12, 12, 20), c(6, 18, 30, 24, 36, 33), c(1, 1, 0, 1, 0, 1), 36) returns the first example. With the survival package, survfit(Surv(entry, stop, status) ~ 1) gives the same curve.

    Python: def lt_km(entry, stop, status, tau): tj = sorted({x for x, s in zip(stop, status) if s == 1}); n = [sum(1 for l, x in zip(entry, stop) if l < t <= x) for t in tj]; d = [sum(1 for x, s in zip(stop, status) if x == t and s == 1) for t in tj]; S = list(itertools.accumulate([1-a/b for a, b in zip(d, n)], operator.mul)); k = [t for t in tj if t <= tau]; w = [b-a for a, b in zip([0]+k, k+[tau])]; return {"time": tj, "n_risk": n, "n_event": d, "surv": S, "rmst": sum(s*v for s, v in zip([1]+S[:len(k)], w))} Needs import itertools, operator; returns the same values as the R function.

    Test (Risk sets follow the open-interval rule): Every risk set equals the number who entered before its event time less the number who left before it. Expected result: TRUE. FALSE shows risk sets counted from the origin. Excel check: =SUMPRODUCT(--(RiskCol=COUNTIF(EntryTimes,"<"&EventCol)-COUNTIF(ExitTimes,"<"&EventCol)))=ROWS(EventCol)

    Common error (Ordinary Kaplan-Meier on delayed-entry data): Leaving out entry times raises survival at every step, 0.8333 against 0.6667 at month 6 in the example, and carries the bias into restricted means, life-years and QALYs.

    Source: StataCorp. Stata Survival Analysis Reference Manual, Release 19. College Station, TX: Stata Press; 2025. Entry sts, Methods and formulas (n_j at risk just before t_j); Therneau TM. survival: Survival analysis. R package version 3.7-0. 2024. Help page Surv. (counting process intervals (start, end]).

    n_j = sum_i I(L_i < t_j <= X_i); S(t) = prod_(t_j <= t) (1 - d_j / n_j); RMST_tau = sum_k S_k * w_k

Try this function

Implementations

  • Excel

    Entry-restricted risk set and survival step from named ranges

    With entry and exit times in EntryTimes and ExitTimes and EventTime, Deaths and SurvPrev named, the first formula returns the number at risk, held in AtRisk, and the second the survival after the step, held in SurvNew.

    =COUNTIF(EntryTimes,"<"&EventTime)-COUNTIF(ExitTimes,"<"&EventTime); =SurvPrev*(1-Deaths/AtRisk)

Assumptions

  • Entry and event times independent in the risk-set correction

    The corrected risk sets estimate survival from the origin only if entry time and event time are independent given that the event comes after entry; the correction fixes the denominators, not the selection.

  • Follow-up intervals open at entry and closed at exit

    Someone entering exactly at t_j is not at risk at t_j, and someone exiting at t_j is; censoring at t_j is treated as after the event.

Worked examples

  • Risk set at month 6 in the six-patient registry example

    At month 6 only A, B and C have entered and none has left, so 3 are at risk; one death gives survival of 2 out of 3, about 0.6667, as in the article.

    N_in = 3; N_out = 0; d_j = 1; S_prev = 1; n_j = 3; S_j = 0.6667
  • Risk set at month 18 in the six-patient registry example

    By month 18 five patients have entered (F enters at 20) and A has left, so 4 are at risk; survival falls from 0.6667 to 0.5, as in the article.

    N_in = 5; N_out = 1; d_j = 1; S_prev = 0.6667; n_j = 4; S_j = 0.5
  • Risk set at month 33 in the six-patient registry example

    All six have entered and A, B, D and C have left, so E and F are at risk; survival falls from 0.375 to 0.1875, as in the article.

    N_in = 6; N_out = 4; d_j = 1; S_prev = 0.375; n_j = 2; S_j = 0.1875
  • Month 6 when entry times are ignored

    Counting everyone from diagnosis puts all six at risk at month 6, so survival is 5 out of 6, about 0.8333 instead of 0.6667, as in the article.

    N_in = 6; N_out = 0; d_j = 1; S_prev = 1; n_j = 6; S_j = 0.8333

Common errors

  • Starting every record at the origin

    In the example the naive curve ends at 0.25 at 33 months instead of 0.1875; in Betensky and Mandel's simulation of two groups with identical survival, one entering 1.5 years after the origin, a Cox model ignoring delayed entry gave a hazard ratio of 5.01 instead of 1.

  • Dropping late entrants instead of using their entry times

    Restricting the analysis to patients with no delay discards most of a registry cohort and, as Brown and colleagues note, introduces other selection biases.

  • Numbers-at-risk tables that only fall

    With late entry the number at risk can rise as well as fall; a table that only decreases suggests entry times were not used.

Sources

  • Risk set at each event time restricted to those who have entered

    Betensky RA, Mandel M. Recognizing the problem of delayed entry in time-to-event studies: better late than never for clinical neuroscientists. Annals of Neurology. 2015;78(6):839-844. doi:10.1002/ana.24538 (author manuscript read). Time-to-event analysis and delayed entry: the risk set at an event time includes only subjects who have entered the study prior to that event time and have not yet had the event or been censored; Simulation studies: with entry at 1.5 years for group 2, the unadjusted Cox model yields a biased hazard ratio of 5.01 instead of the true 1.

    View source →

  • Counting-process intervals open on the left and closed on the right

    Therneau TM. survival: Survival analysis. R package version 3.7-0. 2024. Help page Surv. Argument time2: intervals are assumed to be open on the left and closed on the right, (start, end], the counting process form used for delayed entry.

    View source →

  • Product-limit estimator with the number at risk just before each failure time

    StataCorp. Stata Survival Analysis Reference Manual, Release 19. College Station, TX: Stata Press; 2025. Entry sts, Methods and formulas: n_j is the number at risk of failure just before time t_j and d_j the number of failures at t_j; Remarks: the enter option adds a number-who-enter column.

    View source →

  • Rising numbers at risk and selection from dropping late entrants

    Brown S, Lavery JA, Shen R, et al. Implications of selection bias due to delayed study entry in clinical genomic studies. JAMA Oncology. 2022;8(2):287-291. doi:10.1001/jamaoncol.2021.5153 (author manuscript read). Discussion: due to staggered entry the number of patients at risk does not consistently decrease over time; restricting analyses to patients who are not left truncated greatly reduces available sample size and introduces other selection biases.

    View source →

Canonical Identity