Skip to contents

Introduction

When clinical trials include overall survival (OS) as a secondary or exploratory endpoint, regulators may recommend not only monitoring for early evidence of efficacy and futility, but also for potential harm — that is, evidence that the experimental treatment may be worsening survival relative to control. This article demonstrates how the gsDesign package supports group sequential designs with three boundaries: an efficacy (upper) bound, a futility (lower) bound, and a harm bound, using test.type = 7 (binding) and test.type = 8 (non-binding). For sparse mortality data without a separate futility rule, the section Sparse-event mortality monitoring uses test.type = 6 and exact-binomial boundary calculations.

Regulatory context: FDA guidance on OS monitoring in oncology

The August 2025 FDA draft guidance Approaches to Assessment of Overall Survival in Oncology Clinical Trials (U.S. Food and Drug Administration 2025) discusses OS as a safety endpoint even when another endpoint establishes efficacy. It remains draft, nonbinding guidance, not a requirement to use the method illustrated here. Sections II, III.A, and III.B.2 discuss:

  • Indolent diseases and long survival, where demonstrating OS superiority may be impractical.
  • Prespecified harm assessments, justified thresholds, sufficient events, and precision to assess clinically relevant mortality detriment.
  • Event-driven interim analyses when appropriate, with independent data monitoring. Some small or low-event-rate trials may not warrant interim OS analyses.
  • The distinction between identifying a harm signal and ruling out an unacceptable degree of harm.

The guidance does not prescribe 0.1/Pocock spending or a separate futility boundary in every trial. Types 7 and 8 are useful when both futility and harm are wanted; type 6 can instead supply a non-binding harm bound alone on the lower side. The design and its assumptions require context-specific justification.

The harm bound implemented in gsDesign is a new method that is easy to use — a principled, straightforward extension of the widely used group sequential spending function framework. While we believe this approach is understandable, useful, and flexible, other methods for monitoring potential harm may also be considered. However, there are limitations with this approach. The first example has higher mortality risk than many cases. The later sparse-event example examines what the approach can and cannot establish when deaths are uncommon.

Design framework overview

In a standard two-sided asymmetric group sequential design (test.type = 3 or 4), there are two boundaries:

  • Efficacy (upper) bound: Reject \(H_0\) if the test statistic exceeds this boundary (evidence of treatment benefit).
  • Futility (lower) bound: Stop for futility if the test statistic falls below this boundary (insufficient evidence of treatment benefit).

The harm bound extension (test.type = 7 or 8) adds a third boundary:

  • Harm bound: Signal that the experimental treatment may be harming patients (evidence of a detrimental effect).

The harm bound lies below the futility bound. At each analysis, there are four possible outcomes:

  1. Cross the efficacy bound (above): Stop for efficacy.
  2. Between the efficacy and futility bounds: Continue the trial.
  3. Cross the futility bound but not the harm bound (between futility and harm): Stop for futility.
  4. Cross the harm bound (below): Stop for harm.

The harm bound is intended to flag sufficiently small observed p-values favoring control. The nominal threshold depends on the analysis schedule and spending specification; it is not a fixed p-value cutoff at every look. That is, the harm bound flags evidence that the experimental treatment may be worsening survival — a negative treatment effect on the log hazard ratio scale.

Design with non-binding bounds (test.type = 8)

We demonstrate a survival design using gsSurvCalendar() with test.type = 8 (non-binding futility and harm bounds). The scenario is based on a 1:1 randomized trial monitoring overall survival with:

  • Median control survival: 3 years (36 months), i.e., \(\lambda_C = \log(2)/36\).
  • Target hazard ratio: HR = 0.75 (25% reduction in hazard).
  • Power: 90% (\(\beta = 0.1\)).
  • One-sided \(\alpha\): 0.0125 (e.g., the OS component of a trial with multiplicity adjustment).
  • Enrollment: Uniform enrollment over 18 months.
  • Study duration: 5 years (60 months) with planned analyses at years 1, 2, 3, 4, and 5 from start of enrollment.

The astar parameter controls the total spending for the harm bound under \(H_0\). The defaults are astar = 0.1 (also selected by input astar = 0) and sfharm = sfLDPocock. We specify these explicitly below for clarity. The 10% target describes harm crossing under \(H_0\) if futility stopping is ignored, while still accounting for earlier efficacy and harm stops. Actual harm stopping can be less frequent when futility is followed. The existing cap that keeps harm at or below an active futility boundary can also reduce the attainable spending.

Spending function specification

We specify:

  • Efficacy bound: Lan-DeMets O’Brien-Fleming (sfLDOF) spending function (conservative, spending little \(\alpha\) at early analyses).
  • Futility bound: Hwang-Shih-DeCani (HSD) spending function with \(\gamma = -2\) (moderate \(\beta\)-spending under \(H_1\)).
  • Harm bound: Lan-DeMets Pocock (sfLDPocock) spending function (spending under \(H_0\) for detecting harm).
x8 <- gsSurvCalendar(
  test.type = 8,
  alpha = 0.0125,
  beta = 0.1,
  astar = 0.1,
  calendarTime = c(12, 24, 36, 48, 60),
  sfu = sfLDOF,
  sfl = sfHSD, sflpar = -2,
  sfharm = sfLDPocock,
  lambdaC = log(2) / 36,
  hr = 0.75,
  R = 18,
  minfup = 42
)

Summary

The summary() method provides a concise description of the design:

cat(strwrap(summary(x8), width = 65), sep = "\n")
#> Asymmetric two-sided group sequential design with non-binding
#> futility and harm bounds, 5 analyses, time-to-event outcome with
#> sample size 1148 and 657 events required, 90 percent power, 1.25
#> percent (1-sided) Type I error (sample size/power method:
#> Lachin-Foulkes) to detect a hazard ratio of 0.75. Enrollment and
#> total study durations are assumed to be 18 and 60 months,
#> respectively. Efficacy bounds derived using a Lan-DeMets
#> O'Brien-Fleming approximation spending function (no parameters).
#> Futility bounds derived using a Hwang-Shih-DeCani spending
#> function with gamma = -2. Harm bounds derived using a Lan-DeMets
#> Pocock approximation spending function.

Detailed boundary table

The gsBoundSummary() function produces a tabular summary with columns for each boundary. By default, B-value, Spending, CP, CP H1, and PP are excluded. We note that for the first interim analysis, the efficacy bound is so extreme it is effectively impossible to cross. However, the harm and futility bounds are more moderate, allowing for early stopping if there is evidence of harm or futility. The futility bound is an indicator of why bounds are often non-binding — the futility bound is not intended to be a strict stopping rule, but rather a signal that the trial may be unlikely to succeed if it continues. Crossing the harm bound is a stronger indication that the treatment may be harmful, and the trial should be at least paused with a recommendation to review the safety and other endpoint data.

Conditional power (CP, CP H1) and predictive power (PP) can also be included in the summary. Below we show the full table with all statistics, including conditional and predictive power at each boundary:

gsBoundSummary(x8, exclude = c()) |> lt()

Interpreting the boundaries

The design has five analyses at calendar times of 12, 24, 36, 48, and 60 months. At each analysis, the test statistic (Z-value) is compared against three boundaries:

bounds <- data.frame(
  Analysis = 1:x8$k,
  Month = x8$T,
  Events = ceiling(x8$n.I),
  Harm = round(x8$harm$bound, 2),
  Futility = round(x8$lower$bound, 2),
  Efficacy = round(x8$upper$bound, 2)
)
bounds |>
  lt() |>
  lt_header("Z-value boundaries at each analysis")

Decision rules at an analysis where all three bounds are active:

  • If \(Z >\) efficacy bound: Stop for efficacy (reject \(H_0\)).
  • If futility bound \(< Z \leq\) efficacy bound: Continue the trial.
  • If harm bound \(< Z \leq\) futility bound: Stop for futility.
  • If \(Z \leq\) harm bound: Stop for harm.

When both lower bounds are active, the harm bound is always at or below the futility bound. If futility is skipped but harm is tested, the harm bound is the sole active lower stopping boundary. The harm and futility bounds may coincide when the uncapped harm boundary would exceed the futility boundary; the cap can prevent full harm spending.

Boundary crossing probabilities

We examine the operating characteristics under two scenarios: no treatment effect (HR = 1, i.e., under \(H_0\)) and the design alternative (HR = 0.75). When harm and futility are both active, x8$lower$prob and x8$harm$prob are reported as mutually exclusive stopping outcomes. Thus, the probability of crossing the futility threshold is the sum of the two lower-tail components.

probs <- data.frame(
  Scenario = c(rep("Under H0 (HR=1)", x8$k), rep("Under H1 (HR=0.75)", x8$k)),
  Analysis = rep(1:x8$k, 2),
  Month = rep(x8$T, 2),
  `P(Efficacy)` = c(cumsum(x8$upper$prob[, 1]), cumsum(x8$upper$prob[, 2])),
  `P(Futility only)` = c(cumsum(x8$lower$prob[, 1]), cumsum(x8$lower$prob[, 2])),
  `P(Harm)` = c(cumsum(x8$harm$prob[, 1]), cumsum(x8$harm$prob[, 2])),
  `P(Futility or Harm)` = c(
    cumsum(x8$lower$prob[, 1] + x8$harm$prob[, 1]),
    cumsum(x8$lower$prob[, 2] + x8$harm$prob[, 2])
  ),
  check.names = FALSE
)
probs |>
  lt() |>
  lt_format(columns = 2:7, decimals = 4) |>
  lt_header("Cumulative boundary crossing probabilities")

Under \(H_0\), the actual cumulative probability of stopping for harm is approximately 0.0417 when all active stopping rules are followed. This need not equal astar: trials that stop for futility cannot subsequently stop for harm. The cumulative probability of crossing the futility threshold, inclusive of harm, is approximately 0.9888. Under \(H_1\) (HR = 0.75), crossing the harm bound is very unlikely (4^{-4}), since the treatment is beneficial.

Harm calibration versus actual stopping

For both binding and non-binding designs, harm calibration ignores futility stopping. This avoids making a harm threshold near zero merely to spend a specified amount among the selected trials that survived earlier futility monitoring. Efficacy and earlier harm stopping are still included.

The following example tests harm at all three looks, skips efficacy at IA1, and tests futility only at IA1. It uses the new harm defaults. We compare the target spending, harm crossing without futility, and actual harm stopping:

xh <- gsDesign(
  test.type = 8, timing = c(.5, .75),
  testUpper = c(FALSE, TRUE, TRUE),
  testLower = c(TRUE, FALSE, FALSE),
  testHarm = TRUE
)
harm_without_futility <- gsProbability(
  k = xh$k, theta = 0, n.I = xh$n.I,
  a = xh$harm$bound, b = xh$upper$bound
)
data.frame(
  Analysis = c("IA1", "IA2", "Final"),
  `Harm Z` = xh$harm$bound,
  `Nominal harm p` = pnorm(xh$harm$bound),
  `Cumulative target` = cumsum(xh$harm$spend),
  `Harm without futility` = cumsum(harm_without_futility$lower$prob[, 1]),
  `Harm with futility` = cumsum(xh$harm$prob[, 1]),
  check.names = FALSE
) |> lt() |> lt_format(columns = 2:6, decimals = 4)

At IA2 the probability of an actual harm stop includes only paths that continued past IA1. Thus the cumulative harm probability is \(P_0(Z_1 \le h_1) + P_0(Z_1 > f_1, Z_2 \le h_2)\), where \(h_i\) denotes a harm boundary and \(f_1\) the IA1 futility boundary. Calibration without futility instead uses \(h_1\) in place of \(f_1\) in the second term.

gsBoundSummary() reports actual cumulative stopping probabilities, not the hypothetical probabilities used to calibrate harm spending. Its optional Spending rows report the spending targets, so these two rows need not agree. Also, its p (1-sided) rows use the efficacy direction; the nominal p-value favoring harm is pnorm(xh$harm$bound), the complement of that displayed p-value. print(xh) labels the harm-direction values as Nominal p.

Changing the example to test.type = 7 retains the same harm-calibration convention. Binding status changes efficacy calibration: type 7 accounts for both lower stopping rules in protecting alpha, whereas type 8 ignores them for that purpose. Power and actual stopping probabilities include both rules in either design. Ignoring a binding futility rule in practice can inflate efficacy Type I error, even though it was ignored for harm calibration.

Visualization

All standard plot() types are supported for test.type = 7 and 8 designs, with a third line (or set of lines) shown for the harm bound.

Z-value boundaries

The default plot shows Z-value boundaries at each analysis. Three boundaries are displayed: efficacy (upper), futility (lower), and harm (below futility).

plot(x8)
Z-value boundaries for non-binding harm bound design

Z-value boundaries for non-binding harm bound design

Boundary crossing probabilities

The power plot (plottype = 2) shows cumulative boundary crossing probabilities as a function of the treatment effect. Three sets of lines appear: upper bound (cumulative efficacy crossing probability), 1-(Futility or harm), and 1-Harm. Because the harm boundary is nested below the futility boundary when both are active, crossing the futility threshold includes both futility-only and harm stops. The 1-(Futility or harm) curve therefore subtracts x8$lower$prob + x8$harm$prob, while the 1-Harm curve subtracts harm crossings only. This aggregation is performed only for plotting: the probability arrays stored in x8 remain mutually exclusive so that efficacy, futility-only, and harm outcomes add without double counting. When the underlying treatment effect favors control, the high probability of crossing the harm bound indicates that the harm bound is sensitive and serves its intended purpose.

plot(x8, plottype = 2)
Boundary crossing probabilities for non-binding harm bound design

Boundary crossing probabilities for non-binding harm bound design

Approximate treatment effect at boundaries

The effect size plot (plottype = 3) shows the approximate treatment effect at each boundary. For survival designs, this is expressed as the approximate hazard ratio at the boundary.

plot(x8, plottype = 3)
Approximate treatment effect at boundaries

Approximate treatment effect at boundaries

Conditional power at boundaries

Conditional power (plottype = 4) at each interim analysis is shown for all three boundaries. This is generally not a very useful plot.

plot(x8, plottype = 4)
Conditional power at boundaries

Conditional power at boundaries

Spending function plot

The spending function plot (plottype = 5) shows the three spending functions: \(\alpha\) (efficacy), \(\beta\) (futility), and harm.

plot(x8, plottype = 5)
Spending functions for non-binding harm bound design

Spending functions for non-binding harm bound design

B-values at boundaries

B-values (plottype = 7) are Z-values scaled by \(\sqrt{t}\) where \(t\) is the information fraction. As discussed by Proschan et al. (2006), the expected value of B-values increases linearly with the information fraction under the assumption of a constant treatment effect (proportional hazards). This linear relationship makes B-values useful for visual assessment of treatment effect trends across interim analyses: departures from linearity may suggest non-proportional hazards or other changes in treatment effect over time. Three boundary lines are shown: efficacy, futility, and harm.

plot(x8, plottype = 7)
B-values at boundaries

B-values at boundaries

Design with binding bounds (test.type = 7)

For test.type = 7, both the futility and harm bounds are binding — meaning the computation of the efficacy bound assumes the trial will stop if either bound is crossed. This yields a slightly less conservative efficacy bound (easier to cross), but at the cost of inflated Type I error if the stopping rule is not strictly followed.

We first create a binding design with \(\alpha = 0.0125\) to compare with the non-binding design above:

x7 <- gsSurvCalendar(
  test.type = 7,
  alpha = 0.0125,
  beta = 0.1,
  astar = 0.1,
  calendarTime = c(12, 24, 36, 48, 60),
  sfu = sfLDOF,
  sfl = sfHSD, sflpar = -2,
  sfharm = sfLDPocock,
  lambdaC = log(2) / 36,
  hr = 0.75,
  R = 18,
  minfup = 42
)

Comparing binding and non-binding

comparison <- data.frame(
  Bound = c("Efficacy", "Futility", "Harm"),
  `Binding (type 7)` = c(
    paste(round(x7$upper$bound, 3), collapse = ", "),
    paste(round(x7$lower$bound, 3), collapse = ", "),
    paste(round(x7$harm$bound, 3), collapse = ", ")
  ),
  `Non-binding (type 8)` = c(
    paste(round(x8$upper$bound, 3), collapse = ", "),
    paste(round(x8$lower$bound, 3), collapse = ", "),
    paste(round(x8$harm$bound, 3), collapse = ", ")
  ),
  check.names = FALSE
)
comparison |>
  lt() |>
  lt_header("Comparison of binding vs. non-binding Z-value boundaries")

Note that the efficacy bounds for test.type = 7 (binding) are slightly lower (easier to cross) than for test.type = 8 (non-binding). The maximum number of events for test.type = 7 (639) is also slightly smaller than for test.type = 8 (657), reflecting the assumption that the trial will stop at the lower bounds.

Efficacy bounds at alternate \(\alpha\) levels

The gsBoundSummary() function accepts an alpha argument to display efficacy bounds at one or more alternate \(\alpha\) levels alongside the original design. Each alternate-alpha column retains the testUpper schedule from the design, so an efficacy analysis that was skipped remains inactive and every efficacy characteristic at that analysis, including cumulative crossing probability, is NA. Here we show the non-binding design (x8) with efficacy bounds for both \(\alpha = 0.0125\) (the design level) and \(\alpha = 0.025\):

gsBoundSummary(x8, alpha = 0.025) |> lt()

Alternate-alpha summaries are supported for test.type = 8, where both futility and harm bounds are non-binding. They are not supported for test.type = 7: a binding harm or futility bound is outside the non-binding sequential-p-value framework used for Maurer–Bretz graphical multiple testing.

Sparse-event mortality monitoring

A harm bound without a separate futility bound

Consider a randomized trial with a primary efficacy endpoint other than OS, such as progression-free survival, and relatively few expected deaths. A prespecified mortality harm signal can be useful without requiring a separate rule for stopping because OS benefit appears unlikely. This is a hypothetical illustration, not an FDA-specified disease scenario or recommended protocol.

Use test.type = 6: its lower Z-bound spends under the null hypothesis and is non-binding for efficacy Type I error. Here we interpret that lower bound as harm; there is no additional beta-spending futility bound. For type 6 the inputs are astar = 0.1 and sfl = sfLDPocock, not sfharm. They must be supplied explicitly: type 6 retains its general-purpose defaults of astar = 1 - alpha (selected by input 0) and HSD lower spending. The lower boundary is stored in lower, even if a generic summary labels it as futility. Its clinical interpretation in this example is harm.

We plan analyses at 6, 23, and 40 total deaths, with 1:1 randomization, enrollment of 100 participants per month for 16 months, no dropout, and a constant control mortality hazard of 0.001 per month. An HR of 0.8 is used only as a planning scenario for the expected timing and power, not as evidence of mortality benefit. gsSurvPower() evaluates this fixed event schedule; it does not increase the event target to obtain 90% OS power.

sparse <- gsSurvPower(
  k = 3, test.type = 6,
  alpha = .025, sided = 1,
  astar = .1, sfl = sfLDPocock,
  targetEvents = c(6, 23, 40),
  lambdaC = .001, hr = .8, hr0 = 1,
  gamma = 100, R = 16, ratio = 1, eta = 0
)
# Use the planned final event count as the explicit spending denominator.
event_fraction <- sparse$n.I / tail(sparse$n.I, 1)
sparse_exact <- toBinomialExact(
  sparse, usTime = event_fraction, lsTime = event_fraction
)

The example retains an OS efficacy boundary to illustrate non-binding harm monitoring alongside an alpha-controlled efficacy test. It does not allocate alpha to another primary endpoint or supply a multiple-endpoint testing strategy. If OS is solely a safety endpoint, the efficacy stopping rules and operating characteristics should be specified accordingly rather than claiming an OS efficacy objective that the trial does not have.

Integer death-count boundaries

In the exact-binomial representation the direction reverses: a high count of experimental-arm deaths signals harm. upper$bound is therefore the harm cutoff, while lower$bound is the efficacy cutoff. The probabilities account for stopping at any earlier active boundary.

data.frame(
  `Total deaths` = sparse_exact$n.I,
  `Experimental deaths for harm (at least)` = sparse_exact$upper$bound,
  `Cumulative harm probability under HR=1` =
    cumsum(sparse_exact$upper$prob[, 1]),
  check.names = FALSE
) |> lt() |> lt_format(columns = 3, decimals = 4)

# Efficacy Type I error is evaluated ignoring the non-binding harm boundary.
efficacy_without_harm <- gsBinomialExact(
  k = sparse_exact$k, theta = .5, n.I = sparse_exact$n.I,
  a = sparse_exact$lower$bound, b = sparse_exact$n.I + 1L
)
stopifnot(sum(efficacy_without_harm$lower$prob) <= sparse$alpha + 1e-10)

At the first look, a signal requires all six deaths to be in the experimental arm. At later looks the cumulative cutoff is applied only if no earlier stopping boundary has been crossed. The total null harm-crossing probability is 7.8%, below the 10% target because only integer death counts are available. The efficacy Type I error without harm stopping is 1.96%, below 2.5%.

Sensitivity to excess mortality

For balanced exposure and uncommon events, approximate the probability that a death is in the experimental arm by \(p = HR/(1 + HR)\), where \(HR = \lambda_E/\lambda_C\). More generally the corresponding conditional Poisson model uses \(p = HR\,r/(1 + HR\,r)\) with exposure ratio \(r\). The recursion below is exact for independent binomial event increments with constant \(p\), not automatically exact for arbitrary censored survival data. Risk-set imbalance, differential censoring, or changing hazards can invalidate the simple mapping from randomization ratio and HR. Such settings need appropriate survival-model calculations or simulation.

harm_hr <- c(1, 1.25, 1.5, 2)
harm_oc <- gsBinomialExact(
  k = sparse_exact$k, theta = harm_hr / (1 + harm_hr),
  n.I = sparse_exact$n.I,
  a = sparse_exact$lower$bound, b = sparse_exact$upper$bound
)
data.frame(
  `True mortality HR` = harm_hr,
  `Probability of a harm signal` = colSums(harm_oc$upper$prob),
  check.names = FALSE
) |> lt() |> lt_format(columns = 2, decimals = 3)

Even at HR = 1.5, the chance of a signal is only 38.8% in this example. Failure to cross the boundary therefore cannot establish that a clinically important mortality increase has been excluded. Increasing harm spending can improve sensitivity, but also increases false signals; it cannot replace information from additional deaths and follow-up. A non-binding designation concerns efficacy alpha calculations, not the clinical importance of reviewing a signal.

Detecting harm versus ruling out harm

A harm-monitoring test asks whether data provide evidence of excess mortality relative to HR = 1. A reassurance objective instead asks whether the data can exclude a prespecified unacceptable HR, such as 1.5. The latter value is an illustration, not a universal safety margin. One approach is to require an appropriately constructed upper confidence bound to fall below that margin.

To illustrate the lack of precision, suppose a single fixed analysis has 20 deaths in each arm. Under the same binomial model, transform a one-sided 95% upper bound for the experimental share into an upper HR bound:

p_upper <- binom.test(20, 40, alternative = "less", conf.level = .95)$conf.int[2]
hr_upper <- p_upper / (1 - p_upper)
data.frame(
  `Experimental deaths` = 20, `Control deaths` = 20,
  `One-sided 95% upper HR bound` = hr_upper,
  check.names = FALSE
) |> lt() |> lt_format(columns = 3, decimals = 2)

Despite a balanced observed count, this upper bound exceeds 1.5. This is a fixed-analysis precision illustration, not a sequentially adjusted interval for the monitored trial. Repeated or stopping-selected reassurance assessments require an interval or test valid for that monitoring plan. A vignette example must not turn “no harm signal” into a claim of safety: the harm margin, precision, event target, follow-up, and operating characteristics for ruling out harm all need their own justification.

Practical considerations

Choice of spending functions

The choice of spending functions for the three boundaries should reflect regulatory and scientific considerations:

  • Efficacy: A conservative spending function such as Lan-DeMets O’Brien-Fleming (sfLDOF) is typical, spending very little \(\alpha\) at early interim analyses when limited information is available.
  • Futility: Moderate spending (e.g., HSD with \(\gamma = -2\)) allows early stopping for futility when the treatment effect is clearly absent.
  • Harm: The Lan-DeMets Pocock (sfLDPocock) spending function provides more aggressive spending at early analyses, which is appropriate for harm monitoring since detecting a detrimental effect early is critical for patient safety.

Total harm spending and the number of analyses

Lan–DeMets Pocock spending with total 0.1 allocates cumulative probability \(A(t) = 0.1\log\{1 + (e-1)t\}\) by spending time \(t\). It allocates relatively more spending early than a conservative O’Brien–Fleming choice, helping to flag moderately small nominal p-values favoring harm throughout monitoring. The total 0.1 is a repeated-monitoring probability under the null, not a nominal 0.1 cutoff at every analysis. Neither the total nor the spending function guarantees that any particular nominal p-value will trigger harm stopping at every look.

With more analyses, a user may wish to choose a larger total harm spending (astar) to retain similarly permissive nominal thresholds. This increases the allowed probability of a false harm signal under the null. It does not mean increasing futility spending (beta), which has a different purpose and is calibrated under the design alternative. Inspect the actual thresholds for the intended timing and number of analyses. For example:

harm_nominal <- function(k, astar) {
  d <- gsDesign(k = k, test.type = 8, astar = astar)
  data.frame(
    Analyses = k, Total = astar, Look = seq_len(k),
    `Nominal harm p` = pnorm(d$harm$bound), check.names = FALSE
  )
}
rbind(harm_nominal(3, .1), harm_nominal(6, .1), harm_nominal(6, .2)) |>
  lt() |> lt_format(columns = 4, decimals = 4)

The larger total in this illustration is a sensitivity analysis, not a universal recommendation. The choice requires considering both sensitivity to potential harm and false harm signals.

Interpreting the harm bound

The harm bound is crossed when the observed p-value favoring control is at or below the analysis-specific nominal harm threshold. In terms of the test statistic, a negative Z-value indicates that the hazard rate is higher in the experimental arm than the control arm — i.e., the experimental treatment appears to be worsening survival. When the Z-value falls below the harm bound, this constitutes a statistical signal that the treatment may be harmful, and the trial should be stopped with a recommendation to review the safety data.

The harm spending is computed under \(H_0\) (no treatment effect), ignoring futility stopping but retaining efficacy stopping. It budgets the probability of a false harm signal under those assumptions; following futility can reduce the actual probability further.

Harm bound capping

In the implementation, the harm bound is automatically capped so it never exceeds the futility bound when both are active. This ensures the ordering harm bound \(\leq\) futility bound \(\leq\) efficacy bound at analyses where all three are tested, while allowing harm to be the active lower boundary when futility is skipped.

When to use test.type = 7 vs. test.type = 8

  • test.type = 8 (non-binding) is most often preferred in practice. Regulators will generally expect non-binding bounds, which preserve Type I error control regardless of whether the stopping rules are strictly followed. Since Data Monitoring Committees (DMCs) typically retain discretion to continue or stop a trial based on the totality of the evidence, the non-binding approach ensures that the statistical validity of the efficacy analysis is maintained even if a futility or harm boundary is crossed but the trial continues.
  • test.type = 7 (binding) is appropriate when there is a firm commitment to stop the trial upon crossing any boundary. This provides a small efficiency gain (slightly easier efficacy bounds and fewer required events) but requires strict protocol adherence. If the trial does not stop after crossing a binding boundary, Type I error may be inflated.

In most regulatory settings, test.type = 8 is the safer and more common choice.

Binding and non-binding harm monitoring

When futility and harm are both tested at an analysis, the harm boundary lies below the futility boundary and therefore does not add another stopping region. Crossing probabilities are nevertheless partitioned into mutually exclusive harm, futility, and efficacy outcomes.

When harm is tested at an analysis where futility is skipped, the harm boundary is the active lower stopping boundary. This stopping probability is included in sample-size derivation and, for test.type = 7, in the binding efficacy-bound calculation. For test.type = 8, efficacy bounds continue to use the non-binding Type I error convention, while power and expected information account for actual harm and futility stopping.

Adjusting the boundaries

The boundaries are adjustable through several design parameters:

  • Alternate astar: Controls the Type I error allocated to excess OS harm detection.
  • Alternate spending functions: Different spending functions for efficacy, futility, and harm boundaries change the aggressiveness of each boundary across analyses.
  • Alternate timing of analyses: Changing the calendar times of interim analyses shifts the information available at each look.

Regardless of the statistical design, bounds must be clinically, ethically, and statistically sound. As previously noted, this approach is one option to address the regulatory expectation for OS harm monitoring, but other approaches may also be considered.

References

Proschan, Michael A., K. K. Gordon Lan, and Janet Turk Wittes. 2006. Statistical Monitoring of Clinical Trials: A Unified Approach. Springer.
U.S. Food and Drug Administration. 2025. Approaches to Assessment of Overall Survival in Oncology Clinical Trials: Draft Guidance for Industry. Https://www.fda.gov/media/188274/download.