A bootstrap study of logistic regression for heart-failure mortality

ST 453 final project, April 2026 · R with no modeling packages · 299 patients, 96 deaths

The question

Does the standard confidence interval on a logistic regression coefficient hold up when the data have heavy outliers? One predictor here, serum creatinine, runs up to 9.4 mg/dL against a median of 1.1, so the normal approximation behind the Wald interval is exactly the thing to doubt.

The data

Heart Failure Clinical Records from the UCI Machine Learning Repository: 299 patients from a single cardiology center in Faisalabad, Pakistan (Ahmad et al., 2017; Chicco and Jurman, 2020), 13 variables, and a binary outcome, death during follow-up, at 32.1% (96 of 299). I used four predictors: age, ejection fraction, serum creatinine and follow-up time, each standardized once to mean 0 and SD 1 using the original sample.

Dataset at UCI

What I built

Everything was written by hand in base R, since no packages were allowed. The fit is Newton-Raphson from β = 0, with each step's linear system solved by Gaussian elimination with partial pivoting and back substitution, a routine I had written for an earlier homework. It converges when the step is under 1e-8, which took about 7 iterations on the real data. The covariance matrix comes from the inverse Hessian at the solution, solved column by column.

The bootstrap is a pairs (case) bootstrap: resample the 299 rows with replacement, refit with the same routine, repeat 1,000 times with seeds 1 through 1,000. All 1,000 fits converged. For the BCa interval, the acceleration constant comes from a leave-one-out jackknife, which is 299 more fits.

What the bootstrap showed

Coefficient estimates, Wald standard errors and bootstrap standard deviations, N = 1,000
CoefficientEstimateWald SEBootstrap SDRatio (Bootstrap SD ÷ Wald SE)
Intercept−1.29020.19530.21771.11
Age0.51540.17690.17000.96
Ejection fraction−0.88530.18410.21441.16
Serum creatinine0.74460.18060.28181.56
Time−1.59970.22370.26211.17
Serum creatinine: 95% intervals by method, N = 1,000
MethodLowerUpperWidth
Wald0.39061.09870.7081
Percentile0.33101.44931.1183
BCa0.29801.34321.0452

For three of the five coefficients the Wald inference was essentially right. For serum creatinine the bootstrap SD came out 56% larger than the Hessian standard error, so the normal-approximation interval understates the uncertainty, and the percentile or BCa interval is the honest one. None of the qualitative conclusions changed, every coefficient stays significant under either framework, but the band on serum creatinine is a good deal wider than the textbook would have told you.

Run it yourself

This page carries the same 299 rows and the same estimator, rewritten in JavaScript. Press run and it refits the model on N resampled datasets in your browser, then draws the intervals. The random draws differ from the R run, so the bootstrap intervals will not match the report exactly; the ends can move by a few hundredths or more, depending on N and the seed. The report's values are shown beside them for comparison.

The live demo needs JavaScript and a browser that supports module Web Workers. The report's numbers are in the tables above.

Limitations, in my own words

Only 4 of the 13 variables went into the model, so the omitted ones (diabetes, anaemia, smoking, sex, blood pressure, platelets, sodium, creatinine phosphokinase) may be inflating some effects. Follow-up time is really a censoring problem, and a survival model would be the proper tool. I assumed the log-odds are linear in each predictor and did not test that. It is one center and 299 patients, and no amount of resampling fixes who got recruited. And the serum creatinine result should be replicated with N = 5,000 or 10,000 before anyone leans on it.

Files