AI in Healthcare

Irregular clinical time series

Vitals are measured whenever a nurse decides to check, not on a fixed clock, so the gaps between readings carry as much signal as the readings themselves.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A hospital vitals chart has no fixed timetable — readings arrive whenever someone decides to check.

Think about waiting for a bus that runs "whenever it runs" versus a train that comes every fifteen minutes exactly. You can plan around a train. A random bus forces you to keep watching.

A patient's vitals work like that random bus. A stable patient might get checked once every four hours. A worrying one gets checked every fifteen minutes. The checking itself is a decision, made by a person who was already a little concerned.

Why it exists

Most time series lessons assume evenly spaced data: one temperature reading every hour, one stock price every minute. That assumption makes the maths simpler, and it is often false in a hospital.

A nurse does not check vitals on a timer. She checks more often when something looks wrong, and less often when a patient looks fine. That means the frequency of checking is itself a clue. A patient checked every ten minutes is probably sicker than one checked every six hours.

Force this kind of data onto a fixed grid, and that clue disappears. Copying the last known value forward invents hours of "normal" readings that were never taken.

How it works

Real readings:    08:00  08:15  08:20  ................  11:00  ......................................  19:45
                    78     92    101                        88                                            84

Forced onto a
fixed hourly
grid, gaps
filled forward:   08:00  09:00  10:00  11:00  12:00  13:00 ... 18:00  19:00
                    90     90     90     88     88     88  ...  88     84

Notice the fabricated 88s from 12:00 to 18:00. Nobody measured this patient for over eight hours. The forward-filled grid hides that gap completely, and hides the fact that nobody was worried enough to check.

Where you have already seen it

  • A fitness tracker's heart rate graph looks smooth because it samples constantly. A hospital monitor chart looks jagged for the opposite, more informative reason: humans decided when to look.
  • A doctor glancing at a chart reads the gaps as much as the numbers. A long silent stretch usually means "nothing to worry about happened here."

An honest warning

Treating "missing" and "stable" as the same thing is a common mistake in clinical machine learning. It is easy to make without noticing.

A feature like "minutes since the last reading" is often more useful to a model than the reading's actual value. That single number carries a piece of a nurse's judgement, encoded as a timestamp.

None of this makes a model ready for real patient care on its own. A model built on features like these still needs clinical validation and regulatory review before it touches a real decision.

Remember this

  • Readings arrive when a human decides to check, not on a fixed clock — the timing itself is informative.
  • Forward-filling across a long gap invents data that was never measured.
  • "Time since last observation" is often a stronger feature than the raw value.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pandas

Minimal runnable code

irregular_vitals.py
import pandas as pd

# Heart rate readings for one patient. Nurses measure vitals when they
# judge it necessary, not on a fixed clock -- a sicker patient gets
# checked more often.
vitals = pd.DataFrame({
    "time": pd.to_datetime([
        "2026-01-01 08:00", "2026-01-01 08:15", "2026-01-01 08:20",
        "2026-01-01 11:00", "2026-01-01 19:45",
    ]),
    "heart_rate": [78, 92, 101, 88, 84],
})

# Gap since the previous reading, in minutes. A short gap often means a
# nurse got worried. A long gap usually means the patient looked fine.
vitals["minutes_since_last"] = vitals["time"].diff().dt.total_seconds() / 60
print(vitals.to_string(index=False))
print()

# The naive fix: force this onto an hourly grid and forward-fill the gaps.
hourly = (
    vitals.set_index("time")["heart_rate"]
    .resample("1h")
    .mean()
    .ffill()
)
print("forced onto an hourly grid:")
print(hourly.to_string())
Output
               time  heart_rate  minutes_since_last
2026-01-01 08:00:00          78                 NaN
2026-01-01 08:15:00          92                15.0
2026-01-01 08:20:00         101                 5.0
2026-01-01 11:00:00          88               160.0
2026-01-01 19:45:00          84               525.0

forced onto an hourly grid:
time
2026-01-01 08:00:00    90.333333
2026-01-01 09:00:00    90.333333
2026-01-01 10:00:00    90.333333
2026-01-01 11:00:00    88.000000
2026-01-01 12:00:00    88.000000
2026-01-01 13:00:00    88.000000
2026-01-01 14:00:00    88.000000
2026-01-01 15:00:00    88.000000
2026-01-01 16:00:00    88.000000
2026-01-01 17:00:00    88.000000
2026-01-01 18:00:00    88.000000
2026-01-01 19:00:00    84.000000
Freq: h

What actually happened

Look at minutes_since_last: 15, 5, then a jump to 160, then a jump to 525. The nurse checked twice in five minutes around 8:15, then did not return for almost nine hours. That gap is real clinical information — the patient was judged stable enough to leave alone.

Now look at the hourly grid. From 11:00 to 18:00, every single row says 88.000000. Not one of those numbers was measured. .ffill() copied the 11:00 reading across eight invented hours, silently converting "nobody checked" into "the heart rate stayed exactly at 88."

Line by line, the parts that are not obvious:

  • .diff() on a datetime column returns a Timedelta, not a number — .dt.total_seconds() is required to turn it into something you can do arithmetic with.
  • .resample("1h") creates a row for every hour in the range, whether or not any reading fell inside it. It does not know or care that some of those hours had zero real measurements.
  • .ffill() after .resample() is the specific combination that manufactures data. Each on its own is reasonable; together, on data this irregular, they are not.

Common mistakes

Resampling before checking how sparse the data really is. A patient checked twice in twenty years, when resampled to daily bins, produces thousands of forward-filled rows from two real numbers.

Throwing away the timestamp once you have the value. The value alone loses the "how urgent did this look" signal that the checking frequency carried.

Assuming a long gap means "no event happened." Sometimes it means the monitor was disconnected, or the patient was moved to a different unit's system entirely. A missing observation reason is itself worth recording.

Try it yourself

Add a minutes_since_last column to the hourly-resampled series too — for each hour, how long ago was the real measurement it was filled from? That number, not the filled value itself, is usually the safer feature to hand a model.

What to learn next

Researcher — Mathematics and papers.

Formalising the problem

Regular time series methods assume a fixed sampling interval Δt. Clinical time series are better modelled as a marked point process: a sequence of (time, value) pairs

{ (t_1, x_1), (t_2, x_2), ..., (t_n, x_n) },   t_1 < t_2 < ... < t_n

where the arrival times t_i are themselves a random process, not a fixed grid. Critically, the arrival process is not independent of the underlying patient state — this is sometimes called informative observation time or informatively missing data, distinct from the MCAR/MAR/MNAR taxonomy of missing values covered in what hospital data actually looks like, because here the timing itself is the missing-data mechanism.

Encoding irregularity for a model

Three standard representations, in increasing order of how much timing information they preserve:

1. Last-observation-carried-forward + resample to fixed Δt
   -- loses all timing information, the naive approach shown above

2. (value, time-since-last-observation, mask) triplets per feature per step
   -- keeps timing as an explicit input; used by GRU-D

3. Continuous-time models
   -- treat t as a continuous input, no discretisation at all

GRU-D (Che et al., 2018) augments a GRU with a learned decay term. The hidden state decays toward the population mean as the gap since the last real observation grows:

gamma_t = exp( -max(0, W_gamma * delta_t + b_gamma) )
x_hat_t = m_t * x_t + (1 - m_t) * (gamma_t * x_last + (1 - gamma_t) * x_mean)
  • delta_t — time since the feature was last actually observed
  • m_t — a binary mask, 1 if observed at step t, 0 if missing
  • gamma_t — a learned decay factor in (0, 1], computed per feature
  • x_last, x_mean — the last observed value and the population mean for that feature

As delta_t grows, gamma_t shrinks toward 0, and the imputed value drifts toward the population mean rather than staying frozen at the last reading — a principled generalisation of the forward-fill shown in the developer block.

Neural ODEs and latent ODE models (Rubanova, Chen & Duvenaud, 2019) go further, treating the hidden state as the solution to a differential equation that can be evaluated at any real-valued time, removing discretisation entirely.

Cost

For a patient with n irregular observations across d features:

ApproachSequence length seen by the model
Fixed-grid resampling at interval Δt(t_n - t_1) / Δt, independent of n
Event-based (one step per real observation)n

Fixed-grid resampling at fine Δt can make sequence length explode independent of how much data actually exists — a patient measured twice in a week, resampled to five-minute bins, produces thousands of mostly-fabricated steps.

Key references

  • Che, Z. et al. (2018). Recurrent Neural Networks for Multivariate Time Series with Missing Values. Scientific Reports 8. Introduces GRU-D.
  • Rubanova, Y., Chen, R. & Duvenaud, D. (2019). Latent ODEs for Irregularly-Sampled Time Series. NeurIPS.
  • Little, R. & Rubin, D. (2019). Statistical Analysis with Missing Data (3rd ed.). Wiley. The standard reference on missingness mechanisms, extended here to missingness in time.
  • Shukla, S. & Marlin, B. (2021). Multi-Time Attention Networks for Irregularly Sampled Time Series. ICLR.

Current state

GRU-D and its descendants remain common baselines in clinical ML papers. Transformer-based approaches with continuous-time positional encodings (extending the ideas in rotary position embeddings to real-valued, non-integer positions) are an active area, but a well-engineered "time since last observation" feature fed to a gradient-boosted tree, as in the developer block, remains a strong and far cheaper baseline in most published comparisons.

What to learn next