AI in Finance and Fraud

Point-in-time financial data

Financial numbers change after they are first published, so a model must only ever see the version that existed on that day, never the corrected one.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Point-in-time data means using only the numbers that existed on a given day. Never the corrected numbers that arrived later.

Think about a school re-evaluation. You get your mark sheet in June. In August, after a re-check, one answer is re-graded and your marks change on the official website. Suppose someone later builds a model. It guesses "how confident did this student feel walking out of the exam hall?" Using the August marks would be cheating. The student never saw the August number. They only had the June one.

Financial data behaves the same way, constantly.

Why it exists

A company reports quarterly earnings on one day. Weeks later, accountants sometimes find an error and quietly issue a restatement — a corrected version of the same number. Government statistics offices do this too: a country's first GDP estimate for a quarter is almost always revised, sometimes twice.

Most financial databases store only the latest version of each number. That is the version anyone doing research today usually wants. It is fine for research. It is dangerous for training a model that must act as if it lived through that day.

Suppose your training data quietly uses the corrected, later version of a number. Your model is being handed information from the future. This is called look-ahead bias — the model looks ahead in time without anyone meaning it to. A model trained this way looks brilliant in testing. Then it loses money the moment it runs on live, not-yet-corrected data.

A close cousin is survivorship bias, equally dangerous in the opposite direction. Many stock databases only include companies still trading today. Companies that went bankrupt or got delisted quietly vanish from the record. A model trained only on survivors learns a falsely rosy picture of the past. It never saw the failures.

How it works

What actually happened, day by day:

  Jan 20  ->  Company reports EPS = 1.10        (this is all anyone knew)
  Apr 22  ->  Company reports EPS = 1.35 for Q2
  Jun 10  ->  Auditors correct the Jan number to 1.05   (restatement)

A point-in-time query for "1 Feb" must return 1.10 — the number
that existed on 1 Feb — never the corrected 1.05, which did not
exist yet.

A point-in-time database stores every version of a number, along with the date it became known. You can ask "what did we know on this day?" instead of only "what do we know now?"

A real example you have seen

Financial news channels sometimes report "GDP growth revised from 6.1% to 5.8%." That correction often comes months after the first announcement. Both numbers are real. A trading or lending model must only use the number that was public on the day it decided. Otherwise the backtest is quietly grading itself with an answer key from the future.

Remember this

  • Financial numbers get corrected after they are first published. Training data must use the version that existed on that day, not the final one.
  • Look-ahead bias is when a model learns from information it could not have known yet. It makes backtests look better than reality.
  • Survivorship bias hides past failures from the data. It is equally dangerous, in the opposite direction.
  • Getting this right makes a backtest more honest, not automatically tradeable — real money still needs a risk team's review first.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install pandas

Minimal runnable code

We join daily prices to the most recent earnings report. One join method leaks the future. One does not.

point_in_time_join.py
import pandas as pd

# A company's earnings are announced on report_date.
earnings = pd.DataFrame({
    "report_date": pd.to_datetime(["2024-01-20", "2024-04-22", "2024-07-19"]),
    "eps_first_reported": [1.10, 1.35, 1.20],
})

# We want to attach "the most recently known EPS" to each trade date.
prices = pd.DataFrame({
    "trade_date": pd.to_datetime(["2024-04-15", "2024-05-01", "2024-08-01"]),
    "price": [100.0, 108.0, 95.0],
})

# WRONG: "nearest" can grab a report that had not been published yet.
wrong = pd.merge_asof(
    prices.sort_values("trade_date"), earnings.sort_values("report_date"),
    left_on="trade_date", right_on="report_date", direction="nearest",
)

# RIGHT: only ever look backward in time from the trade date.
right = pd.merge_asof(
    prices.sort_values("trade_date"), earnings.sort_values("report_date"),
    left_on="trade_date", right_on="report_date", direction="backward",
)

print("WRONG (nearest) - trade date 2024-04-15 borrows a report from the future:")
print(wrong[["trade_date", "report_date", "eps_first_reported"]].to_string(index=False))
print()
print("RIGHT (backward) - only reports that already existed are used:")
print(right[["trade_date", "report_date", "eps_first_reported"]].to_string(index=False))
Output
WRONG (nearest) - trade date 2024-04-15 borrows a report from the future:
trade_date report_date  eps_first_reported
2024-04-15  2024-04-22                1.35
2024-05-01  2024-04-22                1.35
2024-08-01  2024-07-19                1.20

RIGHT (backward) - only reports that already existed are used:
trade_date report_date  eps_first_reported
2024-04-15  2024-01-20                1.10
2024-05-01  2024-04-22                1.35
2024-08-01  2024-07-19                1.20

What actually happened

Look at the first row of each table. The trade date 2024-04-15 is 7 days before the April earnings report and 86 days after the January one.

direction="nearest" measures raw distance in days and picks whichever report is closer in time — even if that report is in the future relative to the trade date. It picked the April 1.35 report, which had not been announced yet on 15 April. That is look-ahead bias, created by one wrong argument.

direction="backward" only ever looks at rows where report_date <= trade_date. It correctly picked the January 1.10 report, the only one that actually existed on 15 April.

  • pd.merge_asof is built for exactly this "as of this date" join, and is the standard tool for point-in-time joins in pandas.
  • Both frames must be sorted by the join column first, or the merge raises an error.
  • direction="backward" is the default, but writing it explicitly here makes the intent visible to the next person reading the code.

Common mistakes

Using a plain pd.merge on the nearest date. A regular merge has no concept of time direction at all — you would have to build the "nearest but not future" logic yourself. merge_asof already has it.

Trusting a vendor dataset without asking "point-in-time or latest-known?" Many cheap financial datasets only store the current, corrected values. Ask explicitly, or assume leakage.

Forgetting reporting lag. Even the first report of a number is not known instantly — there is usually a delay between the period ending and the report being published. Use the report's actual publish date, not the period it describes.

Sorting only one of the two frames. merge_asof requires both sides sorted on the join key; sorting only one still raises an error.

Try it yourself

Add a fourth earnings report on "2024-10-15" with eps_first_reported = 0.95, and a trade date "2024-11-01". Run the backward join again and confirm it picks up the October report, not an earlier one.

What to learn next

Researcher — Mathematics and papers.

Framing the problem formally

Let v(x, t) be the value of a data field x as it would have been reported if queried at time t. A vendor's "latest" table only stores v(x, T) for the current time T. A point-in-time table stores the full function, so a backtest at simulated time t_0 < T can correctly query v(x, t_0).

The bias from ignoring this is not random noise — it is a systematic upward bias on any backtested return, because restatements and revisions are, on average, informative and correlated with the outcome you are trying to predict (a beaten earnings estimate tends to get revised up further, not down).

Survivorship bias, quantified

If a strategy only has access to firms S that survived to the present, and expected return conditional on survival exceeds the unconditional expected return —

E[ r | survived ]  >  E[ r ]

— then any backtest over S overstates true historical performance by construction, independent of the model. This cannot be fixed downstream; it has to be fixed by sourcing point-in-time universe membership (which tickers existed, and were investable, on each historical date), not only point-in-time values.

Cost and practice

Building a proper point-in-time store means keeping every historical vintage of every field, not overwriting on revision — an append-only, bitemporal design: one time axis for "the period the data describes," one for "when we learned it." Storage grows roughly linearly with revision frequency rather than staying flat, which is the real engineering cost of doing this correctly.

Key references

  • Ang, A. (2014). Asset Management: A Systematic Approach to Factor Investing. Oxford University Press — chapter on backtesting pitfalls, including point-in-time data.
  • Bailey, D. H., Borwein, J., Lopez de Prado, M., & Zhu, Q. J. (2014). The Probability of Backtest Overfitting. Journal of Computational Finance — formalizes how look-ahead and survivorship bias inflate backtested Sharpe ratios.
  • Vendor documentation from providers such as Compustat and CRSP on "point-in-time" versus "as-reported" datasets are the standard practical references finance teams cite when sourcing this data.

Current state

Point-in-time datasets are commercially expensive precisely because maintaining full revision history is operationally hard, which is why most public tutorials silently skip this and produce backtests that would not survive contact with a live desk. Treat any published finance-ML result that does not state its data vendor's point-in-time policy with real suspicion.

What to learn next

What to learn next

These follow on from what you just read.

  • AI in Finance and Fraud

    Backtesting without fooling yourself

    A backtest that tries enough parameter combinations will find a "winning" strategy in pure random noise — the fix is testing that winner on data it has never seen.

  • AI in Finance and Fraud

    Credit scoring

    A credit model estimates the probability that a borrower will fail to repay, and it is one of the most heavily regulated uses of machine learning that exists.

  • AI in Finance and Fraud

    Reject inference

    A credit model only ever learns from applicants who were approved, so it never sees what would have happened with the people it would have rejected — and that blind spot has a name.