Question Answering

Answering questions about tables

Answering a question about a table often means running a real calculation over rows and columns, not extracting a quoted span of text.

On this page 5
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Answering a question about a table means finding the right row and column, or calculating across several of them.

Think about scanning a train timetable on a board to find one platform number. Your eyes run down the "Train" column, find the right row, then read across to the "Platform" column. You are not reading sentences. You are navigating a grid.

Why it exists

Extractive QA, from earlier lessons, finds a span of text inside a passage. Tables are not passages of connected sentences. They are grids of values, and many table questions need something extractive QA was never built for.

"Which city has the largest population?" needs a comparison across every row, not a single quoted phrase. "What is the total revenue across all four quarters?" needs an actual sum. Neither answer exists as a contiguous span of text anywhere in the table.

Table question answering treats the table as structured data instead — rows, columns, typed values. It runs the right operation: a lookup, a comparison, a sum, a count.

How it works

Table:
  city        population_millions
  Mumbai      20.4
  Delhi       32.9
  Chennai     11.5

Q: "Which city has the largest population?"
   -> not a text span. Requires comparing every row's number.
   -> answer: Delhi (32.9 million)

Q: "How many cities are in the table?"
   -> not a text span either. Requires counting rows.
   -> answer: 3

Where you have already seen it

  • Spreadsheet formulas. A "maximum," "total," or "count" formula is exactly this kind of question, asked in formula form instead of plain English.
  • "Ask a question about your data" features in BI tools like Power BI. They translate a typed question into a query.
  • ChatGPT's Code Interpreter. It answers table questions by writing and running real code over an uploaded spreadsheet.

Remember this

  • Many table questions need a real calculation, not a quoted span of text.
  • Treating a table as rows and columns, not as prose, is what makes those calculations possible.
  • Lookup questions and aggregate questions need different handling under the hood.

What to learn next

Developer — Code and libraries.

This answers lookup and aggregate questions over a small table using pandas, the standard tool for exactly this kind of structured-data question.

Setup

bash
pip install pandas

Answering by querying, not by extracting text

table_qa.py
import pandas as pd

data = {
    "city": ["Mumbai", "Delhi", "Bengaluru", "Chennai", "Kolkata"],
    "population_millions": [20.4, 32.9, 13.2, 11.5, 15.1],
    "state": ["Maharashtra", "Delhi", "Karnataka", "Tamil Nadu", "West Bengal"],
}
df = pd.DataFrame(data)
print(df)
print()

question1 = "Which city has the largest population?"
row = df.loc[df["population_millions"].idxmax()]
print(f"Q: {question1}")
print(f"A: {row['city']} ({row['population_millions']} million)")

question2 = "What is the population of Chennai?"
value = df.loc[df["city"] == "Chennai", "population_millions"].iloc[0]
print(f"Q: {question2}")
print(f"A: {value} million")

question3 = "How many cities are in the table?"
print(f"Q: {question3}")
print(f"A: {len(df)}")
Output
        city  population_millions        state
0     Mumbai                 20.4  Maharashtra
1      Delhi                 32.9        Delhi
2  Bengaluru                 13.2    Karnataka
3    Chennai                 11.5   Tamil Nadu
4    Kolkata                 15.1  West Bengal

Q: Which city has the largest population?
A: Delhi (32.9 million)
Q: What is the population of Chennai?
A: 11.5 million
Q: How many cities are in the table?
A: 5

Line by line

df["population_millions"].idxmax() finds the row index of the largest value in that column — the pandas equivalent of scanning every row and remembering the biggest one seen so far.

df.loc[df["city"] == "Chennai", "population_millions"] is a filter, then a column selection. Read it right to left: filter rows where city equals "Chennai," then pull out the population_millions column from what remains.

len(df) counts rows. This looks almost too simple to mention, and that is the point — a question that sounds like it needs language understanding sometimes needs nothing more than len(), once the question is correctly mapped to the right operation.

None of these three answers involved reading any cell as a sentence. Each was a direct, typed query against structured data — the defining difference from the extractive QA lessons earlier in this section.

Common mistakes

Feeding a table to a text-based QA model as if it were a paragraph. Serialising a table into "city: Mumbai, population: 20.4..." and running extractive QA over it works for pure lookups by luck of matching text, and fails outright on any aggregate question — no span of that serialised text says "Delhi," because the comparison across rows was never performed as a comparison.

Not validating the question maps to a real column. "What is the population of Hyderabad?" against this table should return "not found," not a crash or a silently wrong row — Hyderabad is not in the data at all.

Ignoring type mismatches. A column read from a CSV as text, when it should be numeric, breaks idxmax() and similar operations silently or loudly, depending on the pandas version. Confirm dtypes before running aggregate queries.

Assuming a natural-language question maps cleanly to one pandas operation. Real questions are messier: "Which cities have more than 15 million people?" needs a filter and a list, not a single value. Mapping open-ended natural language reliably to the right sequence of operations is a genuinely hard problem, covered further in the researcher block below.

Try it yourself

Add a question asking for cities with population above 15 million, using df[df["population_millions"] > 15], and print the result.

You should get Mumbai, Delhi, and Kolkata back as a filtered table, not a single value — a different shape of answer from every question above, because this question asks for a list, not a lookup or a count.

What to learn next

Researcher — Mathematics and papers.

Two competing strategies

Semantic parsing to a query language translates a natural-language question into an executable query — SQL, pandas operations, or a custom logical form — then runs that query against the table directly, exactly as the developer block does by hand. This guarantees a mathematically correct answer once the translation is right, and it fails hard, with a wrong or unrunnable query, when the translation is wrong.

End-to-end neural table encoding trains a model to consume the table and question jointly and predict an answer, or an operation plus cell selections, without an explicit intermediate query. Herzig et al. (2020), TAPAS: Weakly Supervised Table Parsing via Pre-training (arXiv:2004.02349), is the reference architecture here: it extends BERT with row and column embeddings so the model has positional awareness of table structure, and predicts both which cells the answer involves and which aggregation operation, if any — count, sum, average, or none — applies to them. Crucially, TAPAS trains with only weak supervision: the training data supplies the final answer, not the reasoning steps or query that produced it, and the model must learn to infer both from that weaker signal alone.

Why weak supervision is the harder engineering problem

Strong supervision — training pairs of (question, exact SQL query) — is expensive to collect at scale, since it requires an annotator to write correct, executable queries for every training example. Weak supervision — training pairs of (question, final answer only) — is far cheaper to collect but underspecifies the reasoning path: many different queries can produce the same final answer by coincidence, and the model must learn to find a generally correct reasoning strategy without ever being shown one directly.

Complexity and scaling limits

Semantic-parsing approaches scale to arbitrarily large tables, since the query executes directly against the full table regardless of size — a SQL SUM over a million rows is the same operation as over ten. End-to-end neural approaches like TAPAS are bounded by the model's input length, since the entire table must fit inside the model's context at once; this makes them impractical for tables beyond a few hundred rows without a preceding row- or column-selection step to shrink the table first.

Key references

  • Herzig, J. et al. (2020). TAPAS: Weakly Supervised Table Parsing via Pre-training. arXiv:2004.02349

Current state and open problems

In production systems, semantic parsing to SQL or pandas — often generated by a general-purpose large language model prompted with the table's schema, rather than a table-specialised architecture like TAPAS — has become the dominant approach for large or arbitrary tables, precisely because of the scaling limit above. TAPAS and similar end-to-end architectures remain relevant for smaller, well-bounded tables and for research on joint table-and-text reasoning. The open problem common to both approaches is reliability: a generated query can be syntactically valid and semantically wrong, executing without error while still answering a different question than the one asked, and detecting that failure mode automatically remains unsolved in general.

What to learn next