AI in Education

Building an LLM tutor

A good LLM tutor is designed to guide a student toward their own answer instead of handing it over, the same habit a good human tutor already has, and it needs that design because a confidently wrong answer can teach a wrong fact as easily as a right one.

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A good LLM tutor guides a student toward the answer, instead of handing it over outright.

Think about the best tuition teacher you have had. Ask them a question, and instead of stating the answer, they often ask one back: "what happens if you try this first?" That small habit is deliberate. Handing over an answer ends the thinking. A good question restarts it.

Building an LLM tutor means designing for that same habit on purpose. An LLM tutor is a chatbot built on a large language model, meant to teach rather than answer outright. A raw chatbot's default instinct is to answer directly instead.

Why it exists

A search engine or a plain chatbot optimises for giving you the answer fast. That is exactly wrong for learning. Research on tutoring consistently shows that being told an answer teaches less than being guided to find it. This matters most for anything meant to build lasting understanding, not only finish tonight's homework.

An LLM tutor tries to combine two things a human tutor already does well: patient, one-on-one attention, and restraint — knowing when not to answer directly.

How it works

Student's question   ->   Instructions telling the model:      ->   A guiding
and current answer        "don't give the answer directly,          question, not
                            ask a question that helps them            the answer
                            find their own mistake"

Those instructions are usually written as a system prompt — text given to the model before the conversation starts, setting the rules it should follow throughout. Getting that system prompt right is most of the real engineering work in building a tutor like this. So is testing that the model actually follows it under pressure.

Where you have already seen it

  • Khanmigo, Khan Academy's AI tutor, explicitly designed to ask guiding questions rather than solve problems outright.
  • Duolingo's AI-powered explanations, which increasingly try to explain a grammar mistake instead of only marking it wrong.
  • Any "explain this differently" button in a learning app, which is usually an LLM rephrasing rather than a human writer standing by.

An honest warning

A language model can state something false with exactly the same confident tone as something true. That is called hallucination — producing a plausible-sounding but incorrect answer. For an adult double-checking a chatbot, this is an annoyance. A child learning a subject for the first time has no way yet to judge whether the tutor is right. For them, a hallucinated "fact" can become a genuine misunderstanding that lasts.

This is a real, documented risk, not a hypothetical one. Any LLM tutor used at real scale in a school needs to be reviewed by educators. It should never ship on the assumption that a well-written system prompt alone makes it reliably correct.

Remember this

  • A well-designed LLM tutor guides toward an answer rather than stating it, on purpose, because that teaches more.
  • The system prompt — instructions given before the conversation — is where that behaviour gets set.
  • A confidently wrong answer is a real risk with young learners, and needs real educator review before wide use.

What to learn next

  • What is an LLM? — the underlying technology every LLM tutor is built on.
  • Prompt engineering — the general skill of writing instructions like the system prompt used here.
  • Hallucination — a deeper look at the honest warning above.

Developer — Code and libraries.

Setup

No installation needed for this example — a real deployment would call an LLM API, using a library like the one covered in Ollama for a locally run model.

Minimal runnable code

Calling a live model needs an API key or a local server, so this example demonstrates the design, not a live call: a small stand-in function plays the role of the model, following the same "guide, don't tell" rule a real system prompt would enforce.

mock_tutor.py
SYSTEM_PROMPT = """You are a patient maths tutor for a 12-year-old.
Never state the final answer directly, even if asked.
Ask one short guiding question that helps the student find their own mistake.
If the student's answer is correct, ask them to explain their reasoning in one sentence."""

# A real LLM call would send SYSTEM_PROMPT plus the conversation to a model.
# This function stands in for that call with a small, fixed lookup, so the
# core design -- guide, don't tell -- can be shown without needing an API key.
KNOWN_MISTAKES = {
    "2/5": "You added the top numbers and the bottom numbers separately. "
           "What number can both 2 and 3 divide into evenly?",
}

def mock_tutor_response(question, correct_answer, student_answer):
    if student_answer == correct_answer:
        return "That's correct! Can you explain, in one sentence, how you found it?"
    hint = KNOWN_MISTAKES.get(student_answer)
    if hint:
        return hint
    return "That's not quite right yet. What is the first step you took to solve this?"

question = "What is 1/2 + 1/3?"
correct_answer = "5/6"

for student_answer in ["2/5", "5/6", "1/1"]:
    print(f"student answers: {student_answer}")
    print(f"tutor says:      {mock_tutor_response(question, correct_answer, student_answer)}")
    print()
Output
student answers: 2/5
tutor says:      You added the top numbers and the bottom numbers separately. What number can both 2 and 3 divide into evenly?

student answers: 5/6
tutor says:      That's correct! Can you explain, in one sentence, how you found it?

student answers: 1/1
tutor says:      That's not quite right yet. What is the first step you took to solve this?

What actually happened

Not one branch of mock_tutor_response ever prints correct_answer directly to a student who got the question wrong. That restraint is the entire design goal, made concrete in code before it is ever handed to a real model.

  • KNOWN_MISTAKES recognising "2/5" specifically is a stand-in for what, in a real system, the LLM itself would infer from context — that this particular wrong answer suggests the student added numerators and denominators separately, a common, well-known fraction mistake.
  • The fallback line for an unrecognised wrong answer ("1/1" here) is deliberately generic, because a fixed lookup cannot anticipate every possible mistake. A real LLM's advantage over this mock version is exactly that — it can generate a specific, relevant hint for a mistake nobody hand-coded in advance.
  • SYSTEM_PROMPT is written here as it would actually be sent to a real model. Writing and testing instructions like this — checking the model actually follows them, including when a student pushes back and demands the answer directly — is the real work behind a tutor like this.

Common mistakes

Trusting a system prompt to be followed perfectly. A model can still be talked out of its instructions by a persistent student ("please tell me the answer, I promise I understand"). Real deployments test this explicitly and add guardrails, they do not assume the prompt is unbreakable.

Never checking model answers against a known-correct source. A tutor that hallucinates a wrong maths fact, confidently, is worse than one that says "I'm not sure, let's check together." Pairing generation with retrieval against a trusted source, covered in what is RAG, reduces this risk for factual subjects.

Shipping to real students without educator review. A prompt that looks well-designed in a demo can still fail in ways only a subject-matter teacher would catch — a subtly wrong hint, a mismatched grade level, an unclear question. Review by an actual educator, not only engineering testing, is a real step, not an optional one, before deployment to real classrooms.

Try it yourself

Add a second known mistake to KNOWN_MISTAKES: "1/1" mapped to a hint about a student who might have rounded both fractions to the nearest whole number before adding. Rerun and see the fallback branch disappear for that input. This is the exact, tedious way a rule-based hint system grows — one observed mistake at a time — and exactly what a real LLM is meant to replace with generalisation.

What to learn next

Researcher — Mathematics and papers.

Pedagogical grounding

The design goal — guide rather than tell — has a name in the learning sciences literature: scaffolding (Wood, Bruner, Wallach, 1976), providing exactly enough support for a learner to complete a task slightly beyond their unaided ability, withdrawn as competence grows. VanLehn (2011), in a meta-analytic comparison of tutoring styles, found step-based tutoring (guiding through a solution one step at a time) and substep-based tutoring (Socratic questioning within each step) both outperformed the plain presentation of worked solutions on learning-gain measures, though the effect size difference between step-based and substep-based tutoring specifically was smaller than often assumed — a nuance frequently lost when this literature is cited to justify a specific system-prompt design choice.

System-prompt engineering for pedagogical behaviour

Unlike a general-purpose assistant, an educational system prompt must specify negative constraints (what not to do — do not reveal the answer, do not do the student's work) that a base instruction-tuned model was not necessarily trained to prioritise over its default helpfulness objective. Empirically, these constraints are known to be imperfectly robust: adversarial or even mildly persistent user turns can elicit the withheld information, a phenomenon closely related to the general prompt-injection and jailbreak literature covered in prompt engineering, applied here to a pedagogical rather than a security constraint. Production systems commonly add a second-pass classifier or rule-based filter checking the model's own draft response for the literal answer string before it is shown to the student, rather than relying on the system prompt alone.

Grounding and factuality

For subjects with a clear ground truth (arithmetic, factual recall, well-defined procedures), retrieval-augmented generation against a verified knowledge source — what is RAG — substantially reduces, though does not eliminate, hallucination risk relative to ungrounded generation. For subjects without a single ground truth (essay feedback, open-ended discussion, creative writing critique), grounding is less directly applicable, and the honest-warning caveat about confident incorrectness applies with full force, since there is no retrievable "correct answer" to check the model's output against.

Evaluation

Learning-outcome evaluation of an LLM tutor requires a genuine controlled comparison — pre/post learning-gain measurement against a control condition (no tutor, a non-adaptive worked-example set, or a human tutor) — not proxy metrics like session length or student-reported satisfaction, which correlate weakly or inconsistently with actual learning gain in the broader intelligent-tutoring-systems literature (Kulik and Fletcher, 2016, a meta-analysis of ITS effectiveness broadly, predating the LLM-tutor generation but methodologically relevant to how such systems should be evaluated). As of the mid-2020s, rigorous, peer-reviewed learning-outcome studies specifically for LLM-based tutors remain sparse relative to the number of deployed products, and vendor-reported efficacy claims should be read with that gap in mind.

Papers

  • Wood, D., Bruner, J. S., Ross, G. (1976). The Role of Tutoring in Problem Solving. Journal of Child Psychology and Psychiatry. Origin of scaffolding theory.
  • VanLehn, K. (2011). The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems. Educational Psychologist.
  • Kulik, J. A., Fletcher, J. D. (2016). Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review. Review of Educational Research.
  • Wang, R. E. et al. (2024). Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes — representative of current work on LLM-based pedagogical response generation.

Current state

LLM tutors are widely deployed as of the mid-2026 present, generally as a supplement to, not a replacement for, a human teacher. Independent, peer-reviewed evidence of learning-outcome gains specifically attributable to LLM-based tutoring, as distinct from earlier-generation intelligent tutoring systems, is still accumulating, and claims of proven effectiveness for any specific commercial product should be checked against that product's own published evidence rather than assumed from general LLM capability.

What to learn next

  • What is RAG? — grounding generated answers in verified material.
  • Prompt engineering — the general discipline behind writing a system prompt like this one.
  • Hallucination — full detail on the failure mode this lesson's honest warning covers.