How Should You Evaluate an AI Tutor? A Practical Scorecard
To evaluate an AI tutor, test whether it improves what learners can do independently while meeting your requirements for curriculum control, human oversight, privacy, safety, accessibility, reliability, and cost. A fluent answer or polished demonstration is not evidence of good tutoring.
Use real learner mistakes and realistic operating conditions. Score what you can verify, record what remains unknown, and run a bounded pilot before making broad outcome claims.
The scorecard below is designed for schools, universities, training providers, employers, and other learning programs. Local law and policy still require qualified review in your jurisdiction.
What is the difference between a chatbot and a tutor?
A chatbot is usually rewarded for producing a helpful response. In education, the most immediately helpful response can be the wrong teaching move: it may complete the task the learner was meant to practise.
A tutor must manage the sequence of help. It should discover what the learner has tried, locate the misconception or missing prerequisite, choose an appropriate explanation or prompt, and return the work to the learner. The interaction should end with evidence—not merely confidence—that the learner can continue.
That is why “Did it give the right answer?” is necessary but insufficient. The evaluation must include the learner's behavior after the response.

How should you score each requirement?
Use the same four-point scale for every question:
| Score | Meaning |
|---|---|
| 0 | Missing, contradicted, or unacceptable |
| 1 | Claimed or demonstrated only in a vendor-controlled example |
| 2 | Documented and reproducible in your test environment |
| 3 | Verified with your learners, content, staff, and operating conditions |
Do not average away a critical failure. Privacy, safeguarding, accessibility, or curriculum-control requirements may be pass/fail gates even when the total score is high.
Which 12 questions should an AI-tutor scorecard ask?
1. Does it begin with the learning job?
Can your team define the learner, intended capability, context, and boundary of help before configuring the system? Reject demonstrations built around vague promises to “personalize learning.”
Test: Give the vendor one bounded outcome from a real program and ask them to show how the tutor is configured for it.
2. Does it elicit the learner's thinking?
Does the tutor ask what the learner has tried or gather an unaided attempt before teaching? A system cannot diagnose reasoning it never sees.
Test: Submit correct, partially correct, confidently wrong, blank, and copied responses. Compare the first move.
3. Does it respond to the actual break?
An explanation can be accurate and still irrelevant. The tutor should distinguish missing knowledge, faulty reasoning, a procedural mistake, and a misread instruction.
Test: Create several errors that lead to the same wrong answer. Check whether the tutor gives the same generic lesson or addresses each cause.
4. Does it scaffold instead of completing?
Look for a deliberate progression: question, cue, hint, partial example, explanation, and—only when appropriate—solution. Your educators should be able to set boundaries for different tasks.
Test: Ask directly for the answer, then pressure the tutor with “just tell me.” Record when and why it gives way.
5. Does it return responsibility to the learner?
A good interaction should require explanation, prediction, comparison, revision, or a fresh independent attempt.
Test: After assistance, give a different but equivalent task with the tutor unavailable. Measure what the learner can do, not how satisfied they felt.
6. Can you control curriculum and sources?
Can staff define outcomes, assistance rules, approved material, terminology, and exclusions? Can a reader distinguish source material from generated explanation?
Test: Include an ambiguous question, outdated source, and material that conflicts with general model knowledge. Inspect the response and citations.
7. Is educator oversight useful rather than performative?
Educators need meaningful information: repeated misconceptions, incomplete work, evidence by outcome, uncertainty, and cases requiring intervention. A transcript archive alone moves the analysis burden back to staff.
Test: Ask an instructor to use the setup and review workflow without vendor help. Measure time, decisions supported, and information still missing.
8. Is the learning evidence interpretable?
Can the organization see what was attempted, what support was given, what the learner did independently, and how the result was calculated? An unexplained “mastery” number is not enough.
Test: Trace one dashboard result back to the underlying learner evidence. Then ask how uncertainty, recency, and conflicting evidence are handled.
For a measurement model, see how to measure whether an AI tutor improves learning.
9. Are privacy and data rights explicit?
Map what is collected, why it is needed, where it goes, who can access it, how long it remains, whether it trains models, and how it is exported or deleted.
The U.S. Department of Education provides a privacy and education-technology resource collection, including a model terms-of-service checklist. Other jurisdictions have different requirements; a global product may need several reviews.
Test: Complete a real deletion and export request in the proposed account configuration. Do not accept a policy description as proof of workflow.
10. Are safety and escalation boundaries clear?
Define what happens with harmful content, sensitive disclosures, repeated failure, suspected misuse, and requests outside the tutor's role. Name the human owner and response time for every escalation.
UNESCO's guidance for generative AI in education emphasizes privacy protection, age-appropriate use, human agency, and ethical validation.
Test: Run agreed adversarial and boundary scenarios. Inspect the event record, user experience, and staff notification—not only the model's words.
11. Can all intended learners use it?
Accessibility includes the complete learning interaction: authentication, input, voice, captions, board content, feedback, navigation, time limits, and alternatives.
The W3C recommends using the current Web Content Accessibility Guidelines as a shared standard, while noting that conformance does not cover every user need.
Test: Include learners and specialists in evaluation. Test keyboard and screen-reader use, captions, contrast, zoom, input alternatives, language, and accommodation workflows.
12. Can it operate reliably at an acceptable cost?
Measure time to first response, end-of-turn latency, interruptions, recovery, completion, support load, staff setup time, and fully loaded variable cost.
Test: Use real network conditions, concurrent sessions, long sessions, provider failure, and reconnects. Ask who owns unfinished work and records when a service fails.
How should you run the evaluation?
Use four stages:
- Document review: policies, architecture, evidence, contracts, data flows, accessibility status, and incident process.
- Scenario testing: educator-written cases across subjects, learner states, risks, and boundary conditions.
- Workflow testing: configuration, enrollment, teaching, review, support, export, and deletion with the actual staff who would own them.
- Bounded pilot: a defined cohort, baseline, primary learning outcome, guardrails, cost ceiling, stop rules, and decision date.
NIST's voluntary AI Risk Management Framework organizes risk work around governance, context mapping, measurement, and management. An education-specific scorecard should connect those disciplines to the learning job rather than treating “AI risk” as a separate compliance page.
What evidence should a vendor provide?
Ask for the method before the headline:
- Which learners, subjects, settings, duration, and product version were tested?
- What comparison or baseline was used?
- Was the outcome measured while help was available or independently afterward?
- Was retention tested later?
- Who selected the measures and analyzed the result?
- What failed, and who dropped out?
- Does the claim apply to your population and implementation?
Evidence from one setting is not universal proof. The What Works Clearinghouse notes that study findings can apply narrowly to the population and setting studied; its current handbooks explain how it reviews evidence quality.
The final question
After all 12 categories, ask:
If the polished AI response disappeared, what new capability or better decision would remain with the learner and the organization?
That is the result an AI tutor should be evaluated to produce.
See how Kuji works with organizations or use this scorecard to structure your next vendor evaluation.

