New York: London: Tokyo:

Evaluating AI Tutors for Exam Preparation: A Procurement Checklist for Education Providers

7 / 100 SEO Score

AI tutors promise to make personalised exam support available beyond the time teachers and tutors can provide individually. The commercial momentum behind that proposition is growing: EU-Startups reports that London-based Medly AI has raised a €6.9 million Seed round for its AI-powered exam-preparation platform. That funding is a useful market signal, but it is not evidence that any particular product improves attainment, reduces educator workload, or meets an institution’s safeguarding obligations.

Schools, tutoring businesses, and training providers therefore need to evaluate AI tutors as educational systems rather than attractive software demonstrations. The central procurement question is not whether a tool can produce fluent answers. It is whether it can deliver dependable, curriculum-relevant support to a defined learner group while keeping educators accountable and student data protected.

Begin with the exam, learner group, and intended job

A procurement exercise should start with a narrowly defined use case. “Personalised learning” is too broad to test or contract for. Specify the qualification, subject, exam board where relevant, learner age, expected usage period, and exact task the product will perform.

For example, the intended job might be helping GCSE learners practise algebra independently, giving formative feedback on essay plans, or identifying gaps before a vocational certification assessment. Each creates different accuracy, safeguarding, and oversight requirements.

Ask the supplier to demonstrate:

  • which curricula, qualifications, specifications, and exam-board versions are supported;
  • how quickly content is updated when specifications or assessment criteria change;
  • whether questions reproduce the format, command words, timing, and mark allocation learners will encounter;
  • how the system handles subjects with contested interpretations or multiple valid answers;
  • whether educators can restrict activity to material already taught; and
  • whether performance data distinguishes genuine mastery from repeated prompting or guessing.

Do not accept a generic mapping to a subject or age range as proof of alignment. Review sample tasks against the current specification and marking guidance. Subject teachers or experienced tutors—not procurement staff alone—should conduct this check.

Test correctness, explanations, and uncertainty separately

An AI tutor can state a correct answer while giving a confusing explanation, or provide polished feedback based on a mistaken interpretation. Evaluation should therefore separate answer accuracy, reasoning quality, feedback usefulness, and uncertainty handling.

Create a test set that reflects real learner behaviour, including misspellings, incomplete working, unusual methods, misconceptions, ambiguous questions, and attempts to obtain answers without doing the work. Include both routine material and boundary cases selected by educators. Suppliers should explain whether they have evaluated the product independently, how recent their testing is, and whether results apply to the version being purchased.

Review whether feedback identifies the learner’s specific error, explains the relevant principle, and offers an appropriate next step. In exam preparation, guidance should also remain consistent with the applicable marking criteria. A long response that overwhelms a learner is not necessarily better than a concise hint.

Uncertainty is a procurement requirement, not a minor interface feature. Ask what happens when the system lacks sufficient context, encounters material outside its supported scope, or receives a challenge to a questionable answer. A responsible design should be able to acknowledge uncertainty, avoid fabricating authority, and route the issue to a teacher, tutor, or support team. Buyers should require a visible method for learners and staff to report errors and a documented process for correction.

Keep educators in control of consequential decisions

AI support should not blur responsibility. Before purchase, define which tasks the system may perform autonomously and which remain human decisions. Giving practice hints presents a different risk from assigning predicted grades, recommending examination entry levels, or deciding that a learner no longer needs intervention.

Educators need practical oversight tools: visibility into learner interactions, the ability to inspect feedback, controls over content and assignments, alerts for predefined concerns, and a way to override recommendations. Administrators should be able to audit changes in settings and understand how learner profiles are generated.

Oversight also has a workload cost. If teachers must inspect every conversation, the product may not expand capacity. If no one checks problematic interactions, accountability becomes nominal. Ask the supplier to show how exceptions are prioritised and estimate the time required for routine administration, incident review, and learner support.

For younger users, safeguarding must be designed into the service. Evaluate age assurance, moderation, inappropriate-content controls, reporting routes, staff permissions, and responses to disclosures or signs of distress. Clarify whether the tutor is designed to avoid emotional dependency, manipulative engagement, or representations that could make a child mistake it for a human confidant. The institution’s designated safeguarding personnel should review escalation procedures before learners gain access.

Examine data practices, accessibility, and operational fit

Request a data-flow map showing what is collected, why it is needed, where it is stored, who can access it, how long it is retained, and which subprocessors are involved. Determine whether student prompts, uploaded work, inferred ability levels, and interaction histories are used to train or improve models. The supplier should provide contractual answers rather than relying on broad privacy language.

Buyers should assess lawful processing, parental or learner communications where applicable, deletion and access procedures, security controls, breach notification, international transfers, and what happens to records when the contract ends. Data minimisation matters: a revision tool should not collect additional personal information merely because it could support future product development.

Accessibility testing should involve learners rather than depend solely on a feature list. Check keyboard navigation, screen-reader compatibility, captions or transcripts, colour contrast, reading-level controls, language support, and operation on the devices and connections learners actually use. Personalisation that works only for confident readers with modern hardware can widen the support gap it claims to address.

Operational review should cover single sign-on, account provisioning, roster synchronisation, learning-platform integration, export formats, role-based permissions, service availability, and support response times. Establish who owns implementation and training. A low licence price can be misleading if staff must manually create accounts, reconcile reports, and answer routine access questions.

Compare the full cost with a credible alternative

Pricing should be assessed against the defined use case, not simply divided by the number of eligible learners. Ask whether charges are based on seats, active users, usage, subjects, institutions, or premium features. Model likely costs for peak revision periods, retakes, staff accounts, integrations, implementation, training, support, and renewal.

Then compare the tool with realistic alternatives: additional tutoring hours, teacher-led revision sessions, an established question bank, or targeted support for the learners most likely to benefit. Include staff time needed to supervise the system and respond to incidents.

Request clear terms for minimum commitments, price changes, unused licences, data export, termination, and service withdrawal. Funding news can suggest that a supplier has resources to develop its product, but buyers still need continuity plans. They should know how learner records will be recovered and teaching will continue if a feature changes or the service becomes unavailable.

Run a bounded pilot before wider adoption

Apparent personalisation is not the same as educational impact. A system may generate varied questions and friendly feedback without improving retention, transfer, or exam performance. The safest route is a time-limited pilot with predetermined success and stop criteria.

Select a defined learner group and document its starting point. Use a reasonable comparison, such as a similar cohort receiving the existing form of support, while recognising that a local pilot is not a controlled research study. Decide in advance which measures matter:

  • learner activation, repeat use, and completion rather than registrations alone;
  • progress on curriculum-aligned assessments or comparable exam questions;
  • retention of learning after a suitable interval;
  • educator time spent assigning, monitoring, correcting, and supporting;
  • frequency and severity of inaccurate or unsuitable responses;
  • safeguarding, privacy, accessibility, and technical support incidents; and
  • differences in participation and outcomes between learner groups.

Collect qualitative evidence from learners and educators, but do not substitute satisfaction for learning progress. Review examples of weak feedback and reported errors, not only dashboard averages. The pilot should also test escalation in practice: submit uncertain questions, report a mistake, request deletion of a learner record, and measure how the supplier responds.

Proceed only if the product demonstrates acceptable educational quality, manageable oversight, safe data handling, equitable access, and value relative to alternatives. If evidence is mixed, narrow the use case or require remediation rather than scaling on the strength of engagement or investor interest. AI tutoring can expand access to practice, but the provider remains responsible for deciding where it belongs—and for ensuring that a human can intervene when it fails.

What Target’s Post-Ulta Beauty Strategy Means for Retailers Managing Brand Partnerships

Target’s transition from Ulta Beauty shop-in-shops to its own Target Beauty Studio concept presents a consequential question for retailers: when should a partner-led category experience […]

Panama Canal Surcharges: A Practical Response Plan for Importers

Panama Canal restrictions create more than a freight-rate problem for importers. When vessel draft limits persist and carriers introduce or increase fees, the effects can […]

How Manufacturers Can Manage Supplier Bottlenecks Before They Constrain Production

A supplier problem becomes a production problem when a missing part—not overall purchasing volume—determines whether a finished product can ship. That is why manufacturers need […]

How to Rework Your About Page for Better AI Visibility

Your ecommerce About page is no longer read only by customers deciding whether they trust your store. Search systems and AI-powered answer engines can also […]

A Small-Business Cloud Setup Checklist: From Workload Requirements to Operational Control

A cloud setup is not complete when servers, storage, and applications are online. For a small business, success also depends on whether the environment supports […]

Choosing a Remote-Work Tool Stack Without Creating Software Sprawl

A remote team needs ways to communicate, coordinate tasks, share documents, schedule work, and review performance. The mistake is treating each need as a separate […]

How to Audit and Reduce Business Overhead Without Weakening Operations

Overhead reduction is not simply a hunt for the largest bills. An expense can be indirect and still protect sales, service quality, compliance, or delivery […]

Evaluating AI Tutors for Exam Preparation: A Procurement Checklist for Education Providers

AI tutors promise to make personalised exam support available beyond the time teachers and tutors can provide individually. The commercial momentum behind that proposition is […]

How Power-Intensive Businesses Can Turn Electricity Flexibility Into an Operating Advantage

For a power-intensive facility, electricity is not merely a bill to negotiate once a year. Its cost can vary with consumption timing, peak demand, tariff […]