Customer Assurance Current public

AI Quality & Safety Testing Summary

Version 1.0 Effective August 16, 2026 Last updated August 24, 2026

Overview

MHLE applies a structured evaluation program before deploying AI models in production and on an ongoing basis during operation. This page summarizes our approach to AI quality assurance, safety testing, and ongoing monitoring.

Our goal is to give users, educators, and institutional partners confidence that the AI capabilities in MHLE meet a defined standard of accuracy, fairness, and safety appropriate for educational use.


Our AI Evaluation Program

1. Pre-Deployment Evaluation

Before any new AI model or major prompt change is deployed to production, our engineering team conducts:

  • Task-fit assessment — Does the model's capability profile match the intended task (analysis, summarization, factual retrieval, embedding)?
  • Regression testing — Automated comparison against a reference output set to detect quality regressions.
  • Adversarial prompt testing — Attempts to elicit harmful, biased, or factually incorrect outputs from the model under realistic student-use conditions.
  • Bias screening — Evaluation for demographic and disciplinary bias across a representative sample of academic note content.
  • Hallucination rate estimation — Rate of factually incorrect claims in model outputs on a curated factual benchmark.

2. Production Monitoring

Once deployed, models are monitored continuously:

  • Latency and error rate tracking — API failures and response-time degradation trigger automated alerts.
  • User feedback signals — Users can flag individual AI outputs as incorrect or unhelpful. These signals feed into our quality review queue.
  • Weekly verification checks — Where enabled, MHLE automatically re-verifies facts in user knowledge graphs against current sources and surfaces outdated claims.

3. Periodic Re-Evaluation

Each active AI model in production is evaluated at least every 6 months against the metrics in section 4 below. Results are logged in our internal AI Benchmark Results Register.


Models Covered

Provider Model Primary Use
Google gemini-2.0-flash Multi-lens note analysis
Google gemini-flash-latest Multi-lens note analysis (alias)
OpenAI gpt-4o Synthesis document generation, deep analysis
OpenAI gpt-4o-mini Lightweight analysis, classification
OpenAI gpt-3.5-turbo Routine summarization
OpenAI text-embedding-3-small Semantic similarity, knowledge graph
OpenAI whisper-1 Voice note transcription
OpenAI tts-1-hd Text-to-speech output
Anthropic claude-sonnet-4-5 Multi-perspective analysis
Anthropic claude-haiku-4-5 Fast classification and labeling
Anthropic claude-opus-4-5 Complex reasoning tasks
Perplexity sonar Factual verification, web-grounded claims

Key Quality Metrics

Metric Definition Target
Task accuracy Correct outputs on held-out evaluation set ≥ 90%
Hallucination rate Factually incorrect claims per 100 outputs ≤ 5%
Bias score Disparity in output quality across demographic groups ≤ 5% gap
Latency (P95) 95th-percentile response time ≤ 30 s (analysis)
Error rate API failures as share of requests ≤ 0.5%

These targets are aspirational; actual results are logged in our AI Benchmark Results Register and are available to enterprise customers on request.


AI Safety Measures

MHLE applies the following safeguards to all AI-generated content:

  1. Output sanitization — AI-generated HTML is sanitized before rendering to prevent injection attacks.
  2. Content moderation — Outputs are screened for content inappropriate for educational settings.
  3. Citation and source labeling — Where AI-generated content draws on verifiable sources, those sources are surfaced to the user.
  4. No automated high-stakes decisions — MHLE AI outputs are advisory only; no admissions, grading, or student-record decisions are made solely by AI.
  5. Human review queue — Flagged outputs are reviewed by the engineering team within 5 business days.

Known Limitations

Users and administrators should be aware of the following limitations:

  • AI outputs may contain errors. The AI analyses MHLE provides are educational aids, not authoritative sources. Users should verify important facts independently.
  • Training data cutoffs. Each model has a knowledge cutoff date; very recent events may not be reflected accurately.
  • Domain specificity. Models may perform better in some academic domains than others. Highly specialized or niche disciplines may see lower accuracy.
  • Language support. Models perform best with English-language content. Other languages are supported but may show higher error rates.

Transparency Commitments

MHLE is committed to increasing the transparency of our AI evaluations over time. Our planned disclosures include:

  • Publishing aggregated accuracy and bias benchmark results annually.
  • Providing enterprise customers with model-specific evaluation reports on request.
  • Notifying users when a model change materially affects output quality or behavior.

Contact

For questions about our AI testing program or to request detailed evaluation reports as an enterprise customer, contact: security@mhle.com

For concerns about specific AI outputs, use the in-app feedback flag or contact: support@mhle.com

Back to Trust Center