AI Quality & Safety Testing Summary
Overview
MHLE applies a structured evaluation program before deploying AI models in production and on an ongoing basis during operation. This page summarizes our approach to AI quality assurance, safety testing, and ongoing monitoring.
Our goal is to give users, educators, and institutional partners confidence that the AI capabilities in MHLE meet a defined standard of accuracy, fairness, and safety appropriate for educational use.
Our AI Evaluation Program
1. Pre-Deployment Evaluation
Before any new AI model or major prompt change is deployed to production, our engineering team conducts:
- Task-fit assessment — Does the model's capability profile match the intended task (analysis, summarization, factual retrieval, embedding)?
- Regression testing — Automated comparison against a reference output set to detect quality regressions.
- Adversarial prompt testing — Attempts to elicit harmful, biased, or factually incorrect outputs from the model under realistic student-use conditions.
- Bias screening — Evaluation for demographic and disciplinary bias across a representative sample of academic note content.
- Hallucination rate estimation — Rate of factually incorrect claims in model outputs on a curated factual benchmark.
2. Production Monitoring
Once deployed, models are monitored continuously:
- Latency and error rate tracking — API failures and response-time degradation trigger automated alerts.
- User feedback signals — Users can flag individual AI outputs as incorrect or unhelpful. These signals feed into our quality review queue.
- Weekly verification checks — Where enabled, MHLE automatically re-verifies facts in user knowledge graphs against current sources and surfaces outdated claims.
3. Periodic Re-Evaluation
Each active AI model in production is evaluated at least every 6 months against the metrics in section 4 below. Results are logged in our internal AI Benchmark Results Register.
Models Covered
| Provider | Model | Primary Use |
|---|---|---|
| gemini-2.0-flash | Multi-lens note analysis | |
| gemini-flash-latest | Multi-lens note analysis (alias) | |
| OpenAI | gpt-4o | Synthesis document generation, deep analysis |
| OpenAI | gpt-4o-mini | Lightweight analysis, classification |
| OpenAI | gpt-3.5-turbo | Routine summarization |
| OpenAI | text-embedding-3-small | Semantic similarity, knowledge graph |
| OpenAI | whisper-1 | Voice note transcription |
| OpenAI | tts-1-hd | Text-to-speech output |
| Anthropic | claude-sonnet-4-5 | Multi-perspective analysis |
| Anthropic | claude-haiku-4-5 | Fast classification and labeling |
| Anthropic | claude-opus-4-5 | Complex reasoning tasks |
| Perplexity | sonar | Factual verification, web-grounded claims |
Key Quality Metrics
| Metric | Definition | Target |
|---|---|---|
| Task accuracy | Correct outputs on held-out evaluation set | ≥ 90% |
| Hallucination rate | Factually incorrect claims per 100 outputs | ≤ 5% |
| Bias score | Disparity in output quality across demographic groups | ≤ 5% gap |
| Latency (P95) | 95th-percentile response time | ≤ 30 s (analysis) |
| Error rate | API failures as share of requests | ≤ 0.5% |
These targets are aspirational; actual results are logged in our AI Benchmark Results Register and are available to enterprise customers on request.
AI Safety Measures
MHLE applies the following safeguards to all AI-generated content:
- Output sanitization — AI-generated HTML is sanitized before rendering to prevent injection attacks.
- Content moderation — Outputs are screened for content inappropriate for educational settings.
- Citation and source labeling — Where AI-generated content draws on verifiable sources, those sources are surfaced to the user.
- No automated high-stakes decisions — MHLE AI outputs are advisory only; no admissions, grading, or student-record decisions are made solely by AI.
- Human review queue — Flagged outputs are reviewed by the engineering team within 5 business days.
Known Limitations
Users and administrators should be aware of the following limitations:
- AI outputs may contain errors. The AI analyses MHLE provides are educational aids, not authoritative sources. Users should verify important facts independently.
- Training data cutoffs. Each model has a knowledge cutoff date; very recent events may not be reflected accurately.
- Domain specificity. Models may perform better in some academic domains than others. Highly specialized or niche disciplines may see lower accuracy.
- Language support. Models perform best with English-language content. Other languages are supported but may show higher error rates.
Transparency Commitments
MHLE is committed to increasing the transparency of our AI evaluations over time. Our planned disclosures include:
- Publishing aggregated accuracy and bias benchmark results annually.
- Providing enterprise customers with model-specific evaluation reports on request.
- Notifying users when a model change materially affects output quality or behavior.
Contact
For questions about our AI testing program or to request detailed evaluation reports as an enterprise customer, contact: security@mhle.com
For concerns about specific AI outputs, use the in-app feedback flag or contact: support@mhle.com