A comprehensive guide for educators — grounding every dashboard metric in cognitive science, explaining the research that drives each feature, and showing how to apply the expanded toolset — including the Discussion Wheel, Simulation Calibration, and Middle School Mode — to improve learning outcomes across any classroom context.
The MHLE Instructor Dashboard is a real-time learning analytics platform built around a core conviction: that what a student writes—not what they click—is the truest signal of cognitive engagement. Every metric in the dashboard is derived from the actual content of student notes and the depth of their interactions with AI-powered analytical lenses, not from passive event tracking or completion percentages.
This white paper is written for educators, instructors, department heads, and academic coaches who want to understand not just how to use the dashboard, but why each feature was designed the way it was, and what scientific literature supports the approach. The dashboard draws on five established bodies of educational research: Bloom's Taxonomy (1956, revised 2001), the SOLO Taxonomy (Biggs & Collis, 1982), Sweller's Cognitive Load Theory (1988), Zimmerman's Self-Regulated Learning framework (2000) — markers for which are now surfaced directly in the Note Rigor Scorer — and the Socratic dialogue tradition that underpins the Discussion Wheel's comprehension scoring.
The result is a toolset that allows instructors to move from gut-feel grading to evidence-based intervention — identifying struggling students before they fall behind, surfacing conceptual gaps across an entire cohort, and generating AI-powered study coaching personalized to each learner's actual knowledge state. Version 2.0 of this document adds documentation for three expanded capabilities: the Discussion Wheel, which visualizes group discussion as a directed comprehension graph; Simulation Calibration & Preview, which lets instructors verify wicked-problem difficulty before committing it to a course; and Middle School Mode, a dedicated engagement model for MS classrooms with age-appropriate scoring and AI behavior.
MHLE surfaces the quality of student thinking, not just the quantity of activity. A student who has logged in 40 times but only skimmed content will score lower than a student who engaged deeply three times — because the platform reads what was written, not just what was visited.
The design of every MHLE metric is grounded in peer-reviewed educational science. This section explains each theoretical framework and maps it directly to the dashboard features it informs.
Bloom's Taxonomy of Educational Objectives (Bloom et al., 1956; revised by Anderson & Krathwohl, 2001) classifies cognitive tasks into six levels of increasing complexity. The revised taxonomy is hierarchical: lower-order thinking (Remember, Understand) is prerequisite to, but insufficient for, higher-order thinking (Analyze, Evaluate, Create).
MHLE's Note Rigor Scorer automatically classifies each student note into one of these six levels using action-verb detection and linguistic pattern matching — with no AI calls, no latency, no cost. The instructor's Deep-Dive panel then surfaces the student's typical Bloom's level alongside a trend line, letting instructors diagnose at a glance whether a student is stuck in rote recall or operating at analysis and above.
Meta-analyses by Hattie (2009) in Visible Learning identify higher-order questioning and feedback on cognitive level as having effect sizes above d=0.60 — among the highest instructional strategies studied. MHLE surfaces this signal automatically so instructors do not need to read every student note to assess depth.
The Structure of Observed Learning Outcome (SOLO) taxonomy (Biggs & Collis, 1982) is a model of learning quality that measures structural complexity — how well-connected and integrated a learner's understanding of a concept is. Unlike Bloom's, which describes the type of cognitive task, SOLO describes the quality of the response produced.
MHLE uses SOLO depth scores (0.0–1.0) as the gating condition for whether a concept is considered fully covered by a student. A concept the student mentions repeatedly at a Unistructural level will appear as Partial (Shallow) in the Coverage Grid — not Full — because depth of understanding matters more than frequency of mention.
The threshold for Full coverage is set at Relational (≥ 0.65): the student must demonstrate that they understand how the concept connects to other ideas, not merely that they can name it. This design decision reflects the educational consensus that surface-level familiarity does not constitute mastery.
Biggs & Tang (2011) in Teaching for Quality Learning at University demonstrate that students assessed only at Unistructural and Multistructural levels consistently underperform on transfer tasks — real-world application of knowledge — compared to students who are pushed toward Relational understanding. MHLE's SOLO gating forces both students and instructors to recognize when coverage is shallow.
Sweller's Cognitive Load Theory (1988, 1994) proposes that working memory has limited capacity. Instructional designs that overload this capacity — through extraneous information, unclear presentation, or poorly sequenced content — impair learning. Three types of cognitive load are distinguished:
The inherent complexity of the subject matter itself. Cannot be reduced without changing the content.
Load imposed by poor instructional design — confusing interfaces, unclear feedback, fragmented information.
Cognitive effort that contributes to schema formation and long-term learning. The load we want to maximize.
MHLE's split-screen layout, single-concept note prompts, and incremental AI lens delivery are designed to minimize extraneous load while maximizing germane load through structured reflection.
The Friction Alerts feature (see Section 3.10) is a direct application of CLT: when the platform detects that a student is looping on the same page, abandoning notes, or bouncing off feature pages within 30 seconds, it infers that the student's working memory may be overwhelmed or that they are experiencing extraneous load from a UI element they cannot navigate. This triggers an alert for instructor intervention.
Zimmerman's model of Self-Regulated Learning (2000) describes expert learners as people who actively plan, monitor, and evaluate their own cognitive processes. Key SRL behaviors include goal-setting, self-monitoring, strategy use, and self-reflection. Research consistently shows that SRL skills are stronger predictors of academic success than baseline ability measures.
MHLE detects SRL behaviors directly within note text. The Note Rigor Scorer's Self-Regulation dimension searches for:
| Marker Type | Example Phrases | SRL Behavior | Weight |
|---|---|---|---|
| Review Flags | "need to review," "unclear," "come back to," "check this" | Self-monitoring, metacognitive awareness of gaps | 1.5 pts |
| Summary Signals | "in summary," "key takeaway," "main point," "bottom line" | Active consolidation, schema formation | 2.0 pts |
| Self-Test Markers | "practice question," "quiz," "exam question," "check understanding" | Retrieval practice, formative self-assessment | 2.5 pts |
| Elaborations | "this reminds me," "connects to," "this is like," "analogous to" | Prior-knowledge activation, schema extension | 2.0 pts |
| Questions | Any explicit "?" character in the note | Elaborative interrogation, deep processing | 0.5 pts |
Dunlosky et al. (2013) in Psychological Science in the Public Interest rated self-testing (retrieval practice) and distributed practice as the two highest-utility learning techniques — far above highlighting and re-reading. MHLE gives self-test markers the highest weight (2.5 pts) in the Self-Regulation score specifically because of this evidence base.
Learning Analytics (Siemens & Long, 2011) is the measurement, collection, analysis, and reporting of data about learners with the purpose of improving learning and the environments in which it occurs. The distinction between summative assessment (measuring learning at the end) and formative assessment (measuring learning during the process to inform next actions) is fundamental to MHLE's design.
Every metric in the Instructor Dashboard is formative, not summative. The Student Health Score, Note Rigor grades, and Gap alerts are designed to trigger instructor action before a student fails — not to confirm failure after the fact. Black & Wiliam's landmark 1998 meta-analysis (Assessment and Classroom Learning) showed that high-quality formative assessment can produce effect sizes of d=0.4 to d=0.7 — comparable to one-on-one tutoring — when used to guide instructional decisions.
Use the Health Score and Rigor Score as conversation starters, not grades. Share the Note Rigor breakdown with students directly — research shows that learners who understand how they are being assessed adjust their strategies accordingly (Sadler, 1989).
The Lens System — MHLE's core AI analysis engine — is grounded in the epistemological conviction that expert thinking across all disciplines requires the ability to examine a problem through multiple analytical frames simultaneously. King & Kitchener's model of Reflective Judgment (1994) describes mature epistemic cognition as the capacity to hold multiple perspectives in productive tension, rather than collapsing to a single authoritative answer.
MHLE's four core lenses translate this into a practical classroom tool:
Examines resource allocation, incentive structures, cost-benefit tradeoffs, and economic consequences of decisions. Trains students to ask: "Who pays? Who profits? What are the tradeoffs?"
Surfaces stakeholder harm, fairness concerns, rights-based reasoning, and systemic justice implications. Trains students to ask: "Who is affected and how? Is this fair?"
Focuses on technical feasibility, systems thinking, failure modes, and implementation constraints. Trains students to ask: "How does this work? What can go wrong?"
Challenges assumptions, examines evidence quality, identifies cognitive biases and counter-evidence. Trains students to ask: "How do we know this? What are we missing?"
The Gap Hunter feature uses lens coverage as a signal of intellectual breadth. A student whose notes consistently receive Financial and Engineering analysis but never trigger Ethical or Skeptic analysis is showing an epistemic blind spot — and the dashboard makes this visible to the instructor without requiring manual review.
"The capacity to think critically… requires that students be willing to engage in reasoned discourse, carefully examine their own assumptions, and consider multiple perspectives."— Paul & Elder, The Miniature Guide to Critical Thinking, 2008
The Instructor Dashboard begins with your Classroom List — a centralized view of every course section you manage. Each classroom card surfaces the live enrollment count, the average class health score, and the number of open alerts requiring your attention.
From the dashboard home, click Create Classroom. You will be prompted to name the classroom, optionally associate it with an organization (for institutional accounts), and define the course subject and learning level. Once created, an invite link is generated that you can paste into your LMS, email, or syllabus. Students who click the link are automatically enrolled upon login.
The Roster view shows every enrolled student with their current Health Score, note count, AI analysis count, and last login recency. You can remove students, export the full roster as a CSV file, or drill down into any individual student's full profile. Enrollment is self-service by default — students join via the link — but instructors on institutional plans can also manually add students by email address.
Share the invite link during your first class session. Students who enroll in the first week tend to build habits earlier and show consistently higher Health Scores by Week 4. Early enrollment also gives the platform time to accumulate baseline note data before you need to act on health alerts.
The Student Health Score is a composite 0–100 index that measures the quality and consistency of a student's learning engagement. It is computed nightly and updated each morning, giving instructors a daily read on every student without requiring any manual review.
The score is the sum of four equally-weighted components, each designed to capture a distinct dimension of engaged learning:
| Component | How It's Calculated | Saturates At | Research Basis |
|---|---|---|---|
| Note Volume | Logarithmic scale: 20 × log₂(n+1) / log₂(11) |
10+ notes | Quantity of retrieval attempts correlates with retention (Roediger & Karpicke, 2006) |
| AI Analysis | Logarithmic scale: 30 × log₂(n+1) / log₂(19) |
18+ analyses | Elaborative interrogation and multi-perspective analysis deepen encoding (Pressley et al., 1992) |
| Recency | 25 pts today; linear decay to 0 over 15 days | Active today | Spaced practice requires consistent return (Cepeda et al., 2006) |
| Lens Coverage | Green=25, Yellow=13, Red=0 based on uncovered notes | 0 uncovered notes | Multi-perspective analysis correlates with transfer (Hmelo-Silver, 2004) |
A critical design choice: if a student has zero notes AND zero AI analyses, their Health Score is automatically set to 0 regardless of login recency. This prevents ghost-presence: a student who logs in daily but never engages with content is not "healthy" academically.
| Score Range | Status | Suggested Action |
|---|---|---|
| 75 – 100 | Strong Engagement | Acknowledge progress. Consider peer mentoring opportunities. |
| 55 – 74 | Moderate Engagement | Check lens coverage. Prompt to run AI analysis on recent notes. |
| 35 – 54 | At Risk | Review deep-dive modal. Consider proactive outreach or coaching referral. |
| 0 – 34 | Critical | Alert has likely already been generated. Direct intervention recommended. |
The Gap Hunter is MHLE's multi-perspective coverage tracker. For each student, it monitors how many of their notes have been processed through each analytical lens (Financial, Ethical, Engineering, Skeptic). A note that has been analyzed through all four lenses is considered fully covered from an epistemic standpoint. A note with no analysis is uncovered.
The Gap Score component of the Health Score translates this into a simple three-tier signal:
All notes have been analyzed through at least one lens. No uncovered notes. Student is engaging with multi-perspective analysis consistently. Gap Score: 25 pts.
1–2 notes have not yet been analyzed. Student is engaging but some notes are sitting unanalyzed. Gap Score: 13 pts.
3 or more notes have no lens analysis. The student may be writing notes but not engaging with deep AI analysis. Gap Score: 0 pts.
No notes have been created at all. This is distinct from Red — the student has not yet begun engaging with the platform content. Gap Score: 0 pts.
The Cohort Analytics view (accessible from each classroom) aggregates Gap data across all students, surfacing your "Top Gaps" — the specific course topics with the lowest cumulative lens coverage across the entire class. This is invaluable for identifying concepts that the whole class has touched but not deeply examined, pointing toward lecture content that may need richer treatment or supplemental activity.
The Note Rigor Scorer is a deterministic, zero-AI, instant-response text analysis engine. It reads the raw content of each student note and scores it across seven dimensions, producing an overall letter grade (A–F) and an Exam Readiness score (0–100). Because it uses no AI, it runs in milliseconds on every note save — providing immediate feedback without latency or API cost. Scores are stored in note.extra_data['rigor_scores'] and surfaced in the instructor deep-dive modal under the dedicated "Note Rigor" tab.
Formative feedback is only valuable if it is fast enough to influence the behavior it is meant to shape. A feedback loop that takes 30 seconds to respond will be ignored by students mid-session. The Rigor Scorer's deterministic, regex-based architecture ensures sub-100ms scoring, making it viable as a live writing coach, not just a retrospective assessment tool.
Measures the percentage of instructor-provided syllabus topics that appear in the student's note. If you have not yet defined a syllabus in MHLE, the system automatically extracts candidate topics from headers, bold text, and capitalized noun phrases within the note itself and measures self-coverage. Coverage is presented as a percentage alongside a scrollable list of "extracted topics" — the exact concepts the scorer identified.
Counts the use of causal, comparative, and application marker phrases. A note that says "Supply increases because demand rises, unlike the monopoly scenario where price is fixed" is demonstrating richer conceptual linking than one that says "Supply rises. Demand rises. Monopoly price is fixed." The scorer detects three categories of linking:
| Link Type | Example Phrases | Learning Behavior |
|---|---|---|
| Causal | "because," "therefore," "leads to," "results in," "causes" | Systems thinking, cause-effect reasoning |
| Comparative | "unlike," "whereas," "in contrast," "however," "similarly" | Analogical reasoning, schema differentiation |
| Applicative | "for example," "such as," "in practice," "applied to" | Transfer, concrete-abstract bridging |
A composite of four sub-metrics that together constitute a "Writing Quality" letter grade (A–F). The readability target is calibrated to the student's persona: Middle School notes are scored against a Flesch-Kincaid Grade 5–8 target, while higher-education notes target Grade 8–18.
| Sub-Metric | Measures | Why It Matters |
|---|---|---|
| Spell Error Rate | Errors per 100 words | Accuracy under cognitive load signals engagement level |
| Readability | Approx. Flesch-Kincaid Grade Level | Sentence complexity appropriate to the learner's level |
| Vocabulary Range | Type-token ratio (unique/total words) | Lexical diversity correlates with domain knowledge depth |
| Idea Density | Content words vs. function words ratio | High idea density signals substantive rather than filler writing |
Classifies the organizational format of the note into one of five structural types, from least to most sophisticated:
| Structure Type | Detection Signal | Learning Implication |
|---|---|---|
| Prose | Flowing paragraphs; no special formatting | Baseline narrative; may lack retrieval affordances |
| List | Bullet or numbered lines detected | Enumerative thinking; better scan-ability |
| Hierarchical | Nested indented lists or markdown headers | Categorical thinking; explicit parent-child relationships |
| Cornell | Cue/Q: labels, "Summary:" or "Key question" headers | Active recall structure; highest retrieval value |
| Relational | Arrow notation (→, ←, ↔, ==>) linking concepts | Explicit causal/directional mapping; systems thinking |
Classifies the note into one of Bloom's six cognitive levels (Remember → Create) by searching for action verbs and linguistic markers. The scorer assigns a dominant level based on a weighted match count — higher-order verbs (Evaluate, Create) receive more weight than lower-order ones (Remember, Understand) to ensure that a single instance of synthesis language lifts the score appropriately. The scorer returns both the Bloom's label and a 1–6 numeric depth value used in the Exam Readiness calculation.
Detects metacognitive language as described in Section 2.4. The scorer scans for five signal types that indicate active, self-directed learning rather than passive transcription:
| Marker Type | Example Phrases | SRL Behavior | Weight |
|---|---|---|---|
| Review Flags | "need to review," "unclear," "come back to," "check this" | Self-monitoring; metacognitive gap awareness | 1.5 pts |
| Summary Signals | "in summary," "key takeaway," "main point," "bottom line" | Active consolidation; schema formation | 2.0 pts |
| Self-Test Markers | "quiz," "practice question," "exam question," "check understanding" | Retrieval practice; formative self-assessment | 2.5 pts |
| Elaborations | "this reminds me," "connects to," "this is like," "analogous to" | Prior-knowledge activation; schema extension | 2.0 pts |
| Questions | Any explicit "?" character in note text | Elaborative interrogation; deep processing | 0.5 pts |
The total Self-Regulation score is capped at 10. A score of 0 indicates passive information recording; a score above 4 indicates active learning behaviors. This is the dimension most commonly underdeveloped in students new to the platform.
A lightweight, optional check that runs only when instructor notes (PROFESSOR notes) are available for the same course. The scorer compares the student's note content against the instructor's authoritative material and flags potential factual contradictions or misapplied concepts. This check makes zero AI calls — it uses pattern matching against the instructor's own phrasing. When no instructor notes are available, this dimension defaults to a neutral pass and does not penalize the student. The flag surfaces in the Note Rigor tab as an informational indicator for the instructor's review, not as a grade deduction.
The Exam Readiness score is a binary-check composite designed to answer a single question: Is this note likely to support successful exam performance? It awards points for meeting six concrete thresholds:
| Check | Threshold | Points |
|---|---|---|
| Cognitive depth at Analyze or above | Bloom's L4+ | 25 pts |
| Conceptual linking score | ≥ 4.0 | 25 pts |
| Topic coverage | ≥ 40% of syllabus | 20 pts |
| Note length | ≥ 120 words | 15 pts |
| Self-regulation score | ≥ 2.0 | 10 pts |
| Spelling accuracy | < 8% error rate | 5 pts |
Use the Exam Readiness score in pre-exam study sessions. Walk students through their lowest-scoring checks. A student who scores 40/100 on Exam Readiness with the deduction at "Cognitive depth" needs a very different coaching conversation than one who deducts at "Topic coverage" — the first needs deeper analysis prompts, the second needs a curriculum map.
The Coverage Grid is the study-group and cohort-level view of knowledge distribution. It presents as a matrix where each row is a course concept and each column is a student (or group member). Each cell is colored to show coverage depth, giving instructors a comprehensive "knowledge heat map" of the entire classroom at a glance.
| Cell Color | SOLO Level | Meaning | Action Required |
|---|---|---|---|
| ● Emerald Full | Relational+ (≥ 0.65) | Student owns this concept at integration depth | None — mastery achieved |
| ● Amber Shallow | Unistructural/Multi (0.35–0.64) | Concept present but not integrated; surface knowledge only | Probe with Socratic questions or lens re-analysis |
| ● Red None | Missing | No notes include this concept | Direct instructional intervention or note prompt |
Each cell also displays a SOLO Level badge (Extended Abstract, Relational, Multistructural, Unistructural) and a depth percentage pulled from the AI-scored concept records. Hovering over a cell shows the full tooltip including depth score, notes referencing this concept, and the student's self-regulation markers around it. Clicking a cell opens a detail panel with the specific note excerpts that generated the score.
The Coverage Grid uses two concept extraction methods simultaneously. Title-based extraction identifies concepts from note titles and structured headings — fast, deterministic, always available. Semantic extraction (powered by GPT-4o-mini) identifies latent concepts embedded in the body of notes even when they are not explicitly named in headers. This dual approach ensures coverage of both explicit and implicit knowledge.
The Student Deep-Dive Modal is the full-fidelity view of a single student's learning profile. Accessible by clicking any student row in the Roster, it aggregates every signal the platform has collected into a unified diagnostic interface organized into eight tabs.
Health score history chart (30-day sparkline), activity breakdown by day of week, lens usage heatmap showing which analyses the student consistently runs vs. avoids, and recent AI analysis summaries.
Interactive visualization of the student's concept adjacency graph — showing which topics they have connected to each other and the strength of those connections. Isolated concepts (no edges) signal siloed memorization.
Full history of the student's completed AI analyses — each lens result, its timestamp, and the source note. Lets the instructor trace exactly which perspectives the student has applied to their notes and identify blind spots in their analytical lens usage.
AI-generated, personalized study recommendations based on comparing the student's extracted knowledge concepts against the course syllabus — a gap-to-action bridge. Also surfaces the Study Coaching on-demand generation button (see Section 3.11).
The full rigor breakdown: overall letter grade, Exam Readiness bar with check badges, per-dimension cards for all seven scorer dimensions (Completeness, Linking, Writing Quality, Structure, Bloom's Level, Self-Regulation, Accuracy Flag), and a per-note trend list showing rigor trajectory over time.
Per-student simulation outcomes: word count per submission, days between generation and submission, and completion counts by education level. The primary source for Simulation Calibration data at the individual level (see Section 3.8).
AI-generated study coaching — personalized recommendations for each student's next learning actions, derived from their gap profile and note patterns. Separate from the Gap Report; this tab focuses on next-action framing rather than coverage visualization.
Per-student comprehension score (0–100), comprehension level (High / Medium / Low), dominant discussion archetype, and the student's contribution pattern across all active group discussions — all derived from the Conversation Wheel service. See Section 3.7 for full details.
The Knowledge Graph adjacency list shows which concepts a student has linked to each other through their note-taking. A graph with many connections and clusters indicates relational understanding. A graph with many isolated nodes — concepts floating alone with no connections — signals that the student is processing content in silos rather than as an integrated system. In educational neuroscience terms (Bransford et al., 2000), isolated concepts are harder to retrieve under exam pressure because they lack associative retrieval cues.
When you see a student with a dense graph around Financial concepts but sparse connections around Ethical concepts, this is a clear signal about their analytical comfort zone — and the precise point to direct AI lens analysis or in-class discussion prompts.
The Discussion Wheel is a classroom-level and per-student visualization of how group discussion flows during instructor-prompted study group debates. It transforms raw chat activity into a directed comprehension graph — turning conversation into measurable evidence of understanding.
When an instructor activates a discussion prompt in a study group, the Conversation Wheel service begins tracking all subsequent messages. Each student appears as a node on the wheel. Directed edges represent conversational replies: a line from Student A to Student B means A responded to B within a 5-minute window. Edge thickness scales with reply frequency, making visible who is driving the discussion, who is responding, and who is isolated. Nodes are colored by each student's dominant discussion archetype.
Every message posted after a prompt is classified into one of five archetypes using the platform's Contribution Classifier:
| Archetype | Color | What It Signals |
|---|---|---|
| Perspective | Purple | Student offers a point of view or interpretive frame |
| Evidence | Green | Student cites data, examples, or sources to support a claim |
| Synthesis | Blue | Student integrates multiple threads into a unified insight |
| Challenge | Red | Student critiques or problematizes another contribution |
| Question | Amber | Student asks a clarifying or probing question |
Each student receives a Comprehension Score derived from a weighted combination of four signals:
| Signal | Weight | How It's Measured |
|---|---|---|
| Archetype Diversity | 30 pts | Shannon entropy of the student's archetype distribution — students who contribute multiple types of thought score higher than those locked in a single mode |
| Novelty Average | 25 pts | Average novelty score of the student's contributions — how distinct their ideas are relative to prior messages in the thread |
| Synthesis/Evidence Ratio | 25 pts | Ratio of high-value (synthesis + evidence) to passive (question-only) contributions — penalizes students who only ask without answering |
| No Misconception Bonus | 20 pts | Full 20 pts if no factual misconception is detected in the student's messages; 0 pts if a misconception is flagged by a lightweight GPT-4o-mini check |
Comprehension levels: High (≥ 70), Medium (40–69), Low (< 40). When OpenAI is unavailable, the misconception detection step falls back to misconception=False — meaning the student receives the full 20-pt no-misconception bonus and the overall comprehension score is computed normally from the remaining three signals.
When a discussion stalls — defined as a period of low activity after the prompt was activated — the platform generates a Socratic nudge: a single, targeted question that re-engages the group without repeating the original prompt verbatim. Nudges are generated by GPT-4o-mini grounded in the discussion so far, with a deterministic fallback when OpenAI is unavailable. The nudge surfaces to the instructor as a suggested message they can post with one click.
Review comprehension scores at the end of a discussion session. A student with a high Note Rigor score but a Low comprehension score is demonstrating an important gap: they can produce polished written output but struggle to engage dynamically with ideas in real-time — a common pattern in students who have strong memorization habits but underdeveloped analytical discourse skills. This pairing is more diagnostically rich than either signal alone.
The Simulation Engine generates "Wicked Problem" scenarios — open-ended, morally and intellectually complex situations that have no single correct answer. The difficulty and cognitive register of each scenario is calibrated to the student's education level, ranging from Middle School Curious Explorer (Grades 5–6) through PhD / Doctoral. Instructors have two tools for understanding and controlling this calibration.
From the instructor dashboard, per-classroom and per-student simulation stats are available. These include:
| Stat | What It Shows | How to Use It |
|---|---|---|
| Completed Simulation Count | Number of full simulation cycles (generate → submit solution → receive critique) completed per education level | Identify students who are generating scenarios but not submitting solutions — the critique step is where the deepest learning occurs |
| Prior-Period Comparison | Change in completion count vs. the prior period | Spot engagement trends — a student with declining simulation completions may be experiencing difficulty or disengagement with the wicked problem format |
| Education Level Distribution | Breakdown of completed simulations by calibration level | Verify that the class's simulation difficulty is well-matched to their current performance. If most students are completing at an undergraduate level but your course targets graduate-level reasoning, consider adjusting the classroom wicked problem level override |
Instructors and global admins can invoke a dry-run preview of a Wicked Problem scenario at any education level before committing it to the course. The preview endpoint generates a full sample scenario using the course's existing syllabus and notes — without persisting any simulation record to the database and without consuming a student's quota.
To use the preview: navigate to the course's Simulation Settings, select the target education level and optional learning objectives, and click Preview Scenario. The generated scenario appears alongside its calibration label, which describes the role, reading level, stakeholder count, and critique persona for that level. This allows instructors to sanity-check that the scenario complexity is appropriate before students encounter it.
Run a preview at your target education level in Week 1, before students generate their first simulation. Confirm the scenario feels appropriately challenging — not so easy it feels trivial, not so complex it triggers immediate disengagement. If the preview feels off, adjust the classroom's Wicked Problem Level override rather than changing the course structure. The calibration change takes effect immediately for all subsequent generations.
When a classroom is tagged as a Middle School classroom (via the grade_level field), the instructor dashboard switches to a purpose-built engagement model designed for early and late middle school learners (Grades 5–8). This mode changes both the health scoring formula and the AI's behavior for students in that classroom.
Standard health scoring (100 pts across 4 components) is replaced by a 5-signal model calibrated for middle school engagement patterns:
| Signal | How It's Computed | Pedagogical Rationale |
|---|---|---|
| Login Recency (25 pts) | 25 pts if active today; 20 pts within 3 days; 15 pts within 7 days; 8 pts within 14 days; 2 pts otherwise | Habit consistency is the primary predictor of MS learning outcomes — more so than single-session depth |
| Note Count (20 pts) | 2 pts per note, capped at 20 pts (10 notes) | Volume of written output is a stronger baseline signal for MS learners than AI analysis count |
| Micro-Synthesis Rate (20 pts) | Proportion of notes with a completed 3-bullet "What did I just learn?" micro-summary, scaled to 20 pts | Micro-synthesis directly trains metacognitive consolidation; age-appropriate alternative to full lens analysis |
| Curiosity Proxy (20 pts) | Net new unique topics discovered in the recent 14-day window vs. the prior 14-day window, scaled to 20 pts | Topic breadth growth measures intellectual exploration; rewards students who are actively discovering rather than passively consuming |
| Brain Break Engagement (15 pts) | 15 pts if Brain Break check-in recorded within 24 hrs; 8 pts within 7 days; 3 pts otherwise | Emotional self-awareness check-ins predict sustained engagement in MS populations |
A dismissible Brain Break modal appears for MS students on whichever event comes first: page load, first analysis, or first note open — once per session. Students select one of three states:
High-energy state. AI analysis for this session adds an extra layer of nuance and open-ended questions to challenge the student at the edge of their current understanding.
Neutral baseline. AI analysis proceeds at standard depth for the student's persona.
Low-energy state. AI analysis for this session applies a hard ~45-word-per-lens cap and simpler vocabulary to reduce cognitive load. The platform reduces, but does not eliminate, analytical depth.
For all Middle School personas (ms_early and ms_late), Bloom's Taxonomy is explicitly capped at the Application level (L3) in the AI analysis system prompt. The platform does not generate Analyze, Evaluate, or Create-level prompts for MS students — these cognitive demands are developmentally premature and increase cognitive overload rather than productive challenge. All MS AI feedback is also routed through a growth mindset reframing layer that frames verdicts around effort and process rather than ability.
After every AI analysis for an MS student, a background process generates a 3-bullet "What did I just learn?" micro-summary alongside an editable "explain it to a friend" summary. These are stored on the note record and surfaced in the dashboard's MS roster view. Instructors can see the micro-summary completion rate per student as part of the 5-signal health breakdown, making it easy to identify which students are consolidating their learning and which are analyzing without reflecting.
The Brain Break state is stored at the user level and becomes stale after 12 hours (BRAIN_STATE_STALE_AFTER = timedelta(hours=12)). The instructor dashboard surfaces each student's current brain_state alongside their last check-in timestamp in the MS roster view. A student who checked in as "Exhausted" late in the evening may well be in a Fired Up or Meh state the next morning — the timestamp is essential context when interpreting the reported state.
Before enabling Middle School Mode, consider sending parents and guardians the
MHLE Parent & Guardian Guide for Middle School Mode.
It explains, in plain language, what the Brain Break check-in is, why Bloom's Taxonomy is capped at the Application level, what data is stored and for how long, and how to exercise COPPA/FERPA rights — all without requiring families to read this whitepaper. Direct link: /whitepaper/ms-parent-guide
(also reachable at /parents/middle-school).
MHLE generates two categories of alerts for instructors, derived from distinct behavioral signals:
| Alert Type | Trigger | Severity | Recommended Response |
|---|---|---|---|
| HEALTH_DROP | Score falls >10 points over 7 days | Critical | Review the deep-dive modal for the drop cause. Was it recency decay, loss of AI analysis activity, or increased uncovered notes? Each requires a different intervention. |
| INACTIVE | No login for 5+ consecutive days | Warning | Send a personal check-in. Research on dropout precursors (Tinto, 1987) shows that a brief personal contact during the first disengagement event is significantly more effective than intervention after sustained absence. |
Friction alerts are not about absence — they are about a present student who is visibly struggling with the platform or the content. They are derived from behavioral patterns in the student's session logs.
| Alert Type | Trigger Pattern | What It May Signal |
|---|---|---|
| URL_LOOP | Same page visited 3+ times in 4 hours without a meaningful action (note created, analysis run) | Student may not understand where to start, may be confused by the interface, or may be experiencing decision paralysis about what to do next. |
| NOTE_ABANDONED | Note created but no AI analysis run within 60 minutes | Student is writing but not analyzing. They may not know the lens system exists, may distrust the AI tool, or may be under time pressure and saving analysis "for later" (which research shows usually never happens). |
| FEATURE_BOUNCE | High-value feature page (Knowledge Graph, Simulation, Socratic Coach, Discussion Wheel) visited and exited in <30 seconds, 3+ times in a week | Student is aware of the feature but cannot engage with it — possibly because the required prerequisite (e.g., enough notes to populate a graph, or an active study group discussion) is not yet met, or the interface is unclear for their use case. |
Friction alerts are best used as prompts for a micro-intervention: a 60-second screen share, a targeted help message, or a classroom announcement that addresses the bottleneck you see across multiple students simultaneously. If 40% of your students are triggering NOTE_ABANDONED, that is a class-wide signal — address the analysis workflow in your next session rather than contacting students individually.
The Study Coaching feature generates a personalized study plan for each student by comparing what they have written about against what they should know — using the course syllabus as the benchmark. It is accessible both as an instructor-facing view (inside the Deep-Dive modal) and as a student-facing recommendation panel.
The process works in three steps:
Share the Study Coaching output with students before major assessments, not just before finals. Weekly coaching reviews normalize the act of self-diagnosis and reduce study anxiety by converting a vague feeling of "I don't know enough" into a specific list of addressable gaps. Students who engage with coaching recommendations at least once per week show significantly more consistent Health Score improvement trajectories.
The Case Workspace Audit Panel provides instructors with real-time visibility into collaborative group decision-making activities. MHLE's case workspaces follow a structured four-phase model — Frame → Analyse → Decide → Reflect — and each phase is tracked individually.
The audit panel shows:
The Frame-Analyse-Decide-Reflect structure is derived from problem-based learning (PBL) design principles (Hmelo-Silver, 2004) and the Kolb Experiential Learning Cycle (1984). Research on collaborative problem solving shows that structured scaffolding of the problem-framing phase — the "Frame" step — produces significantly better quality decisions downstream because groups that skip framing often solve the wrong problem.
Every classroom in MHLE supports one-click export of the full student roster and associated engagement metrics as a CSV file. The export is designed to comply with the Family Educational Rights and Privacy Act (FERPA) in the following ways:
The exported columns include: Student Name, Enrollment Date, Health Score (current), Note Count, AI Analysis Count, Last Login (days ago), Gap Status, Overall Rigor Grade, Exam Readiness Score, and Active Alerts count.
The Coach Dashboard extends the Instructor Dashboard's individual-student focus into a team and organizational lens. It is designed for enterprise coaches, department leaders, and learning & development professionals who manage cohorts of learners across organizational contexts rather than traditional academic classrooms.
The Coach Dashboard's signature metric is the Team Alignment Score — a measure of how much shared conceptual vocabulary and semantic overlap exists within a group of learners (2–6 participants). It is computed using Jaccard similarity across the participants' knowledge graph concept sets.
Alignment is not equivalent to agreement. A high alignment score means the team shares a common conceptual language — they are working from the same intellectual map. A low alignment score is not necessarily bad: it may reflect productive epistemic diversity. The Coach Dashboard helps you distinguish between "diverse perspectives" (healthy) and "knowledge siloes" (problematic) by surfacing which specific concepts are shared vs. isolated.
The comparison engine identifies three categories of insight when comparing participants' knowledge graphs:
Concepts that all compared participants own at Relational depth or above. These are the team's "common ground" — the basis for productive collaboration. High shared concept counts indicate team readiness for complex, interdependent work.
Concepts where participants hold opposing relationship types in their knowledge graphs — where one person's notes link Concept A as a cause of Concept B, and another's link them in reverse. These are rich sites for structured debate and re-examination.
Concepts present in the course benchmark or held by most participants but completely absent from a specific individual's knowledge graph. These are the most actionable coaching targets — specific, named gaps in a specific person's understanding.
Historical tracking of the Team Alignment Score over time. A rising trend indicates that team members are converging on shared frameworks through their learning activities — measurable evidence of collaborative knowledge building.
Coaches can generate downloadable alignment reports in PDF or Markdown format. These reports summarize the team's shared concepts, individual blind spots, alignment trend, and recommended next steps — suitable for inclusion in performance reviews, learning program assessments, or organizational capability reports.
The Coach Dashboard visualizes knowledge distribution across semantic "zones" — thematic clusters of concepts that tend to co-occur across learners' notes. These zones emerge organically from the platform's clustering algorithms, not from pre-defined categories. Seeing that your engineering team's knowledge overwhelmingly clusters in a "Technical Execution" zone while the "Strategic Context" zone is sparse is an immediate signal about organizational capability gaps — one that would take months to surface through traditional training assessments.
The Coach Dashboard applies the same Bloom's classification from the Note Rigor Scorer to every participant, then aggregates the results across the cohort. This produces a distribution — what percentage of your team is operating at Remember/Understand vs. Analyze/Evaluate on each knowledge domain. Enterprise coaches can use this to make targeted decisions about where to invest training resources: if 80% of participants are at Apply level on a topic but none have reached Evaluate, structured critique workshops will be more valuable than additional instructional content on that topic.
This section provides concrete, week-by-week guidance for integrating the MHLE Instructor Dashboard into a semester-length course.
Instructor Actions:
What to Watch On the Dashboard:
Research on habit formation (Clear, 2018; Lally et al., 2010) consistently identifies the 18–66 day window as the critical period for locking in learning behaviors. Weeks 2–4 are therefore the highest-leverage window for instructor intervention.
Instructor Actions:
During Weeks 2–4, pay particular attention to the Discussion Wheel as a complementary signal to note-taking data. A student whose comprehension scores in group discussion suggest surface-level engagement — Low comprehension despite high note counts — may be memorizing effectively but struggling to think on their feet. This is a productive coaching target early in the course, before it becomes a pattern that persists into high-stakes assessments.
Suggested Weekly Routine (15 minutes):
Minutes 1–3: Scan the classroom roster for any HEALTH_DROP or INACTIVE alerts generated since your last review. Note any students showing two consecutive weeks of decline.
Minutes 4–7: Open the Cohort Analytics panel. What are the current "Top Gaps"? Are there any concepts that more than 30% of the class is showing as Red (no coverage) or Amber (shallow)?
Minutes 8–12: Deep-dive on 1–2 students of concern. Is the Health Score drop driven by recency (they stopped logging in) or by lens coverage (they're writing notes but not analyzing)? Each requires a different response.
Minutes 13–15: Plan one class activity or discussion prompt for the coming week based on the coverage gaps you identified. Close the loop between the dashboard signal and your instructional response.
At the midpoint of the course, run a comprehensive Rigor Score review for the whole class. Export the roster CSV and sort by Exam Readiness Score. The bottom quartile of students on this metric — regardless of their grade in the course — are the students most likely to underperform on high-stakes assessments, even if they appear engaged.
This is the ideal moment to schedule brief individual Study Coaching conversations. Use the coaching tab in each student's deep-dive modal to anchor the conversation: "Your notes show strong Financial lens coverage, but I'm not seeing Ethical framing on any of the regulatory topics — let's talk about why that matters for the upcoming case study."
Two weeks before a major assessment, use the Coverage Grid to identify which concepts the class understands at depth (Relational level, Emerald cells) versus which ones are showing as Amber (shallow) or Red (missing). These shallow and missing concepts should receive the majority of your remaining class time — not concepts already at Full coverage.
This approach embodies the principle of deliberate practice (Ericsson et al., 1993): time on task is only valuable when it is targeted at the edge of current competence, not at areas of existing mastery.
Before your major assessment, also use the Simulation Calibration stats (Section 3.8) to verify that the simulation difficulty is well-matched to the class's current performance level. If most students are completing simulations at an undergraduate register but your course targets graduate-level reasoning, adjust the classroom's Wicked Problem Level override now — not after students have submitted low-quality solutions. The calibration preview (dry-run) is particularly useful in the week before an assessment to confirm that the scenario complexity is appropriate without consuming any student quota.
MHLE's system includes a Portfolio Learning Artifact System that transforms the accumulated course learning artifacts into professional portfolio pieces with AI-generated summaries. At course end, encourage students to:
For courses with group projects or collaborative learning components, assign students to MHLE Study Groups. The Coverage Grid then becomes a group planning tool: groups can see at a glance where they have collective blind spots, identify the strongest member on each concept for peer mentoring, and use the SOLO depth display to determine which concepts need more work before a group deliverable is due.
Assign students to study groups in Week 2 before deep content exposure. Early group formation, combined with visible gap tracking via the Coverage Grid, creates social accountability for coverage — students are more motivated to fill in their own gaps when they can see that the group depends on their individual coverage. This is the distributed expertise principle from Vygotsky's Zone of Proximal Development applied to peer learning.
The power of any learning analytics platform raises important ethical questions. MHLE's design reflects a set of explicit commitments about how data is used, what data is surfaced, and who is accountable for decisions driven by that data.
We recommend that instructors be fully transparent with students about what the dashboard measures and why. Students who understand that their Health Score is a formative learning tool — not a grade — engage with it more productively than students who perceive it as surveillance. Consider sharing the Health Score formula with your class in Week 1 and framing it as a self-management tool that you happen to also be able to see.
The Health Score, Rigor Score, and Gap metrics in MHLE are inputs to instructor judgment, not replacements for it. A student with a Health Score of 45 may have extraordinary external circumstances. A student with a Rigor Score of "C" may be writing notes in a second language. Always use dashboard signals to prompt inquiry, never to render automatic conclusions about student ability or effort.
MHLE is built with FERPA principles in operation throughout:
The Note Rigor Scorer's spell-error rate component uses a system dictionary that may disadvantage students writing in English as a second language, using technical domain vocabulary that departs from standard spelling, or using names and terminology from non-dominant cultural contexts. Instructors should be aware of this limitation and apply professional judgment when interpreting Writing Quality scores for students in these categories.
Similarly, the Action Verb detection for Bloom's classification performs best with analytical writing in standard academic English. Students whose cognitive depth is high but whose expression is less conventional may be undercounted by the Bloom's classifier. Again, the deep-dive modal exists precisely to give instructors the full picture beyond any single metric.
Students have full access to their own Rigor Scores, Health Score history, Knowledge Graph, and Study Coaching recommendations within MHLE. The platform is deliberately designed so that students see exactly what their instructor sees. There are no "hidden" instructor-only metrics that create an information asymmetry unknown to the student — only the instructor's ability to view aggregate classroom data and to receive alerts on the student's behalf.