2026 Data4Good Case Competition

Regional Champions, Region 1 (Feb 2026)

Problem: Generative AI is increasingly used in education for tutoring, content creation, grading support, and doubt solving. A key risk is hallucinated or incorrect answers that can quietly undermine learning outcomes and trust. Most learning environments cannot rely on continuous human review or expensive infrastructure, creating the need for a low-cost, automated, and auditable system that verifies AI answers before they reach learners.

Approach: We framed the task as a three-class classification problem using Question, Context, and Answer. We combined ensemble ML models using semantic and structural features with an LLM-as-a-Judge framework using a two-step, relevance-first reasoning flow with few-shot calibration. The decision logic checks relevance first, then evaluates correctness against context.

Solution: We delivered a deployment-ready, safety-first evaluation pipeline combining an ensemble ML baseline with an LLM-as-a-Judge layer, wrapped in a batch-based, auditable inference pipeline. The system prioritizes catching harmful errors, produces consistent decisions, and can be integrated as a pre-publish safety check in learning platforms.

Live Demo: LLM as a Judge

All case competitions

Anushka Mathur | Marketing Analytics Portfolio