Teaching debugging through production incidents
A traditional coding assignment tells you what is wrong. A real production incident does not. I ideated, designed and built Debug Simulator at Scaler so learners could investigate a broken system, explain its root cause, and see whether their fix actually worked. To mimic the actual job that a developer would do in real life.
Opens the live product. You may need to sign up for a Scaler account first.
The learner’s workspace
My role
Full Stack Designer. Took the idea through design, build and production. Own the learner experience, admin system, authoring pipeline, agent architecture and releases.
April – June 2026
Team
Scaler trains working professionals and school graduates for software and business careers. Learners pay for a job outcome, so course time needs to build skills they can use at work. Instructors and content teams supply the subject briefs; the authoring pipeline turns them into playable cases.
Outcomes
In production
55+ live cases
Since July 2026
864 completed attempts
by 362 learners
Learner rating
4.5 / 5
across 624 reviews
Time per case
~23 minutes
Case authoring
~30 minutes
from about one working day
Adoption
Part of the curriculum
The simulator is used as part of the curriculum, with cases lasting about 23 minutes on average. 95% of positive reviews from learners.
Teaching the unthaught
Traditional coding assignments hand over the bug and ask only for a fix, so the investigation never happens. Learners could write the fix but had not practised finding the fault, reading partial evidence in an unfamiliar system, or explaining why the failure happened now.
A normal assignment gives the input, expected output and function to write. A real incident might only say that checkout has been slow since this morning. The developer has to understand the system, read logs and metrics, connect clues, explain why the failure happened now, and communicate a root-cause analysis.
Harder assignments would have been cheaper to build, but kept the same limitation. Learners could ask an AI to solve a stated problem and skip the investigation. The format itself had to change.
Four stages, one investigation
Understand. Read the incident or watch a video, then work out how the healthy system behaves.
Investigate. Open components, read logs, follow alerts and answer the coach’s questions.
Submit. Explain the root cause, why it happened now, and the immediate and long-term fixes.
Review. Get feedback on search behaviour, evidence, reasoning, explanation, speed and fix quality.
Investigate a connected system
From a log line to a reasoning question
01 · Read the evidence
02 · Explain the pattern
The investigation asks for more than naming a broken component. Learners explain the causal chain in their own words and separate an immediate mitigation from a lasting fix. Feedback then makes the quality of that reasoning visible.
The system responds to the fix
A chosen fix changes the simulated system. Treat a symptom and the indicators dip, then return. Fix the cause and they settle. Learners see the consequence of their reasoning.
An end score alone cannot show how a system behaves after an intervention. Applying the action closes that gap. The review explains what was incomplete and how the learner can approach the next incident differently.
A direct feedback loop
When a learner reports a bug or a confusing part of the experience, I contact them, run a short interview, fix the issue, deploy and tell them what changed. The reported resolution turnaround is three hours. Keeping this loop manual helps connect production decisions to a learner’s actual experience.
What made the difference
Changed the exercise. A learner gets an incident and partial evidence instead of a stated bug, so finding the cause is the work.
Made the fix provable. The simulated system reacts to the chosen fix, so treating a symptom looks different from solving the cause.
Made new cases cheap to produce. The agent pipeline cut authoring from about a working day to 30 minutes, with a person still at the release gate.

