Back

Teaching debugging through production incidents

A traditional coding assignment tells you what is wrong. A real production incident does not. I ideated, designed and built Debug Simulator at Scaler so learners could investigate a broken system, explain its root cause, and see whether their fix actually worked. To mimic the actual job that a developer would do in real life.

Try Debug Simulator

Opens the live product. You may need to sign up for a Scaler account first.

The learner’s workspace

Understand stage of a live case
Actual case screen: signup pipeline nodes on the left, and a triage panel asking which first move to make when two dashboards report different signup totals.

My role

Full Stack Designer. Took the idea through design, build and production. Own the learner experience, admin system, authoring pipeline, agent architecture and releases.

April – June 2026

Team

Kishan VagaleKishan Vagale talks about Claude Code, Codex, Grok, Aside Browser, Figma, Illustrator, Photoshop, Notion, Jira, Sheets, Mixpanel, User research, Roadmaps, Design systems.Idea to production
SudhanvaSudhanva talks about Figma, YouTube, Vibe coding.Frontend UI
Chanchal MishraChanchal Mishra talks about React, TypeScript, Next.js, Tailwind CSS, Storybook, GitHub, Claude Code, Design systems, Prototyping.Single sign-on

Scaler trains working professionals and school graduates for software and business careers. Learners pay for a job outcome, so course time needs to build skills they can use at work. Instructors and content teams supply the subject briefs; the authoring pipeline turns them into playable cases.

Outcomes

In production

55+ live cases

Since July 2026

864 completed attempts

by 362 learners

Learner rating

4.5 / 5

across 624 reviews

Time per case

~23 minutes

Case authoring

~30 minutes

from about one working day

Adoption

Part of the curriculum

The simulator is used as part of the curriculum, with cases lasting about 23 minutes on average. 95% of positive reviews from learners.

Teaching the unthaught

Traditional coding assignments hand over the bug and ask only for a fix, so the investigation never happens. Learners could write the fix but had not practised finding the fault, reading partial evidence in an unfamiliar system, or explaining why the failure happened now.

A normal assignment gives the input, expected output and function to write. A real incident might only say that checkout has been slow since this morning. The developer has to understand the system, read logs and metrics, connect clues, explain why the failure happened now, and communicate a root-cause analysis.

Harder assignments would have been cheaper to build, but kept the same limitation. Learners could ask an AI to solve a stated problem and skip the investigation. The format itself had to change.

Four stages, one investigation

Understand. Read the incident or watch a video, then work out how the healthy system behaves.

Investigate. Open components, read logs, follow alerts and answer the coach’s questions.

Submit. Explain the root cause, why it happened now, and the immediate and long-term fixes.

Review. Get feedback on search behaviour, evidence, reasoning, explanation, speed and fix quality.

Investigate a connected system

The Coffee-Break Slowdown system map
Actual incident map with checkout service, order database, catalog table, reporting refresh, read replica and analytics warehouse. A Check this prompt points at the checkout service.

From a log line to a reasoning question

01 · Read the evidence

02 · Explain the pattern

01 · Read the evidence
Cropped checkout logs showing latency climbing at 09:00 and recovering at 09:18.

The investigation asks for more than naming a broken component. Learners explain the causal chain in their own words and separate an immediate mitigation from a lasting fix. Feedback then makes the quality of that reasoning visible.

The system responds to the fix

A chosen fix changes the simulated system. Treat a symptom and the indicators dip, then return. Fix the cause and they settle. Learners see the consequence of their reasoning.

Conceptual system response
Qualitative curves contrast temporary relief followed by recurrence with a sustained recovery. No measured scale.

An end score alone cannot show how a system behaves after an intervention. Applying the action closes that gap. The review explains what was incomplete and how the learner can approach the next incident differently.

Making the next case in about 30 minutes

Instructors provide a structured Markdown brief. An agent pipeline I built on Claude Managed Agents, on Anthropic’s API platform, authors the case, reviews teaching and functionality, generates narration and art, and opens a pull request. I test on staging before releasing.

Case authoring and human release pipeline
Six stages: instructor brief, author and content review, narration and art, browser reviews, pull request, and human staging and release.

Generation time

~30 minutes

versus about one working day by hand

Generation cost

$8–15 per case

depending on review and retry cycles

Voice assets

~100 per case

Writing 50 cases by hand would have made content production the bottleneck. I designed a Markdown template that instructors complete with Claude, which asks questions until the context is sufficient. Uploading the brief through the admin interface starts the pipeline on Claude Managed Agents.

  • The author builds the case using reusable modules for logs, graphs, questions and system nodes.
  • The content reviewer checks correctness and readability.
  • ElevenLabs and an in-house voice generate the coach’s narration.
  • A functional reviewer plays the case in a browser and checks interactions, calculations, logs and facts. A final reviewer repeats the browser review after edits.
  • GPT Image 2 generates the case poster within a controlled visual style.
  • The GitHub step checks conflicts and opens a pull request. A Slack notification brings me back into the release process.

I kept a human as the final release gate. A clean pull request still goes through staging, testing and a deliberate production release.

A direct feedback loop

When a learner reports a bug or a confusing part of the experience, I contact them, run a short interview, fix the issue, deploy and tell them what changed. The reported resolution turnaround is three hours. Keeping this loop manual helps connect production decisions to a learner’s actual experience.

What made the difference

Changed the exercise. A learner gets an incident and partial evidence instead of a stated bug, so finding the cause is the work.

Made the fix provable. The simulated system reacts to the chosen fix, so treating a symptom looks different from solving the cause.

Made new cases cheap to produce. The agent pipeline cut authoring from about a working day to 30 minutes, with a person still at the release gate.

Back
Kishan Vagale

What are you curious about?

Checking limits