Skip to content
CodeWiki
Practice
Paths
Tracks
Cheatsheets
Playground
Glossary
AI era
Search
⌘K
English
English
Chinese
Practice
Paths
Tracks
Cheatsheets
Playground
Glossary
AI era
English
English
Chinese
Practice
/
Quiz
/
Coding in the AI era
/
LLM application evals
Which evaluation-set design gives the strongest release evidence for a customer-support a…
from LLM application evals
Node 24
advanced
2 min
Which evaluation-set design gives the strongest release evidence for a customer-support assistant?
A versioned mix of sampled task categories, costly incidents, boundary cases, adversarial cases, and predefined user slices
The ten demonstrations used while editing the system prompt
A large synthetic set generated by the same prompt and model under evaluation
Only the most common happy-path question, repeated with different order IDs
Check
Ask AI about this kata
Report an error
previous kata
Before an LLM judge blocks releases, what is the most useful calibration step?
Quiz
next kata
What does this release gate print?
Predict the output
A new version is available
Reload