AI
Evaluate language model systems
LLM evaluation is the practice of checking model or application outputs against cases you care about: correctness, format, refusal, and regression after a change.
For engineers shipping prompts, RAG, or fine-tunes who need a repeatable check rather than a vibe.
Start here
Read the OpenAI evals README, then write ten cases from your own task before you look at a public leaderboard.
Recommended resources
Start here · IntermediateOpenAI Evals
A concrete framework for writing task cases and running them again after a change.
github.comReference · IntermediateOpen LLM Leaderboard
Shows how public benchmarks compare base models, and where those scores stop describing your application.
huggingface.coRAG · IntermediateRAGAS documentation
The evaluation set to use when the system retrieves documents before it answers.
docs.ragas.ioContext · IntermediateFine-tuning guide
Training and evaluation belong in the same loop. This is the training side.
huggingface.coRelated topics
Evaluate a RAG system · Fine-tune a language model · Learn PyTorch