Prompt
How do I evaluate code generation output from an LLM?
Latest observation
Evaluating code generation from an LLM works best if you judge it on multiple dimensions, not just “does it run?”
1) Define what “good” means for your task
Start by deciding the target:
- Correctness: Does it solve the problem?
- Safety: Is it secure and free of dangerous behavior?
- Readability: Is it maintainable by humans?
- Efficiency: Does it meet performance constraints?
- Reliability: Does it handle edge cases and failures?
- Style/compliance: Does it follow your coding standards?
A code model can be excellent at one and poor at another.
2) Use test-based evaluation
The most practical approach is to run code against tests.
a) Unit tests
- Write tests for expected behavior and edge cases.
- Include normal cases, boundary conditions, and invalid inputs.
- Measure pass rate.
b) Hidden tests
If the model can “overfit” visible examples, keep some tests hidden.
- This helps measure generalization.
c) Property-based tests
Useful when there are many valid outputs or input combinations.
- Example: sorting output must be ordered and preserve elements.
d) Integration tests
If code touches databases, APIs, files, or other services, test those interactions too.
3) Check compile/run success separately from correctness
Track:
- Syntax/compile rate
- Runtime success rate
- Test pass rate
A model may produce code that compiles but is logically wrong, or code that is mostly correct but has small syntax issues.
4) Evaluate on representative benchmark tasks
Use a dataset or benchmark close to your real use case. Examples of evaluation dimensions:
- Simple algorithmic tasks
- Code completion tasks
- Bug fixing tasks
- Real-world repo tasks
- API usage tasks
The closer the tasks are to production, the more meaningful the results.
5) Measure human quality aspects
Even if tests pass, humans may still reject the code. Review for:
- Clarity and naming
- Structure and modularity
- Idiomatic style
- Comments and documentation
- Avoiding unnecessary complexity
- Correct use of libraries/frameworks
A lightweight code review rubric helps.
6) Check robustness and edge cases
Ask:
- What happens with empty inputs?
- Null/None values?
- Large inputs?
- Unexpected formats?
- Network failures?
- Timeouts?
- Permission issues?
Good code generation should anticipate these.
7) Evaluate safety and security
For generated code, review for:
- Injection risks
- Unsafe deserialization
- Hardcoded secrets
- Unsafe shell execution
- Path traversal
- Weak crypto
- Missing authentication/authorization checks
Security review matters especially for backend, data, and DevOps code.
8) Use task success metrics
Depending on the use case, consider:
- Exact match for short code snippets
- Functional correctness via tests
- Pass@k: whether at least one of k samples is correct
- Latency
- Token cost
- Number of edits required by a human
- Bug rate after deployment
9) Compare against baselines
To know if the model is actually useful, compare it to:
- A human baseline
- Another LLM
- A simpler heuristic/template
- Existing tooling/autocomplete
10) Run an error analysis
For failed outputs, categorize failures:
- Wrong algorithm
- Missing edge cases
- API misuse
- Hallucinated functions/libraries
- Off-by-one errors
- Incomplete code
- Style issues
- Security issues
This helps you improve prompts, training, or post-processing.
11) Use a scorecard
A simple rubric might be:
- Correctness: 0–5
- Robustness: 0–5
- Readability: 0–5
- Security: 0–5
- Efficiency: 0–5
Then compute averages across many tasks.
12) If you want an automated setup
A good pipeline is:
- Generate code
- Format/lint it
- Compile it
- Run unit tests
- Run hidden tests
- Static analysis/security checks
- Optional human review
Practical recommendation
If you’re starting out, the best high-signal evaluation is:
- A curated test suite
- Hidden edge-case tests
- Manual review of a sample of outputs
- Error breakdown by failure type
If you want, I can also give you:
- a rubric/template for evaluating code generation,
- a Python script to automate evaluation, or
- advice tailored to a specific language or use case.