Prompt
What should I use to evaluate code generation outputs?
Latest observation
Oct 9, 2026
OpenAI APIWeb search: off
To evaluate code generation outputs, use a mix of automatic correctness checks and human/behavioral evaluation. The best choice depends on your task, but a solid setup usually includes:
1. Functional correctness
- Unit tests / integration tests: Most important for generated code.
- Execution-based evaluation: Run the generated code against hidden or public test cases.
- Pass@k: Measures whether at least one of the top k generated samples passes the tests.
2. Static quality checks
- Linting / formatting: e.g., ESLint, pylint, black, gofmt.
- Type checking: e.g., mypy, TypeScript compiler, Rust compiler.
- Compilation success: Basic check for syntactic validity.
3. Robustness and reliability
- Edge-case test suites: Inputs near boundaries, malformed inputs, stress tests.
- Mutation testing: See whether tests catch small injected bugs.
- Fuzzing: Good for generated parsers, validators, and input-handling code.
4. Semantic and structural quality
- Code similarity / exact match: Useful only in narrow settings; often misleading.
- AST-level comparisons: Better than raw text when structure matters.
- Documentation/comment consistency: If code should match an explanation or prompt.
5. Human evaluation
- Correctness review
- Readability / maintainability
- Security review
- API/design fit
6. Domain-specific metrics
- For SQL generation: execution accuracy, result-set match.
- For code repair: patch correctness and regression tests.
- For code translation: compiled output + behavioral equivalence.
- For refactoring: preserve tests + style improvement.
Practical recommendation
If you want a default evaluation stack, use:
- Compile/type-check
- Run tests
- Measure pass@k
- Add human review for ambiguous cases
- Add static analysis for quality/security
If you tell me your exact setting — e.g. code completion, problem-solving, code repair, or text-to-SQL — I can recommend the best metric(s) and evaluation pipeline.