Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024
[ICLR'25] BigCodeBench: Benchmarking Code Generation Towards AGI