BENCHMARK AUDIT: TRACK 04 • CANONICAL BLIND HUMANEVAL (164 PROBLEMS) • AUTHOR: GANESH NALLASIVAM
Download Track 04 Report (.PDF)
Canonical Blind HumanEval Benchmark
Full 164-problem blind execution benchmark evaluating Qwen2.5-Coder-14B-Instruct (4-bit) running on Apple Silicon via AtacamaODR (Port 8088). Every test case was executed in an isolated sandboxed subprocess with strict 5.0-second timeouts against official OpenAI canonical assertions. Zero synthetic mock data. Zero prompt overfitting regexes. 100% physically executed code.
Pass@1 Accuracy
86.0%
141 of 164 passed.
Total Problems Evaluated
164 / 164
100% completed suite.
Average Turnaround
13.8 s
Total runtime: 37.7 min.
Execution Failures
23
19 assertion, 4 syntax errors.
Industry Standard Python Coding Accuracy (HumanEval Pass@1)
Grounded Empirical Comparison
| Model Architecture | Parameters / Quant | Host Environment | Pass@1 Accuracy | Status |
|---|---|---|---|---|
| Qwen2.5-Coder + AtacamaODR | 14B (4-bit MLX) | Apple Silicon (24GB UMA) | 86.0% (141/164) | Physically Verified |
| Claude 3.5 Sonnet (Frontier Reference) | Unknown (Cloud) | Anthropic Cloud API | 93.7% | Published Lab Metric |
| GPT-4o (Frontier Reference) | Unknown (Cloud) | OpenAI Cloud API | 90.2% | Published Lab Metric |
| Llama 3.1 70B Instruct | 70B (16-bit / 8-bit) | Enterprise Cloud Host | 80.5% | Published Lab Metric |
| Qwen2.5-Coder-7B-Instruct | 7B (4-bit) | Local Apple Silicon | 79.9% | Published Model Card |
Audit Methodology & Error Breakdown
Empirical Execution Telemetry
All 164 problems were piped via HTTP to http://127.0.0.1:8088/v1/chat/completions, parsed for Python function blocks, concatenated with canonical test cases, and executed in temporary isolated files via python3 subprocesses with 5.0-second timeouts.
Passed Unit Tests:
141 Problems
Exit Code 0 (Pass@1)
Logical / Assertion Errors:
19 Problems
Canonical assertion mismatch
Syntax / Parsing Errors:
4 Problems
Unterminated strings / syntax
Dataset: OpenAI Canonical HumanEval.jsonl (164 tasks) Artifact: humaneval_execution_results.json