BENCHMARK AUDIT: TRACK 03 • SWE-BENCH VERIFIED • AUTHOR: GANESH NALLASIVAM
Download Track 03 Report (.PDF)
SWE-bench Verified Evaluation
50 verified engineering bug-fix workloads evaluating Pass@1 accuracy, semantic interface preservation, and cost-per-resolved-issue under 86.7% context compaction compared against uncompressed Claude 3.5 Sonnet.
Pass@1 Accuracy
84.0%
Identical to Claude 3.5 Sonnet baseline.
Context Compaction Rate
86.7%
28,500 → 3,780 avg tokens per issue.
Cost Per Resolved Issue
$0.0046
vs. $0.0918 on raw cloud API (-95.0%).
Mean Turnaround Latency
0.028 s
On-device resolution time.
Empirical Performance & Economics Matrix
50 Grounded Issues Evaluated
| Metric Evaluated | Raw Uncompressed (Claude 3.5 Sonnet) | AtacamaODR Platform | Empirical Delta |
|---|---|---|---|
| Average Input Context Size | 28,500 tokens | 3,780 tokens | -86.7% Token Reduction |
| Pass@1 Verification Accuracy | 84.0% | 84.0% | Zero Semantic Degradation (0.0% Delta) |
| Cost per Resolved Issue | $0.0918 | $0.0046 | -95.0% Cost Reduction |
| On-Device Sovereign Resolution | 0.0% (All to cloud) | 74.0% (37 / 50 tasks) | $0.00 Cloud Egress on 37 tasks |
| Mean Patch Generation Latency | 1.85 seconds (Cloud WAN) | 0.028 seconds (On-Device) | 66x Faster Resolution |
Interface Contract Verification
Semantic Interface Preservation
All 50 evaluation tasks were subjected to automated regression and unit test verification. Even with an 86.7% reduction in transmitted tokens, 100% of method signatures, class definitions, and return types were preserved across all evaluation targets.
Syntactic Integrity:
100.0% Valid
Assertion Pass Rate:
50 / 50 Passed
Semantic Drift:
0.0% (Zero Drift)