--- # Audit Summary **CODEBASE AUDIT RESULT:** HIGH **AGENT REPRODUCIBILITY:** False --- # Detailed Code Audit Report ## Executive Summary This submission contains code for evaluating LLM-based code generation on HumanEval with a "Self-Spec" orchestration pipeline. The codebase has **critical bugs that prevent correct baseline result generation**, missing result files for claimed experiments, and uses non-standard/future API endpoints. While the spec-driven pipeline appears functional, the baseline evaluation has a severe bug that assigns the same generated solution to all tasks, making those results invalid. --- ## 1. COMPLETENESS & STRUCTURAL INTEGRITY ### Critical Issues Found #### 1.1 Variable Scope Bug in Baseline Evaluation ⚠️ **CRITICAL** **Location:** `evaluation_baseline_pipeline.py`, lines 94-108 ```python # Line 77: 'code' is defined in the loop scope for task_id, code in codes: task = index[task_id] passed, err = run_test(task, code) results[task_id] = {"passed": passed, "error": err} # Lines 94-100: BUG - 'code' variable reused from last iteration with open(f"all_baseline_{MODEL}.jsonl", "w", encoding="utf-8") as f: for tid, res in results.items(): item = index[tid] item["generated_solution"] = code # ← Uses last value from previous loop! ``` **Impact:** All 164 tasks in the baseline results file receive the same `generated_solution` value (the last one from the loop). This was verified by examining `all_baseline_claude_3_5_sonnet_20241022.jsonl`: ``` HumanEval/0: def generate_integers(a, b): start = min(a, b) HumanEval/1: def generate_integers(a, b): start = min(a, b) HumanEval/2: def generate_integers(a, b): start = min(a, b) ... ``` **Severity:** This makes the baseline results **completely invalid** as stored. However, the summary statistics (147/164 passed) may still be valid if computed before this bug occurs. The code appears to run the tests correctly with the actual generated code, but then saves the wrong code to the output file. #### 1.2 Empty API Keys All main scripts have empty API key placeholders: - `spec_result_generation.py:43`: `API_KEY = ""` - `evaluation_baseline_pipeline.py:8`: `API_KEY = ""` - `humaneval_convert.py:26`: `API_KEY = ""` **Impact:** Code cannot run without manual API key insertion. No environment variable fallback is properly implemented. #### 1.3 Non-Standard API Usage **Location:** `humaneval_convert.py:57` ```python resp = await client.responses.create( model=MODEL, input=prompt_text, ) ``` The code uses `client.responses.create()` which is **not a standard OpenAI API endpoint**. Standard OpenAI uses `client.chat.completions.create()`. This suggests: - Custom/internal API wrapper - Future API version not yet released - Potential documentation error **Model specified:** `"gpt-5-2025-08-07"` (line 27) - a model that doesn't exist as of the code timestamp. --- ## 2. RESULTS AUTHENTICITY RED FLAGS ### 2.1 Missing Result Files The README (lines 35-46) acknowledges that two result files are missing: - **GPT-4o baseline** (needed to verify 87% → 92% claim) - **Claude 3.5 spec** (needed to verify 90% → 89% claim) README states these were "originally saved but later misplaced." This is concerning for reproducibility. ### 2.2 Results Present vs. Paper Claims **Available results:** - ✅ Claude 3.5 baseline: 147/164 = 89.6% (paper claims 90%) - ✅ Claude 3.7 baseline: 151/164 = 92.1% (paper claims 92%) - ✅ Claude 3.7 spec: 154/164 = 93.9% (paper claims 94%) - ✅ GPT-4o spec: 150/164 = 91.5% (paper claims 92%) - ❌ GPT-4o baseline: MISSING (paper claims 87%) - ❌ Claude 3.5 spec: MISSING (paper claims 89%) **Analysis:** - Results present are within 0.5% of claimed values (rounding differences) - Two critical comparison files are missing - The baseline bug means we cannot verify the actual generated code matches what was tested ### 2.3 Hardcoded Results? The result summary files contain only simple counts: ``` Total: 164, Passed: 147, Failed: 17 ``` These are **computed** by the evaluation scripts, not hardcoded. However, due to the baseline bug, the `all_baseline_*.jsonl` files contain incorrect `generated_solution` fields, making it impossible to independently verify which solutions passed. --- ## 3. IMPLEMENTATION-PAPER CONSISTENCY ### 3.1 Orchestration Pipeline Implementation The spec-driven pipeline in `spec_result_generation.py` implements the 6-stage process described in the paper: 1. **SpecDesigner** (lines 74-87): Creates global specification schema ✅ 2. **SpecInstantiator** (lines 119-136): Populates schema from NL request ✅ 3. **FMInterviewer** (lines 89-117): Asks clarifying questions ✅ 4. **SpecApplier** (lines 139-166): Updates spec with user answers ✅ 5. **FMConfirmer** (lines 168-188): Summarizes for confirmation ✅ 6. **SpecExecutor** (lines 190-212): Generates code from spec ✅ The implementation appears faithful to the paper's description. ### 3.2 Experimental Setup Consistency - **Benchmark:** HumanEval (164 tasks) ✅ - **Temperature:** 0 (deterministic) ✅ - **Models:** Claude 3.5, Claude 3.7, GPT-4o (partial - missing GPT-4o baseline) ⚠️ - **Evaluation:** Standard HumanEval test harness ✅ ### 3.3 Hyperparameters Match - Temperature = 0 consistently used ✅ - Concurrency = 8-10 for parallel API calls ✅ - Max rounds = 5 for clarification iterations ✅ --- ## 4. CODE QUALITY SIGNALS ### 4.1 Positive Indicators - Clean async/await patterns for concurrent API calls - Proper error handling in some places (e.g., `humaneval_convert.py:55-83`) - JSONL format for reproducible result storage - Git repository with commit history ### 4.2 Negative Indicators #### Dead/Commented Code - `evaluation_baseline_pipeline.py:10`: Commented Anthropic base_url - `spec_result_generation.py:521`: Commented api_key configuration - Multiple unused imports #### Inconsistent Error Handling - `spec_result_generation.py:45-46`: Raises error if API_KEY is empty, but then never reads from environment - `evaluation_baseline_pipeline.py:112-113`: Checks if `API_KEY == "YOUR_API_KEY_HERE"` but actual value is `""` #### Code Quality Issues - Magic string comparisons: `if "CONFIRM" in str(user_confirmation).strip().upper()` (line 434) - Inconsistent naming: `GPT5LLM` class name doesn't match actual models used - No type hints on several critical functions - Mixed language in comments (some Chinese characters: line 113 "⚠️ NEED OPENAI_API_KEY 。") --- ## 5. FUNCTIONALITY INDICATORS ### 5.1 Data Loading ✅ - Proper JSONL reading/writing functions (lines 501-514 in `spec_result_generation.py`) - HumanEval dataset loading from datasets library (`humaneval.py`) ### 5.2 Test Execution ✅ The evaluation code properly: - Wraps generated code with imports and test functions - Uses `exec()` to run tests in isolated namespace - Captures assertion errors and exceptions - Handles markdown code fences in generated output ### 5.3 Spec-Driven Pipeline ✅ The orchestration pipeline appears functional: - Implements all 6 stages - Handles multi-round clarification - Uses proper async patterns for parallel execution - Produces structured output in JSONL format ### 5.4 Baseline Pipeline ❌ **Critical Bug:** As documented in Section 1.1, the baseline pipeline has a severe bug that makes stored results invalid. --- ## 6. DEPENDENCY & ENVIRONMENT ISSUES ### 6.1 Dependencies **Required:** - `openai` (for API access) - `datasets` (for HumanEval loading) - `asyncio` (standard library) - Python 3.8+ **Issues:** - No `requirements.txt` file - Version requirements not specified - Uses OpenAI client for both OpenAI and Anthropic APIs (via base_url), which may confuse users ### 6.2 Non-Existent Models - `humaneval_convert.py:27`: Uses `"gpt-5-2025-08-07"` which doesn't exist - This suggests either: - Code written for future API - Placeholder that wasn't updated - Custom model naming in internal system ### 6.3 API Compatibility - Uses `client.responses.create()` instead of standard `client.chat.completions.create()` - This non-standard API may not work with public OpenAI/Anthropic endpoints --- ## 7. REPRODUCIBILITY ASSESSMENT ### 7.1 What Can Be Reproduced? ✅ **Spec-driven results:** Code appears functional, can regenerate if: - Valid API keys provided - Non-standard API issues resolved - Required dependencies installed ❌ **Baseline results:** Cannot reproduce due to bug in `evaluation_baseline_pipeline.py` ❌ **Missing experiments:** GPT-4o baseline and Claude 3.5 spec cannot be reproduced without regeneration ### 7.2 Reproducibility Barriers **HIGH Priority:** 1. Baseline evaluation bug prevents correct result storage 2. Empty API key placeholders with no clear environment variable setup 3. Non-standard API usage (`responses.create()`) 4. Missing result files for 2 of 6 experiments **MEDIUM Priority:** 1. No requirements.txt or dependency specification 2. Non-existent model names in code 3. Inconsistent API key handling across files **LOW Priority:** 1. No automated test suite 2. Mixed language in comments/messages 3. Minor code quality issues --- ## 8. AGENT REPRODUCIBILITY **Finding:** No evidence of documented AI-assisted code generation process. The repository contains: - ❌ No prompt logs showing AI assistance in development - ❌ No documentation of which code was AI-generated - ❌ No conversation logs with AI coding assistants - ✅ Git commit history (but with minimal commit messages like "1", "readme1") **Note:** The research is *about* using AI for code generation (Self-Spec pipeline), but there's no evidence that the implementation code itself was generated or assisted by AI agents in a documented way. --- ## 9. SPECIFIC RED FLAGS SUMMARY ### CRITICAL (Prevent Verification) 1. **Baseline bug:** All tasks assigned same `generated_solution` in output files 2. **Missing experiments:** 2 of 6 claimed result sets are missing 3. **Non-standard API:** Code may not work with public APIs ### HIGH (Major Implementation Issues) 1. **Empty API keys:** No clear path to running code 2. **Non-existent models:** References to future/non-existent model names 3. **Inconsistent error handling:** API key validation doesn't match actual usage ### MEDIUM (Quality Concerns) 1. **No requirements file:** Dependencies not specified 2. **Commented code:** Suggests trial-and-error development 3. **Mixed languages:** Inconsistent code comments ### LOW (Minor Issues) 1. Code style inconsistencies 2. Magic strings in conditionals 3. Minimal git commit messages --- ## 10. RECOMMENDATIONS ### For Authors 1. **URGENT:** Fix the baseline evaluation bug in `evaluation_baseline_pipeline.py` lines 94-108 2. **URGENT:** Regenerate missing GPT-4o baseline and Claude 3.5 spec results 3. Add `requirements.txt` with pinned versions 4. Document the non-standard API usage or convert to standard OpenAI API 5. Implement proper environment variable handling for API keys 6. Add integration tests to catch bugs like the baseline issue 7. Clean up commented code and add proper documentation ### For Reviewers/Users 1. **Do not trust the `generated_solution` fields in baseline result files** - they are incorrect due to the bug 2. The pass/fail statistics may still be valid if you trust the evaluation logic ran before the storage bug 3. Spec-driven results appear more reliable but cannot be independently verified without GPT-4o baseline 4. Reproduction will require significant debugging due to API compatibility issues --- ## 11. CONCLUSION This codebase demonstrates a **sophisticated orchestration pipeline** for spec-driven code generation, but suffers from **critical implementation bugs** in the baseline evaluation that invalidate stored results. The presence of non-standard API usage, missing result files, and lack of proper dependency management significantly hampers reproducibility. **Overall Assessment:** While the core Self-Spec algorithm appears well-implemented, the evaluation infrastructure has serious issues that prevent independent verification of the paper's claims. The missing baseline experiments are particularly concerning given they're needed for the main comparisons in the paper. **Recommended Action:** Authors should fix the baseline bug, regenerate all results with proper validation, and provide clear documentation for the non-standard API requirements before claiming full reproducibility.