# Audit Summary **CODEBASE AUDIT RESULT:** HIGH **AGENT REPRODUCIBILITY:** False --- # Detailed Code Audit Report for Submission 64 ## 1. COMPLETENESS & STRUCTURAL INTEGRITY ### Critical Issues Identified: **1.1 Missing Ablation Study Implementation (CRITICAL)** - The paper claims ablation study results (Table 2) showing: - Full Debate: 74.3% - Without Uncertainty Phase: 66.3% - Without Role-Switch Phase: 65.0% - Without Both: 62.0% - **Finding:** No code exists to run these ablation experiments - The `debate.py` file only implements the full protocol with all components - There is no mechanism to disable role-switching or uncertainty phases - **Impact:** Cannot verify the claimed ablation results **1.2 Incomplete Project Configuration** - No `requirements.txt`, `setup.py`, or dependency specification file - Line 11 in `debate.py` has empty project ID: `project=""` - **Impact:** Cannot run the code without manual configuration **1.3 Missing Test Set Code** - Paper claims results on "OpenBookQA test set (500 questions)" - Code uses validation set: `openbook_val = openbook["validation"]` - Comment indicates 500 examples, which is the validation set size - **Impact:** Results may not be on the claimed test set ### Positive Findings: - Main debate implementation appears complete with all phases - No TODO comments, placeholder functions, or `pass` statements - Proper error handling with retries (5 attempts with 60s wait) - Core logic is well-structured and readable ## 2. RESULTS AUTHENTICITY RED FLAGS ### Critical Discrepancy: **2.1 Reported vs. Actual Accuracy (SEVERE)** - **Paper claims:** 74.3% accuracy on OpenBookQA test set - **JSON results show:** 64.27% accuracy (259/403 correct) - **Discrepancy:** 10.03 percentage points lower than claimed - **Analysis:** The stored results file (`openbook_debate_transcript.json`) contains 403 debates, but only 259 were correct **2.2 Incomplete Dataset Coverage** - OpenBookQA validation set has 500 examples - JSON file only contains 403 debate transcripts - **Missing:** 97 examples (19.4% of the dataset) - **Possible explanations:** - API failures that were never retried - Manual dataset filtering not mentioned in paper - Code was interrupted before completion **2.3 No Hardcoded Results** - Accuracy is genuinely computed from judge responses - No evidence of cherry-picking specific seeds - Results appear to be authentic API responses - The discrepancy suggests honest errors rather than fabrication ## 3. IMPLEMENTATION-PAPER CONSISTENCY ### Major Inconsistencies: **3.1 Model Specifications** - **Paper claims:** "Gemini-2.5 as debaters, Gemini-2.0-lite as judge" - **Code shows:** - `STRONG_MODEL = "gemini-2.5-flash"` (not "Gemini-2.5") - `WEAK_MODEL = "gemini-2.0-flash-lite-001"` (close match) - **Issue:** Model version mismatch - "flash" variant vs. base model **3.2 Temperature Settings** - **Code implementation matches paper:** - Debaters: temperature=0.7 ✓ - Judge: temperature=0.3 ✓ **3.3 Dataset Split** - **Paper:** "OpenBookQA test set" - **Code:** Uses validation set (`openbook["validation"]`) - **Discrepancy:** Wrong dataset split used **3.4 Ablation Studies** - **Paper:** Comprehensive ablation study with 4 variants - **Code:** Only implements full debate protocol - **Missing:** 3 ablation variants not implemented ## 4. CODE QUALITY SIGNALS ### Strengths: - Clean, readable code structure - Proper exception handling with retry logic - Comprehensive transcript logging for reproducibility - Detailed prompt engineering visible in code - No dead or commented-out code (except CommonsenseQA dataset) - Appropriate use of external libraries (datasets, tqdm, google.genai) ### Weaknesses: - No modularization for ablation experiments - Empty project ID requires manual configuration - No documentation on API setup requirements - No validation of results against expected ranges - CommonsenseQA evaluation code is commented out but not removed ## 5. FUNCTIONALITY INDICATORS ### Working Components: - ✓ API integration with proper error handling and retries - ✓ Dataset loading from HuggingFace - ✓ Complete 5-phase debate protocol implementation - ✓ Answer extraction and matching logic - ✓ Accuracy computation - ✓ Full transcript storage in JSON ### Missing Components: - ✗ Ablation study variants (3 missing implementations) - ✗ Code to analyze uncertainty percentages - ✗ Code to analyze role-switch effectiveness - ✗ Statistical significance testing - ✗ Baseline comparison implementations ### Evidence of Real Execution: - 403 complete debate transcripts with full conversations - Realistic judge responses and reasoning - Natural variations in debate outcomes - Proper answer matching logic with fallbacks - This is genuine LLM output, not fabricated ## 6. DEPENDENCY & ENVIRONMENT ISSUES ### Required Dependencies: - `google.genai` (Google Gemini API client) - `datasets` (HuggingFace datasets library) - `tqdm` (progress bars) - Standard library: `json`, `time` ### Configuration Requirements: - Google Cloud project with Vertex AI enabled - Project ID must be manually configured (line 11) - API credentials must be set up - Location: "us-east1" ### Potential Issues: - No version pinning for dependencies - Requires paid Google Cloud API access - Rate limiting may cause failures (handled with retries) - Large token costs (~100-300 tokens per question × 9 API calls) ## 7. CRITICAL FINDINGS SUMMARY ### Severity: HIGH The codebase has THREE CRITICAL ISSUES that prevent verification of paper claims: 1. **Results Discrepancy (10% accuracy gap):** - Reported: 74.3% - Computed: 64.27% - This is a major inconsistency 2. **Missing Ablation Study Code:** - Paper claims 4 experimental variants - Only 1 variant implemented - Cannot verify claimed improvements from role-switching and uncertainty 3. **Dataset Split Confusion:** - Paper claims "test set" - Code uses "validation set" - Results may not be comparable to paper's claims ### Additional Issues: - 97 missing examples from the dataset (19.4% incomplete) - Model version discrepancy (flash vs. base) - No dependency specifications - Requires manual configuration ## 8. ASSESSMENT ### What Works: The core debate implementation is functional and appears to have been genuinely executed against the Gemini API. The stored transcripts contain realistic multi-turn conversations with proper role-switching and uncertainty quantification phases. ### What Doesn't Work: The paper's main scientific claims about the effectiveness of role-switching and uncertainty phases cannot be verified because: 1. No ablation study code exists 2. The reported accuracy doesn't match the stored results 3. The dataset split used differs from what's claimed ### Reproducibility Status: **Partially Reproducible** - The full debate protocol can be reproduced (with proper API setup), but: - The specific 74.3% accuracy cannot be reproduced from provided artifacts - Ablation study results cannot be reproduced at all - The discrepancy between reported and actual results raises serious concerns ### Possible Explanations: 1. **Honest Error:** Authors may have run additional experiments not included in the submission, leading to version mismatch 2. **Dataset Confusion:** Results from different runs or datasets may have been mixed up 3. **Incomplete Submission:** Full experimental code may exist but wasn't included in the submission 4. **Computational Issues:** 97 missing examples suggest the full run never completed ## 9. RECOMMENDATIONS To address these issues, the authors would need to: 1. Provide the actual code that generated the 74.3% accuracy result 2. Implement and share all 4 ablation study variants 3. Clarify which dataset split was actually used (test vs. validation) 4. Explain the 10% accuracy discrepancy 5. Provide complete results for all 500 questions 6. Add proper dependency specifications and setup instructions 7. Include the ablation study analysis code ## 10. AGENT REPRODUCIBILITY ASSESSMENT **Finding:** No evidence of AI-assisted code generation documentation The submission contains: - No prompt files - No AI interaction logs - No Claude/GPT/Copilot usage documentation - No comments indicating AI assistance The code appears to be human-written based on: - Consistent coding style - Thoughtful error handling - Domain-specific prompt engineering - No typical AI code generation artifacts **Conclusion:** AGENT REPRODUCIBILITY: False