# Audit Summary **CODEBASE AUDIT RESULT:** HIGH **AGENT REPRODUCIBILITY:** True --- # Detailed Code Audit Report - Submission 258 ## Executive Summary This submission presents code for an energy-guided program selection system that ranks LLM-generated code variants by energy consumption. The submission includes **only the final two phases** (Phase 4 reranking and Phase 5 evaluation) of what is clearly a multi-phase experimental pipeline. Critical components for actual reproducibility are **completely missing**, including: 1. **No code generation implementation** (Phase 1) 2. **No correctness filtering pipeline** (Phase 2) 3. **No energy measurement/execution code** (Phase 3) 4. **No actual candidate program files** While the submission documents extensive use of LLMs (ChatGPT 5, Gemini 2.5) in the research process and the analysis scripts are functional, the absence of core experimental infrastructure means the results cannot be independently reproduced. --- ## 1. COMPLETENESS & STRUCTURAL INTEGRITY ### Missing Critical Components (HIGH severity) **Phase 1-3 Implementation Absent:** - The submitted code contains only `phase4_reranker.py` (246 lines) and `phase5_evaluation.py` (250 lines) - Phase 3 is referenced throughout as the source of experimental data but **no implementation exists** - Code references in JSON output files point to `filtered_candidates-codellama:70b-instruct-v5/` directories that **do not exist** in the submission - The metadata CSV explicitly lists 17 candidate files per task (e.g., `candidate_0.py` through `candidate_19.py`) - **none are present** **Evidence from the code:** ```python # phase5_evaluation.py line 15-16 --phase3_dirs ./outputs/phase3 ./outputs/phase3_rerun \ --candidates_dir ./filtered_candidates-codellama:70b-instruct-v5 \ ``` These directories are required parameters but **do not exist** in the submission. **Missing Infrastructure:** 1. **Code generation system**: No implementation of the Code Llama-70B inference pipeline 2. **Correctness filtering**: No unit test framework or validation system 3. **Energy measurement harness**: No CodeCarbon integration or execution framework 4. **Candidate programs**: The 170+ generated Python programs (17 candidates × 10 tasks) are absent 5. **Task specifications**: No problem statements, test cases, or input generators **What IS Present:** - Pre-computed experimental results (256 JSON files with runtime/energy data) - Aggregated CSV summaries from Phase 3 execution - Phase 4 and 5 analysis scripts that **only process pre-existing data** ### Structural Integrity Assessment **Positive aspects:** - The two existing Python files are well-structured with clear documentation - No placeholder functions, TODOs, or hardcoded results in the analysis code - Proper error handling in the reranking script (lines 78-79, 86-88 in phase4_reranker.py) - All imports reference standard libraries (pandas, numpy, scipy, matplotlib) **Red flags:** - The experimental results exist **only as pre-computed data files** - No entry point to actually run the experiment from scratch - No README or setup instructions - Dependency typo: `seabron` instead of `seaborn` in requirements.txt (line 16) --- ## 2. RESULTS AUTHENTICITY RED FLAGS ### Evidence of Computation (MEDIUM severity) **Positive indicators:** - Results show realistic variance across runs (CV: 8.62% runtime, 10.75% energy) - 256 raw JSON files with individual run data demonstrate multiple executions - Statistical values are computed (not hardcoded) in the analysis scripts - Wilcoxon test p-values are calculated programmatically (phase5_evaluation.py lines 187-192) **Concerning aspects:** - All experimental results exist **only as pre-generated outputs** - The submission provides **no mechanism to verify** these results were actually computed - Candidate programs that supposedly produced these results are entirely absent - No git history or version control artifacts to verify development process **Results consistency check:** The numbers in the output files match the paper claims: - Overall energy savings vs Top-1: 44.69% (matches paper's 44.69%) - Matrix multiplication at scale 10⁶: 98.3% savings (matches paper's >98%) - Average runtime penalty: 0.0012s (matches paper's 0.0012s) - Median runtime penalty: 0.0s (matches paper's zero) This consistency is **both positive** (no cherry-picking of specific runs) **and concerning** (suggests the paper was written after finalizing these specific result files). ### Assessment The results appear to be **authentically computed** rather than fabricated, but the submission provides **no way to independently verify or reproduce** them. The analysis scripts correctly process the data, but without the execution infrastructure, we must trust the pre-computed outputs. --- ## 3. IMPLEMENTATION-PAPER CONSISTENCY ### Methodology Claims vs. Code **What the paper claims:** 1. "Uses Code Llama-Instruct (70B) model" - **No implementation present** 2. "Implements automated testing with lightweight unit tests" - **No test framework present** 3. "Employs CodeCarbon (v2.2.3)" - **No integration code present** 4. "Five repeated runs per candidate" - **Results files confirm this** 5. "Single-threaded execution (OPENBLAS_NUM_THREADS=1...)" - **Cannot verify** 6. "Tested across three input scales (10⁴, 10⁵, 10⁶)" - **Results files confirm this** **Verifiable consistency:** - The CSV data structure matches the described experimental protocol - 30 combinations analyzed (10 tasks × 3 scales) - correct - Correctness filtering applied (correct_fraction field in data) - Statistical tests match methodology (Wilcoxon signed-rank) **Cannot verify:** - Whether CodeCarbon was actually used (no implementation to inspect) - Whether candidates were truly LLM-generated (no prompts, no generator code) - Whether the filtering/execution protocol matches claims - Hyperparameters, sampling strategies, or model configurations ### Hyperparameters The paper mentions "temperature and nucleus sampling" for generation diversity, but: - No configuration files exist - No code shows these parameters - Cannot verify if results match the claimed experimental conditions --- ## 4. CODE QUALITY SIGNALS ### Analysis Scripts Quality (phases 4 & 5) **Strengths:** - Clean, well-documented code with docstrings - Proper pandas/numpy usage for data analysis - No dead code or excessive comments - Logical structure and clear variable naming - Appropriate use of pathlib for file handling - Proper command-line argument parsing **Example of quality code:** ```python def calculate_metrics(df, group_name="Overall"): """Calculates comparison metrics for a given dataframe slice.""" num_cases = len(df) if num_cases == 0: return None # Safe division handling with NaN replacement energy_savings_vs_top1 = ( (df['energy_median_kwh_top1'] - df['energy_median_kwh_energy']) / df['energy_median_kwh_top1'].replace(0, np.nan) ).mean() ``` **Minor issues:** - Some magic numbers without constants (e.g., 0.1-second sampling in paper not in code) - Limited input validation beyond file existence checks - Typo in requirements.txt (`seabron` instead of `seaborn`) **Code statistics:** - Total Python LOC: 496 lines (minimal - only 2 analysis scripts) - No commented-out code blocks - No unused imports - Appropriate ratio of comments to code ### Missing Code Quality Concerns Since 60-80% of the experimental pipeline is missing, we cannot assess: - Quality of the code generation prompts - Robustness of the energy measurement harness - Reliability of the correctness testing - Error handling in execution framework --- ## 5. FUNCTIONALITY INDICATORS ### What Works **Phase 4 Reranking (`phase4_reranker.py`):** - ✅ Reads aggregated CSV data correctly - ✅ Filters for correct candidates (correct_fraction == 1.0) - ✅ Applies three selection strategies (Energy-Guided, Top-1, Best-Time) - ✅ Computes relative savings and penalties - ✅ Generates multiple output formats (detailed, summary, per-task) - ✅ Handles edge cases (division by zero, empty dataframes) **Phase 5 Evaluation (`phase5_evaluation.py`):** - ✅ Loads and processes selection data - ✅ Identifies divergent cases between strategies - ✅ Computes aggregate statistics correctly - ✅ Generates visualizations (histograms, boxplots) - ✅ Performs Wilcoxon signed-rank tests - ✅ Analyzes measurement consistency (CV calculation) - ✅ Assesses ranking stability across scales ### What's Missing **Cannot verify actual experimental functionality:** - ❌ No way to generate new candidates - ❌ No way to test candidate correctness - ❌ No way to measure energy consumption - ❌ No way to reproduce the core experimental loop **The submission can:** - Re-analyze existing results with different groupings - Generate new visualizations from existing data - Compute alternative statistical tests on existing data **The submission cannot:** - Generate new experimental data - Validate the correctness of the existing data - Reproduce the experiment on new tasks - Verify the energy measurements were accurate --- ## 6. DEPENDENCY & ENVIRONMENT ISSUES ### Requirements Analysis **requirements.txt contents:** ``` numpy==1.26.4 codecarbon==2.2.3 # Listed but not used in submitted code psutil>=5.9.0,<6.0.0 # Listed but not used in submitted code rich>=13,<14 # Listed but not used in submitted code platformdirs>=4,<5 # Listed but not used in submitted code pydantic>=2.6,<3 # Listed but not used in submitted code requests>=2.31,<3.0 # For Ollama API - not used in submitted code pandas # ✅ Used in phase 4 & 5 ollama # Listed but not used in submitted code matplotlib # ✅ Used in phase 5 scipy # ✅ Used in phase 5 seabron # ❌ TYPO: should be "seaborn" ``` **Issues:** 1. **Typo:** `seabron` instead of `seaborn` - this would cause import failures 2. **Unused dependencies:** Most packages are for missing Phase 1-3 implementations 3. **No version pinning:** pandas, matplotlib, scipy lack version constraints 4. **Missing test:** No way to verify the environment works without Phase 1-3 code **Positive aspects:** - All packages are standard, well-maintained libraries - No conflicting dependencies - Reasonable resource requirements for analysis scripts - No exotic or deprecated packages --- ## 7. AGENT REPRODUCIBILITY ### LLM Usage Documentation The submission includes **exemplary documentation** of LLM usage in `Conversations with LLMs.txt`: **Documented conversations:** 1. Hypothesis and study plan generation (GPT-5) 2. Support prompts (Gemini 2.5) 3. State of the art review (GPT-5) 4. Implementation planning (GPT-5) 5. Phase 2 implementation (Gemini 2.5 Pro) 6. Phase 2 & 3 implementation (GPT-5) 7. **Phase 3 rerun, Phase 4, and Phase 5 implementation (Gemini 2.5 Pro)** 8. Study diagram generation (Gemini 2.5 Pro) 9. Results/Discussion/Conclusion drafting (GPT-5) 10. Content refinement (Gemini 2.5 Pro, multiple rounds) **Assessment:** - ✅ **AGENT REPRODUCIBILITY: True** - The researchers transparently documented all LLM interactions - Links to 11 specific ChatGPT/Gemini conversation threads provided - Clear attribution of which AI generated which components - Covers hypothesis formation, implementation, and writing **However:** - The actual prompts and full conversations are external (on ChatGPT/Gemini platforms) - No guarantee the links will remain accessible long-term - Cannot verify the generated code matches what's claimed without accessing these conversations - The **actual generated candidate programs** (the core experimental artifact) are still missing --- ## 8. REPRODUCIBILITY ASSESSMENT ### What CAN be reproduced With the submitted code, a researcher can: 1. ✅ Re-run the Phase 4 reranking on the provided Phase 3 outputs 2. ✅ Re-run the Phase 5 evaluation and generate new visualizations 3. ✅ Verify the statistical calculations match the paper 4. ✅ Compute alternative metrics on the existing data 5. ✅ Understand the analysis methodology ### What CANNOT be reproduced Without the missing components, a researcher **cannot**: 1. ❌ Generate new program candidates using Code Llama-70B 2. ❌ Run correctness filtering on candidates 3. ❌ Measure energy consumption of programs 4. ❌ Verify the pre-computed results are accurate 5. ❌ Apply the method to new programming tasks 6. ❌ Reproduce the experiment independently ### Reproducibility Score: **15%** - **15%** - Can reproduce the analysis of pre-computed results - **0%** - Cannot reproduce the actual experiment - **0%** - Cannot verify result authenticity - **0%** - Cannot extend to new tasks --- ## 9. SPECIFIC RED FLAGS SUMMARY ### Critical Issues (HIGH severity) 1. **Missing Core Implementation (60-80% of pipeline)** - No code generation (Phase 1) - No correctness testing (Phase 2) - No energy measurement (Phase 3) - Only post-processing analysis provided 2. **Absent Experimental Artifacts** - 170+ candidate programs referenced but not included - No task specifications or test suites - No execution harness or measurement infrastructure 3. **Unverifiable Results** - All data pre-computed with no way to regenerate - Cannot validate measurement accuracy - Must trust the provided CSV/JSON files ### Medium Issues 4. **Dependency Typo** - `seabron` instead of `seaborn` would cause import errors - Suggests code wasn't tested after requirements.txt was written 5. **Incomplete Requirements** - Many dependencies listed for missing code - No version constraints on critical packages (pandas, matplotlib) ### Low Issues 6. **Documentation Gaps** - No README explaining how to use the submitted code - No setup instructions - LLM conversation links may become inaccessible --- ## 10. POSITIVE ASPECTS Despite the major gaps, the submission has strengths: 1. **Transparent LLM Documentation** - Exemplary disclosure of AI usage throughout research process - Links to 11 specific conversations covering all phases - Clear attribution of AI-generated vs. human work 2. **Quality Analysis Code** - Well-written, documented, and structured Phase 4/5 scripts - Proper statistical methodology (Wilcoxon tests) - Good data processing practices (NaN handling, error checks) 3. **Comprehensive Output Data** - 256 raw JSON files with individual run results - Multiple aggregation levels (overall, per-task, per-scale) - Visualization outputs demonstrate results 4. **Results Consistency** - Numbers match paper claims exactly - Statistical values are computed, not hardcoded - Realistic variance in measurements 5. **No Evidence of Fabrication** - Results show expected patterns (variance, scale effects) - No suspicious cherry-picking of specific runs - Proper statistical testing applied --- ## 11. RECOMMENDATIONS ### For Acceptance **Cannot recommend acceptance** in current form because: - Core experimental pipeline (60-80%) is completely missing - Results cannot be independently verified or reproduced - No way to validate the energy measurements were accurate - Cannot extend or apply the method without Phase 1-3 code ### For Authors to Address To make this work reproducible, the authors must provide: 1. **Critical (Required):** - Complete Phase 1-3 implementation code - All 170+ candidate program files - Task specifications and test suites - Execution and measurement harness - Step-by-step reproduction instructions 2. **Important:** - Fix typo in requirements.txt (`seabron` → `seaborn`) - Add version constraints to all dependencies - Include README with setup/usage instructions - Archive LLM conversation content locally (links may break) 3. **Nice to have:** - Docker container for reproducible environment - Scripts to regenerate all results from scratch - Validation that candidate programs match LLM outputs --- ## 12. CONCLUSION This submission presents **analysis code for a legitimate experiment**, but **lacks the core infrastructure** needed for reproducibility. The Phase 4 and 5 analysis scripts are well-written and correctly process the provided data, matching the paper's reported results. However, the absence of: - Code generation implementation - Correctness testing framework - Energy measurement system - The actual candidate programs ...means the work **cannot be independently verified or reproduced**. The extensive documentation of LLM usage is commendable and warrants the **AGENT REPRODUCIBILITY: True** designation, but this alone does not compensate for the missing experimental infrastructure. **Severity: HIGH** - Major implementation gaps prevent independent verification of the core scientific claims. The submission provides post-hoc analysis of pre-computed results rather than a reproducible experimental pipeline. --- ## Appendix: File Inventory ### Provided Files - `phase4_reranking/phase4_reranker.py` (246 lines) - `phase5_evaluation/phase5_evaluation.py` (250 lines) - `requirements.txt` (16 lines, with typo) - `Conversations with LLMs.txt` (35 lines) - `258_methods_results.md` (46 lines) - 256 JSON result files from Phase 3 execution - 278 CSV aggregation/analysis files - 2 PNG visualization files ### Missing Files - Phase 1: Code generation implementation (LLM inference, prompt engineering) - Phase 2: Correctness filtering (test framework, validation) - Phase 3: Energy measurement harness (CodeCarbon integration, execution loop) - ~170 candidate program files (candidate_0.py through candidate_19.py for 10 tasks) - Task specifications and problem statements - Unit test suites for correctness validation - Input data generators - README or documentation **Total submitted code: 496 lines Python** **Estimated missing code: 2,000-5,000 lines** (based on pipeline complexity)