# Code Audit Report: sub_287 **Date:** 2025-10-14 **Auditor:** Claude Code --- ## Executive Summary This submission presents a multi-agent pharmacovigilance system called "Echo" that claims to extract drug-symptom associations from Reddit discussions. However, the code exhibits **critical structural deficiencies** that make it impossible to reproduce the paper's reported results. The implementation consists of disconnected prototype scripts with hardcoded outputs, incomplete core functionality, missing LLM integration, no actual data, and fundamental inconsistencies with the paper's methodology. --- ## Risk Level: CRITICAL --- ## Key Red Flags 1. **Hardcoded Proposer Report - Results Not Computed** - **Evidence:** echo.py:525-533 - The `show_proposer_report()` function displays a generic hardcoded template regardless of input - **Impact:** The paper claims "Proposer generates mechanistic hypotheses by synthesizing biomedical literature" with specific examples for each drug-symptom pair, but the code shows this is impossible - all reports display identical placeholder text. This is a smoking gun for fabricated results. 2. **No LLM Integration - Core Claims Unverifiable** - **Evidence:** explorer_agent.py:158-218 - The `simple_extraction()` function uses basic regex pattern matching, not Claude 3.5 Sonnet/Haiku as claimed in the paper - **Impact:** Paper reports comparing "Claude 3.5 Haiku" vs "Claude 3.5 Sonnet" extractors with different recall scores (0.08 vs 0.27), but the code contains no API calls, no model selection logic, and no LLM integration whatsoever. The claimed 640 associations extracted by Claude 3.5 Sonnet cannot be reproduced. 3. **Incomplete Verifier Agent - Missing Core Methods** - **Evidence:** verifier_agent.py:68-69 - File ends abruptly with incomplete `FDAAdverseReactionExtractor` class; bare `except: return` with no error handling - **Impact:** Paper claims "Verifier computes ground-truth score: ln(FAERScount + 1)" and cross-references FDA databases, but the class is missing critical methods like `get_product_name_variants()`, `get_faers_term_counts()`, and `_filter_by_year()` that are called but never defined (lines 11-12, 12-17, 64). 4. **Empty Reddit API Credentials** - **Evidence:** explorer_agent.py:46-47 - `client_id=""` and `client_secret=""` are empty strings - **Impact:** The Explorer agent cannot connect to Reddit. Without credentials, the system cannot collect any data from Reddit, making the entire pipeline non-functional from the start. 5. **No Actual Data or Results Files** - **Evidence:** Directory listing shows only 5 Python files with no JSON outputs, no processed data, no FAERS database files, no Reddit scraping results - **Impact:** Paper reports analyzing 187 Reddit posts and generating 640 associations with specific metrics, but there is zero evidence this pipeline was ever executed. No intermediate data files exist. 6. **Proposer Agent Mislabeled - Does Not Match Paper Description** - **Evidence:** proposer_agent.py:1-43 - The file analyzes confounders, not generating mechanistic hypotheses; no literature synthesis, no biomedical knowledge integration - **Impact:** Paper claims "Proposer generates mechanistic hypotheses for flagged associations by synthesizing biomedical literature" with specific examples like "immune-mediated disruption of hypothalamic sleep-wake circuits," but the code only counts confounders. The actual Proposer functionality is entirely missing. 7. **Novelty Score Calculation Contradicts Paper** - **Evidence:** echo.py:299 - `novelty_score = (avg_temporal + avg_confidence + avg_community) / 3` - **Impact:** Paper states "Verifier score = 0 indicates no recorded FAERS instances" as the novelty metric, but code computes novelty as average of three Analyzer metrics (temporal/confidence/community). These are completely different definitions and would yield entirely different rankings. 8. **Analyzer Agent is Just a Data Aggregator** - **Evidence:** analyzer_agent.py:9-72 - Only aggregates pre-existing JSON files; no metric calculation, no temporal analysis, no confidence scoring, no community metric computation - **Impact:** Paper claims Analyzer "quantifies association strength with three metrics per evidence instance" but the code simply reads these values from files. There is no implementation of how temporal weight (0-1), confidence weight (0-1), or community support (0-1) are actually computed. 9. **Explorer Uses Hardcoded Symptom List and Fixed Scores** - **Evidence:** explorer_agent.py:181-198 - Searches for fixed symptom keywords, assigns hardcoded `temporal_weight=0.5` and `confidence=0.6` to all extractions - **Impact:** Paper reports precise varied metrics (e.g., "temporal 1.0, confidence 1.0, community 0.8" for specific associations), but the extraction code assigns identical fixed values to everything. The reported granular differences are fabricated. 10. **No Integration Between Agents** - **Evidence:** All four agent files are standalone scripts with no imports between them; echo.py UI is completely disconnected from agent implementations - **Impact:** The paper describes a coordinated "multi-agent system" where agents work together, but there's no orchestration code, no data flow between agents, and no way to execute the full pipeline. Each script would need to be run manually with correct file paths. --- ## Confidence Assessment **Can this code reproduce the paper's results?** No **Reasoning:** - **Completeness:** Critical components are missing (full Verifier implementation, actual Proposer logic, LLM integration, API credentials) or incomplete (truncated classes, undefined methods) - **Critical Functionality:** Core claims are impossible - Proposer shows hardcoded text, Explorer lacks LLM calls, metrics are either hardcoded or read from non-existent files, novelty scoring contradicts paper definition - **Consistency with Paper:** Fundamental mismatches - the code cannot generate the specific results shown in the paper (640 associations, varied metrics, mechanistic hypotheses, FAERS-based novelty scores) - **Code Quality:** This appears to be early-stage prototype code with placeholder implementations, not the system that generated the paper's results --- ## Detailed Findings ### Completeness & Structural Integrity **CRITICAL ISSUES:** 1. **Verifier Agent Incomplete (verifier_agent.py:1-69)** - The `FDAAdverseReactionExtractor` class is truncated at line 68-69 with a bare `except: return` statement - Methods `get_product_name_variants()`, `get_faers_term_counts()`, and `_filter_by_year()` are called (lines 11-12, 12-17, 64) but never defined - The `search_drug_labeling()` method has incomplete error handling that returns `None` implicitly - Functions `get_variants_and_terms()` and `get_side_effect_score()` depend on a fully-formed `extractor` object that doesn't exist 2. **No Entry Point for Complete Pipeline** - No main script that orchestrates all four agents - Each agent file has its own `if __name__ == "__main__"` block, but they don't call each other - No configuration file specifying how agents should be connected - README mentions `src/` directory structure, but files are in root directory 3. **Missing Dependencies** - No `requirements.txt`, `setup.py`, or any dependency specification - Uses external libraries: `praw`, `tkinter`, `pandas`, `requests` without version constraints - No indication of Python version requirements 4. **No Data Files** - The paper reports analyzing 187 Reddit posts and generating 640 associations - Zero JSON files, zero CSV files, zero intermediate data products in the repository - `OUTPUT_DIR = "reddit_data"` referenced in explorer_agent.py:41 doesn't exist - Analyzer expects an `--aggregate_folder` with drug JSON files that don't exist ### Results Authenticity RED FLAGS **CRITICAL ISSUES:** 1. **Hardcoded Proposer Reports (echo.py:525-533)** ```python dummy_report = f""" DRUG-INDUCED SYMPTOM ANALYSIS PRIMARY HYPOTHESIS: NEUROINFLAMMATORY CASCADE The compound may accumulate in brain regions... """ ``` - The function completely ignores the input `row_data` parameter - Every drug-symptom pair would show identical generic text - Paper claims specific mechanistic hypotheses for each association (e.g., "Pembrolizumab → Daytime somnolence: immune-mediated disruption of hypothalamic sleep-wake circuits"), but the code makes this impossible 2. **Fixed Metric Values in Explorer (explorer_agent.py:192-197)** ```python "temporal_weight": 0.5, "age": 20, "severity": "not specified", "confidence": 0.6, ``` - All extractions receive identical hardcoded scores - Paper reports varied metrics: "temporal 1.0, confidence 1.0" vs "temporal 0.8, confidence 0.8" - No mechanism exists to produce the granular variation reported in results 3. **Novelty Score Redefinition** - Paper definition: "Verifier score = 0 indicates no recorded FAERS instances" (novelty when score = 0) - Code definition (echo.py:299): `novelty_score = (avg_temporal + avg_confidence + avg_community) / 3` - These would produce completely different rankings and novel signals - The specific examples in the paper cannot be generated by this formula 4. **No Evidence of Actual Execution** - No output files, no logs, no cached results - Code expects input files that don't exist (aggregated drug JSONs for Analyzer) - The 640 associations reported in the paper exist nowhere in the repository ### Implementation-Paper Consistency **MAJOR INCONSISTENCIES:** 1. **Explorer Agent - LLM Claims vs. Implementation** - **Paper:** "Implementations compared: keyword-based baseline, Claude 3.5 Haiku, Claude 3.5 Sonnet" with different recall scores (0.22, 0.08, 0.27) - **Code:** Only simple regex pattern matching (explorer_agent.py:167-198), no LLM API calls, no model selection - **Impact:** The core comparison in the paper's results cannot be reproduced 2. **Analyzer Agent - Metric Computation** - **Paper:** "Quantifies association strength with three metrics per evidence instance: Temporal weight (0–1), Confidence weight (0–1), Community support (0–1)" - **Code:** Simply aggregates these values from pre-existing JSON files (analyzer_agent.py:39-45): `extraction.get('temporal_weight')`, `extraction.get('confidence')`, `extraction.get('community_metric')` - **Missing:** No implementation of how temporal proximity, certainty language, or upvotes/replies are actually scored 3. **Verifier Agent - FAERS Integration** - **Paper:** "Cross-references associations with FDA sources (e.g., FAERS, labels)" and "Computes a ground-truth score: ln(FAERScount + 1)" - **Code:** Incomplete implementation; the score calculation exists (verifier_agent.py:28) but critical supporting methods are missing - **Missing:** The retrospective analysis requiring "pre-2017 labels/FAERS" filtering logic is incomplete 4. **Proposer Agent - Completely Different Functionality** - **Paper:** "Generates mechanistic hypotheses for flagged associations by synthesizing biomedical literature" - **Code:** Counts confounders across drug-symptom pairs (proposer_agent.py:1-43), no literature access, no hypothesis generation - **Impact:** The agent that should produce the mechanistic insights (a key innovation claim) doesn't exist 5. **Community Metric Calculation** - **Paper:** "Community support (0–1): forum engagement derived from upvotes and supportive replies" - **Code (explorer_agent.py:174):** `community_metric = post_data.get("score", 0) + post_data.get("num_comments", 0)` - raw sum, not normalized to 0-1 range - Posts could have hundreds of upvotes/comments, yielding values >> 1.0 6. **Confounder Identification** - **Paper:** "Identifies potential confounders from text (e.g., comorbidities, demographics, concurrent treatments)" - **Code:** Analyzer simply reads `extraction.get('confounders', [])` but Explorer never populates this field - **Missing:** No NLP logic to extract comorbidities, no demographic parsing, no concurrent treatment identification ### Code Quality **SIGNIFICANT ISSUES:** 1. **Bare Exception Handling (verifier_agent.py:68-69)** ```python except: return ``` - Catches all exceptions silently, returns None - Impossible to debug when API calls fail - Violates basic Python best practices 2. **Unused Imports** - `import pandas as pd` in echo.py:6 - pandas is never used in the file - This suggests copy-paste from a different codebase 3. **Incomplete File Structure** - The verifier_agent.py file ends at line 69 mid-class definition - Missing methods that are called: `_filter_by_year()`, `get_product_name_variants()`, `get_faers_term_counts()` - Appears to be a fragment of a larger file 4. **Hardcoded Visual Elements** - The UI (echo.py) is well-polished with detailed styling and color schemes (lines 17-85) - Suggests significant effort went into appearance while core functionality remains incomplete - Classic "demo-ware" pattern prioritizing visuals over substance 5. **No Error Handling in Critical Paths** - Explorer's `main()` function (explorer_agent.py:220-263) catches exceptions per subreddit but continues silently - Analyzer processes files without validating JSON structure or required fields - Could fail silently and produce partial/incorrect results 6. **Code Organization Mismatch** - README describes `src/` directory structure with files in `src/explorer_agent.py`, etc. - Actual files are in root directory - Suggests README was written for a different version of the code ### Functionality **CRITICAL GAPS:** 1. **Reddit Data Collection Impossible** - Empty API credentials (explorer_agent.py:46-47) mean Reddit connection will fail immediately - No fallback mechanism, no instruction on how to obtain credentials - The paper claims "187 Reddit posts (identified via SerpApi)" but there's no SerpApi integration in the code 2. **No LLM Processing** - Paper's core innovation is using Claude 3.5 models for extraction - Zero Anthropic API calls, no OpenAI integration, no LLM invocation anywhere - The `simple_extraction()` function is the only extraction logic and it's purely rule-based 3. **Analyzer Requires Pre-Processed Data** - Expects aggregated JSON files organized by drug name - No code to create this structure from Explorer output - Explorer outputs per-post JSONs (explorer_agent.py:200-212), not per-drug aggregation 4. **No Hypothesis Generation** - The Proposer that should generate mechanistic hypotheses is actually a confounder counter - The actual hypothesis text shown in UI is hardcoded dummy text - Paper's specific examples (e.g., "Paclitaxel → Laryngeal edema: Cremophor EL–mediated hypersensitivity reactions") cannot be generated 5. **UI Disconnected from Pipeline** - echo.py UI requires loading a JSON file (echo.py:276-284) - No script generates this JSON in the required format - Expected structure (drug → symptom → entries) is not produced by any agent 6. **No Verification Against FAERS** - While Verifier has FDA API endpoint URLs (verifier_agent.py:48-49), the incomplete implementation means no actual verification - The retrospective validation claims (pre-2017 analysis) cannot be performed without the missing filtering logic ### Dependencies & Environment **ISSUES:** 1. **No Dependency Specification** - No `requirements.txt` file - No version pinning for any library - Unclear which Python version was used 2. **External API Dependencies** - Reddit API (PRAW) requires credentials - not provided - FDA API endpoints - no API key management (may not require keys but should be documented) - Anthropic Claude API - never used despite being central to paper claims - SerpApi - mentioned in paper but not in code 3. **Standard Library Conflicts** - `tkinter` may not be available in all Python installations (especially server/cloud environments) - No check for tkinter availability before UI launch 4. **Reproducibility Barriers** - Even with correct dependencies, missing credentials and incomplete code make reproduction impossible - No documentation on required environment setup - No example data or test cases to verify correct installation --- ## Recommended Actions 1. **Request Complete Codebase** - The submitted code appears to be a fragment or demo version - Ask authors for the actual implementation used to generate paper results - Request version control history to verify development timeline 2. **Request Processed Data Files** - Ask for the 187 Reddit posts (raw JSON) that were analyzed - Request intermediate outputs from each agent (Explorer extractions, Analyzer metrics, Verifier scores) - Request the final 640 drug-symptom associations with all computed metrics 3. **Clarify LLM Integration** - How were Claude 3.5 Sonnet/Haiku actually used? What were the prompts? - Request API call logs or at minimum the prompt templates - Explain discrepancy between paper's LLM usage and code's regex patterns 4. **Request Complete Verifier Implementation** - The submitted verifier_agent.py is incomplete - Need full implementation of FAERS querying and scoring logic - Need pre-2017 filtering implementation for retrospective validation 5. **Request Actual Proposer Agent** - Current proposer_agent.py does not match paper description - Need the code that generates mechanistic hypotheses from literature - Request bibliography/literature database used for synthesis 6. **Explain Novelty Score Calculation** - Resolve contradiction between paper definition (FAERS-based) and code definition (metric average) - Request the actual scoring method used for paper results 7. **Request Data Pipeline Documentation** - How do the four agents actually connect? - What is the execution order and data flow? - Request the orchestration script that runs the full pipeline 8. **Independent Verification** - Request access to the interactive UI with actual results loaded - Request screenshots or video demonstration of the system in operation - Consider requesting raw data for independent re-analysis --- ## Files Reviewed - `echo.py` (557 lines) - Interactive UI with hardcoded Proposer reports - `explorer_agent.py` (266 lines) - Reddit scraper with regex extraction, no LLM integration - `analyzer_agent.py` (105 lines) - Simple data aggregator, no metric computation - `verifier_agent.py` (69 lines, incomplete) - Truncated FDA verification agent - `proposer_agent.py` (43 lines) - Confounder counter, not hypothesis generator - `README.md` - Documentation with structural mismatches - `287_methods_results.md` - Paper methods and results summary --- ## Final Assessment This code submission cannot reproduce the paper's results. The implementation exhibits multiple **CRITICAL** red flags: 1. **Hardcoded outputs** (Proposer reports) masquerading as computed results 2. **Missing core functionality** (LLM integration, actual hypothesis generation, complete Verifier) 3. **Fundamental inconsistencies** between paper claims and code implementation 4. **No actual data** despite claims of analyzing 187 posts and generating 640 associations 5. **Incomplete files** (truncated Verifier class, missing methods) The code appears to be either: - An early prototype that was abandoned before paper completion - A post-hoc reconstruction attempting to approximate the described system - A demonstration UI built on top of manually generated results The gap between the sophisticated system described in the paper and the incomplete placeholder code submitted is too large to be explained by minor omissions or documentation issues. **Substantial additional evidence is required** to verify the paper's claims. **Recommendation:** Request complete working code, intermediate data files, and independent verification before accepting paper.