# CODE AUDIT REPORT: Submission 269 ## TEAM-PHI: Multi-Agent PHI De-identification Framework **Date:** 2024 **Auditor:** Autonomous Code Auditing System --- ## EXECUTIVE SUMMARY This submission presents a multi-agent framework (TEAM-PHI) for evaluating PHI de-identification models. The code is **INCOMPLETE and NOT FULLY FUNCTIONAL** as submitted. While the overall architecture is reasonable and the methodology appears sound, the code contains critical missing components that prevent independent execution. **Overall Assessment:** HIGH severity issues prevent reproduction **Agent Reproducibility:** FALSE (no evidence of AI-generated code prompts) --- ## DETAILED FINDINGS ### 1. COMPLETENESS & STRUCTURAL INTEGRITY #### CRITICAL Issues: **1.1 Missing Data Files (CRITICAL)** - **Location:** All scripts reference `"folder_contain_clinical_notes"` or similar - **Issue:** No clinical notes data is provided in the submission - **Evidence:** - GPT_Deid.py line 27: `input_dir = "folder_contain_clinical_notes"` - Llama_Deid.py line 10: `input_dir = "folder_contain_clinical_notes"` - All evaluation scripts reference this directory - **Impact:** Cannot execute any part of the pipeline without data - **Justification:** The Reproducibility Statement acknowledges this: "Due to the sensitive nature of medical data and privacy regulations, these clinical notes are not publicly available" - **Severity:** CRITICAL for execution, but JUSTIFIED due to privacy constraints **1.2 Hardcoded Placeholder Credentials (CRITICAL)** - **Location:** - GPT_Deid.py lines 5, 13-14 - GPT_Eval.py lines 7, 12, 14 - GPT_vote.py lines 5, 13-14 - **Evidence:** ```python endpoint = "your_Azure_endpoint" subscription_key = "your_Azure_subscription_key" api_version = "your_api_version" ``` - GPT_vote.py has empty strings: `endpoint = ""` - **Impact:** GPT-based scripts will fail immediately on execution - **Severity:** CRITICAL - Code cannot run as written **1.3 Missing Intermediate Results (HIGH)** - **Location:** Evaluation agents expect processed De-id outputs - **Evidence:** - GPT_Eval.py line 23: `"De-id Agents' Final Output/gemma2_deid_outputs.json"` - No such directory or files exist in submission - **Impact:** Cannot run evaluation stage independently - **Severity:** HIGH - Breaks pipeline modularity **1.4 Missing Model Directories (HIGH)** - **Location:** All local model scripts - **Evidence:** - Llama_Deid.py line 6: `local_model_dir = "llama-3-8B-Instruct"` - LPPA_Deid.py line 6: `local_model_dir = "108mix4ktest1"` - Gemma_Deid.py line 9: `MODEL_DIR = "gemma-2-9b-it-mlx-q4"` - **Impact:** Scripts expect pre-downloaded models in specific directories - **Severity:** HIGH - Expected for large models but not documented **1.5 Hardcoded Results in Voting Scripts (HIGH)** - **Location:** All voting scripts - **Evidence:** - GPT_vote.py lines 30-32: `llm_results = """\n\n"""` - Llama_vote.py lines 30-32: `llm_results = """\ndeid agents and evaluation agents' output\n"""` - Gemma_Mistral_vote.py lines 16-18: `llm_results = r"""\n\n"""` - **Impact:** Voting scripts require manual insertion of results - **Severity:** HIGH - Not automated as described in paper **1.6 Missing Requirements File (MEDIUM)** - **Location:** Root directory - **Evidence:** No requirements.txt file found, despite README mentioning it - **README Quote:** "**['requirements.txt' file could be used to download the python packages automatically]**" - **Impact:** Dependency versions not specified (except in README) - **Severity:** MEDIUM - Versions are listed in README but not in standard format #### Positive Observations: ✓ **No TODO comments or placeholder functions** - All functions are fully implemented ✓ **No hardcoded results** in de-identification or evaluation outputs ✓ **Proper error handling** in Gemma/Mistral scripts with try-except blocks ✓ **Complete implementation** of all core algorithms (JSON parsing, normalization, evaluation) ✓ **Main entry points exist** for all components --- ### 2. RESULTS AUTHENTICITY #### NO RED FLAGS DETECTED: ✓ **No hardcoded experimental results** - All metrics computed programmatically ✓ **No cherry-picked seeds** - Temperature and sampling parameters are reasonable and consistent ✓ **No manual result insertion** - All outputs generated from model inference ✓ **Proper metric computation:** - Precision calculated as: `total_correct / total_pairs` - Coverage calculated as: `total_pairs / total_tokens` - Recall-proxy calculated using average across models - All metrics derived from actual evaluation, not hardcoded **Assessment:** Code demonstrates authentic computational workflow. Results would be genuinely computed if data were available. --- ### 3. IMPLEMENTATION-PAPER CONSISTENCY #### Strong Alignment: ✓ **Guideline prompts match paper description** (De-identification agents follow specified entity types) ✓ **Evaluation methodology matches paper** (JSON output format: `{"Number of Correct Pairs": N}`) ✓ **Multi-agent voting implemented as described** (independent and cross-informed modes) ✓ **Normalization rules comprehensive:** - Title removal for PERSON entities (lines 46-57 in Eval scripts) - Date format standardization (DATE_PAT regex in Gemma/Mistral Eval) - Deduplication logic properly implemented ✓ **Metrics match paper definitions:** - Precision, Coverage, Recall-Proxy all implemented correctly - Entity-specific metrics for PERSON and DATE_TIME categories #### Minor Inconsistencies: ⚠ **Temperature settings vary:** - De-id models use temperature 0.5-1.0 - Evaluation agents mostly use 0.0 (deterministic) - Voting uses 0.0 (Gemma/Mistral) or 1.0 (GPT/Llama) - **Assessment:** MINOR - Reasonable design choices, though not fully documented ⚠ **Max tokens vary by stage:** - De-id: 256-4096 tokens - Evaluation: 512-1024 tokens - Voting: 8-4096 tokens - **Assessment:** MINOR - Appropriate for different tasks --- ### 4. CODE QUALITY SIGNALS #### Positive Indicators: ✓ **Sophisticated JSON parsing** with multiple fallback strategies (Gemma/Mistral Deid) ✓ **Extensive normalization logic** showing deep understanding of PHI matching challenges ✓ **Consistent code structure** across similar agents (GPT/Llama variants follow same patterns) ✓ **Proper use of transformers library** with chat templates and generation parameters ✓ **No excessive code duplication** - Similar functionality appropriately parameterized ✓ **Clear separation of concerns** - De-id, Evaluation, and Voting are distinct modules #### Quality Concerns: ⚠ **Platform-specific implementations:** - Gemma/Mistral use `mlx_lm` (Apple Silicon specific) - Llama uses standard `transformers` - **Assessment:** MINOR - Appropriate for stated hardware (README mentions MacBook + H100) ⚠ **Limited error handling in some scripts:** - GPT/Llama de-id scripts have no try-except blocks - Could fail silently on malformed JSON - **Assessment:** LOW - Production code would need more robust error handling ⚠ **Inconsistent commenting:** - Some scripts well-commented (Gemma/Mistral with detailed docstrings) - Others minimal (GPT/Llama basic scripts) - **Assessment:** LOW - Stylistic, doesn't affect functionality --- ### 5. FUNCTIONALITY INDICATORS #### Evidence of Real Development: ✓ **Multiple implementation approaches** (MLX vs transformers) suggest iterative development ✓ **Sophisticated regex patterns** for date parsing and normalization indicate domain expertise ✓ **Detailed normalization rules** (title removal, punctuation stripping) show understanding of PHI matching challenges ✓ **Fallback parsing strategies** in Gemma/Mistral scripts demonstrate debugging experience ✓ **CSV output options** with argparse in advanced evaluation scripts ✓ **Sleep timers** to prevent API rate limiting (practical consideration) #### Pipeline Stages: 1. **De-identification:** ✓ Complete implementation 2. **Evaluation:** ✓ Complete implementation (requires De-id outputs) 3. **Voting:** ✓ Complete but requires manual result insertion 4. **End-to-end automation:** ✗ Missing - No master script to run full pipeline --- ### 6. DEPENDENCY & ENVIRONMENT ISSUES #### Dependencies (from README): ✓ Standard packages with reasonable versions: - python==3.8.10 - openai==0.28.1 (older version, pre-1.0 API) - transformers==4.33.3 - torch==1.10.0+cu111 ⚠ **Potential Issues:** - MLX framework (Apple-specific) not in README requirements - `mlx_lm` package used but not listed - Old OpenAI API version (0.28.1) - newer code uses new import style **Assessment:** MEDIUM - Code uses newer OpenAI import (`from openai import AzureOpenAI`) but README lists old version #### Computational Resources: ✓ **Realistic resource requirements:** Dual H100 GPUs mentioned for 70B models ✓ **Mixed hardware approach:** MacBook for smaller models, server for large ones ✓ **HIPAA-compliant infrastructure:** Azure for GPT models (appropriate for medical data) --- ### 7. ARCHITECTURAL ASSESSMENT #### Framework Design: ✓ **Well-structured three-stage pipeline:** 1. Multiple De-id models run in parallel 2. Multiple Evaluation agents assess each model 3. LLM voting aggregates judgments ✓ **Modular design:** Each component can theoretically run independently ✓ **Standardized I/O:** JSON format used consistently throughout ✓ **Sensible normalization:** Title removal, date standardization, deduplication #### Missing Components: ✗ **No orchestration script** to run full pipeline ✗ **No data preprocessing scripts** (clinical notes to JSON) ✗ **No results aggregation scripts** (collect evaluations into voting input) ✗ **No ground-truth comparison scripts** (despite paper claiming F1=0.62±0.01 for Llama-70B) --- ## SPECIFIC RED FLAGS BY SEVERITY ### CRITICAL (Prevents Execution): 1. **Placeholder credentials** in GPT scripts - Code will fail immediately 2. **Missing data directory** - No input to process (justified by privacy) 3. **Missing intermediate outputs** - Evaluation scripts expect non-existent files ### HIGH (Major Gaps): 4. **Manual result insertion required** for voting stage - Not automated as implied 5. **Missing ground-truth evaluation code** - Paper reports F1 scores but no code to compute them 6. **Missing model download instructions** - Expected directories not documented 7. **Inconsistent dependency specifications** - MLX not in requirements, old OpenAI version listed ### MEDIUM (Quality Issues): 8. **No requirements.txt file** - Despite being mentioned in README 9. **No master orchestration script** - Each stage must be run manually 10. **Limited documentation** on expected file formats and directory structure ### LOW (Minor Issues): 11. **Inconsistent error handling** across scripts 12. **Variable commenting quality** across files 13. **Platform-specific code** (MLX) without clear documentation --- ## VERIFICATION AGAINST PAPER CLAIMS ### Claims We Can Verify: ✓ **Multi-agent architecture:** Code shows 8 De-id models (GPT-3.5, GPT-4o, Llama-8B, Llama-70B, Mistral-7B, Gemma-2, LPPA4k, LPPA5k) ✓ **6 Evaluation agents:** Code shows GPT-3.5, GPT-4o-mini, Llama-70B, Llama-8B(?), Gemma-2, Mistral-7B ✓ **JSON output format:** Consistently enforced across all stages ✓ **Normalization strategies:** Title removal, date parsing implemented ✓ **Two voting modes:** Independent and cross-informed guidelines present ✓ **Metrics computation:** Precision, Coverage, Recall-Proxy calculated correctly ### Claims We Cannot Verify: ✗ **F1 scores (Table 8):** No ground-truth comparison code provided ✗ **Human evaluation (Table 9):** No evaluation interface or results ✗ **100 clinical notes dataset:** Not provided (justified) ✗ **Total compute cost (200 GPU-hours):** No logging or timing code ✗ **Specific performance numbers:** Cannot reproduce without data and complete pipeline --- ## AGENT REPRODUCIBILITY ASSESSMENT **Question:** Does the submission document AI usage with prompts that could regenerate the code? **Answer:** FALSE **Reasoning:** 1. ✗ No ChatGPT conversation logs found 2. ✗ No AI generation prompts documented 3. ✗ No comments indicating AI-assisted development 4. ✗ The Reproducibility Statement mentions "AI-generated results" referring to the *outputs of their LLM-based system*, not the code itself 5. ✓ Code quality and sophistication suggest human development with domain expertise 6. ✓ Multiple implementation styles (basic vs advanced) suggest iterative human refinement **Conclusion:** This code was written by humans, not generated by AI assistants. The submission does not provide prompts for regenerating the code. --- ## AUTHENTICITY ASSESSMENT ### Evidence Code is GENUINE (Not Fabricated): 1. **No hardcoded results** - All metrics computed from actual model outputs 2. **Realistic implementation complexity** - JSON parsing with multiple fallbacks shows real-world problem-solving 3. **Platform-specific optimizations** - MLX for Apple Silicon shows practical hardware considerations 4. **Incomplete but consistent** - Missing pieces are infrastructure (data, credentials) not algorithms 5. **Sophisticated domain knowledge** - PHI normalization rules show medical informatics expertise ### Evidence Results COULD Be Reproduced: 1. ✓ All metrics have proper computational logic 2. ✓ No cherry-picking mechanisms visible 3. ✓ Sampling parameters are reasonable 4. ✓ Evaluation methodology is sound 5. ✗ **BUT:** Cannot verify without data access **Assessment:** The code appears to represent genuine research work. Results would be computed authentically if the pipeline could run. --- ## RECOMMENDATIONS FOR REPRODUCIBILITY ### To Make Code Executable: 1. **Provide sample data** (synthetic or public dataset) to demonstrate pipeline 2. **Create master orchestration script** that runs all stages 3. **Add placeholder credentials detection** with informative error messages 4. **Include ground-truth evaluation scripts** to match Table 8 results 5. **Create requirements.txt** with all dependencies including MLX 6. **Document model download process** and expected directory structure 7. **Add data preprocessing scripts** showing clinical note → JSON conversion 8. **Include results aggregation scripts** for voting stage 9. **Provide example outputs** at each pipeline stage ### To Match Paper Claims: 10. **Include timing/logging code** to verify compute cost claims 11. **Add validation scripts** comparing automatic rankings to gold-standard 12. **Include human evaluation interface** or documentation 13. **Document hardware requirements** more clearly in code comments --- ## OVERALL ASSESSMENT ### Strengths: - ✓ Core algorithms are complete and well-implemented - ✓ No evidence of fabricated results - ✓ Methodology matches paper description - ✓ Code shows genuine domain expertise - ✓ Modular architecture is sound - ✓ Results would be authentic if pipeline could run ### Critical Weaknesses: - ✗ Cannot execute as submitted (missing data, credentials, models) - ✗ Missing orchestration and integration code - ✗ Ground-truth evaluation code absent - ✗ Manual intervention required for voting stage - ✗ Incomplete documentation ### Final Verdict: **CODE QUALITY:** Good - well-implemented algorithms with domain expertise **COMPLETENESS:** Incomplete - missing infrastructure and integration layers **EXECUTABILITY:** Non-functional - cannot run without data, credentials, and models **AUTHENTICITY:** High confidence - genuine research, not fabricated **REPRODUCIBILITY:** Low - major barriers prevent independent verification **AGENT REPRODUCIBLE:** FALSE - no AI generation prompts provided --- ## SEVERITY SUMMARY | Category | Count | Examples | |----------|-------|----------| | CRITICAL | 3 | Missing data, placeholder credentials, missing intermediate files | | HIGH | 4 | Manual voting, missing ground-truth code, model directories | | MEDIUM | 3 | No requirements.txt, dependency inconsistencies, no orchestration | | LOW | 3 | Error handling, documentation, platform-specific code | **Total Issues:** 13 (3 Critical, 4 High, 3 Medium, 3 Low) --- ## CONCLUSION This submission represents **genuine research with incomplete code infrastructure**. The core algorithms are well-implemented and show no signs of result fabrication. However, critical components required for independent execution are missing: 1. Data (justified by privacy constraints) 2. Credentials (expected to be user-provided) 3. Integration scripts (should have been included) 4. Ground-truth evaluation (claimed in paper but absent) The code could produce authentic results if completed, but **cannot currently verify paper claims** without substantial additional work. The submission earns credit for algorithmic completeness but loses points for lack of end-to-end reproducibility. **Recommendation:** Request authors provide (1) synthetic/public sample data, (2) complete orchestration scripts, and (3) ground-truth evaluation code to enable independent verification. --- **Report Generated:** 2024 **Audit System Version:** 1.0