CS 4650 Final Project
We investigate how LLMs handle grade-school math word problems using the GSM8K benchmark. We compare a simple majority-class baseline against zero-shot and chain-of-thought (CoT) prompting, apply a lightweight verification step, and perform error analysis to categorize why models fail.
The main contribution is empirical: we develop an error taxonomy (arithmetic errors, multi-step reasoning failures, quantity misinterpretation) and analyze where different prompting strategies break down.
src/
evaluate.py # answer extraction + exact-match scoring
baselines.py # majority-class and random baselines
prompting.py # zero-shot and chain-of-thought prompting
verification.py # heuristic and re-prompt answer verification
fine_tuning.py # stub (out of scope)
notebooks/
data_exploration.ipynb
error_analysis.ipynb
results/ # prediction JSON files (generated by scripts)
python -m venv venv
source venv/bin/activate
pip install -r requirements.txtAll commands run from the project root.
Baseline:
python -m src.baselines --subset_size 100Prompting (mock mode — no GPU needed):
python -m src.prompting --mode mock --prompt_type zero_shot --subset_size 100
python -m src.prompting --mode mock --prompt_type cot --subset_size 100 --output results/cot_preds.jsonPrompting (real model):
python -m src.prompting --mode hf --prompt_type cot --subset_size 50 --model_name gpt2Verification:
python -m src.verification --input results/prompted_preds.json --strategy heuristic
python -m src.verification --input results/prompted_preds.json --strategy reprompt --mode mockEvaluation:
python -m src.evaluate --predictions results/baseline_preds.json
python -m src.evaluate --predictions results/prompted_preds.json- Full fine-tuning (compute-intensive, not feasible for our timeline)
- Symbolic/formal verification
- Multi-dataset evaluation
- Novel architectures
See requirements.txt. Just the basics: datasets, transformers, torch, matplotlib, pandas, jupyter.