Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

9 Commits

Folders and files

Repository files navigation

Diagnosing and Improving Numerical Reasoning in LLMs

CS 4650 Final Project

Overview

We investigate how LLMs handle grade-school math word problems using the GSM8K benchmark. We compare a simple majority-class baseline against zero-shot and chain-of-thought (CoT) prompting, apply a lightweight verification step, and perform error analysis to categorize why models fail.

The main contribution is empirical: we develop an error taxonomy (arithmetic errors, multi-step reasoning failures, quantity misinterpretation) and analyze where different prompting strategies break down.

Repo Structure

src/
  evaluate.py        # answer extraction + exact-match scoring
  baselines.py       # majority-class and random baselines
  prompting.py       # zero-shot and chain-of-thought prompting
  verification.py    # heuristic and re-prompt answer verification
  fine_tuning.py     # stub (out of scope)
notebooks/
  data_exploration.ipynb
  error_analysis.ipynb
results/             # prediction JSON files (generated by scripts)

Setup

python -m venv venv
source venv/bin/activate
pip install -r requirements.txt

Running Experiments

All commands run from the project root.

Baseline:

python -m src.baselines --subset_size 100

Prompting (mock mode — no GPU needed):

python -m src.prompting --mode mock --prompt_type zero_shot --subset_size 100
python -m src.prompting --mode mock --prompt_type cot --subset_size 100 --output results/cot_preds.json

Prompting (real model):

python -m src.prompting --mode hf --prompt_type cot --subset_size 50 --model_name gpt2

Verification:

python -m src.verification --input results/prompted_preds.json --strategy heuristic
python -m src.verification --input results/prompted_preds.json --strategy reprompt --mode mock

Evaluation:

python -m src.evaluate --predictions results/baseline_preds.json
python -m src.evaluate --predictions results/prompted_preds.json

What's Intentionally Out of Scope

  • Full fine-tuning (compute-intensive, not feasible for our timeline)
  • Symbolic/formal verification
  • Multi-dataset evaluation
  • Novel architectures

Dependencies

See requirements.txt. Just the basics: datasets, transformers, torch, matplotlib, pandas, jupyter.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages