Description
transform_projects() in transform.py discards the highlights returned by the LLM. Instead of reading item["highlights"], it reads a non-existent "type" key, so the expression always evaluates to an empty list:
# transform.py:339
"highlights": [item.get("type", "")] if item.get("type") else [],
"type" is not a field on the Project model in models.py, and it is never requested in prompts/templates/projects.jinja, so no resume can populate it.
The sibling transform_work() handles the same field correctly at transform.py:218:
"highlights": item.get("highlights", []),
Impact
Every bullet point under every project on a resume is silently dropped before evaluation.
The rest of the plumbing is already complete — Project.highlights exists in models.py, and convert_json_resume_to_text() already renders project.highlights at transform.py:812. This single assignment is the only break in the chain.
As a result the evaluator receives only project names and URLs, losing all project descriptions, metrics, impact statements, and any demo/deployment links written in bullet text. This under-scores the self_projects category (weight 30/100).
Evidence
The raw LLM response already contains the highlights, confirming this is purely a transform-layer loss (model gemini-2.5-flash, structured output):
{"name": "Multimodal Earnings Call Intelligence System",
"highlights": [
"Built a multimodal sentiment analysis pipeline for 1,038 earnings calls, fusing FinBERT text embeddings with prosodic audio features using a turn-level gating and attention pooling model.",
"Improved macro-F1 from 0.699 to 0.741 (+5.1 points) over a text-only baseline through cross-modal fusion, validated using cross-validation and significance testing."
],
"url": "https://github.com/.../earning-calls"}
But the text actually passed to the evaluator contains none of it:
=== PROJECTS ===
1. Multimodal Earnings Call Intelligence System
URL: https://github.com/.../earning-calls
2. Corporate Compliance Environment
URL: https://github.com/.../corporate-compliance-env
3. AI-Powered Geopolitical Risk Intelligence System
URL: https://github.com/.../forsyt
Steps to Reproduce
- Run
python score.py <resume.pdf> --role software_engineering_intern on a resume whose projects have bullet points.
- Inspect
cache/resumecache_<name>.json — every entry under projects has "highlights": [].
- Equivalently, print
convert_json_resume_to_text(resume_data) — the === PROJECTS === section contains only names and URLs.
Environment Info
- OS: macOS (Darwin 25.5.0)
- Python: 3.11
- Model: gemini-2.5-flash
- Commit: 70fd3ea
Description
transform_projects()intransform.pydiscards thehighlightsreturned by the LLM. Instead of readingitem["highlights"], it reads a non-existent"type"key, so the expression always evaluates to an empty list:"type"is not a field on theProjectmodel inmodels.py, and it is never requested inprompts/templates/projects.jinja, so no resume can populate it.The sibling
transform_work()handles the same field correctly attransform.py:218:Impact
Every bullet point under every project on a resume is silently dropped before evaluation.
The rest of the plumbing is already complete —
Project.highlightsexists inmodels.py, andconvert_json_resume_to_text()already rendersproject.highlightsattransform.py:812. This single assignment is the only break in the chain.As a result the evaluator receives only project names and URLs, losing all project descriptions, metrics, impact statements, and any demo/deployment links written in bullet text. This under-scores the
self_projectscategory (weight 30/100).Evidence
The raw LLM response already contains the highlights, confirming this is purely a transform-layer loss (model
gemini-2.5-flash, structured output):{"name": "Multimodal Earnings Call Intelligence System", "highlights": [ "Built a multimodal sentiment analysis pipeline for 1,038 earnings calls, fusing FinBERT text embeddings with prosodic audio features using a turn-level gating and attention pooling model.", "Improved macro-F1 from 0.699 to 0.741 (+5.1 points) over a text-only baseline through cross-modal fusion, validated using cross-validation and significance testing." ], "url": "https://github.com/.../earning-calls"}But the text actually passed to the evaluator contains none of it:
Steps to Reproduce
python score.py <resume.pdf> --role software_engineering_internon a resume whose projects have bullet points.cache/resumecache_<name>.json— every entry underprojectshas"highlights": [].convert_json_resume_to_text(resume_data)— the=== PROJECTS ===section contains only names and URLs.Environment Info