Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

persona-eval-annotation

An annotation platform for evaluating RAG-generated summaries through pairwise tournament comparisons. Annotators log in with unique tokens, answer background questions, submit queries, and rank four summaries per query in a bracket-style tournament (A vs B, C vs D, then winners face off).

Setup

pip install -r requirements.txt

Environment Variables

Required

Variable Description
TOGETHER_API_KEY API key for TogetherAI. Used for summary generation.
DATABASE_URL PostgreSQL connection string. Provided automatically by Railway's Postgres plugin. For local development see below.

Optional

Variable Default Description
SEMANTIC_SCHOLAR_API_KEY (none) API key for Semantic Scholar. Not required, but recommended for higher rate limits.
OPENALEX_EMAIL (none) Email for OpenAlex polite pool (higher rate limits). Only used when RAG_RETRIEVAL_API=openalex.
RAG_RETRIEVAL_API semantic_scholar Retrieval backend. Set to openalex to use the OpenAlex API instead of Semantic Scholar.
RAG_MODEL_1 meta-llama/Meta-Llama-3-8B-Instruct-Lite First TogetherAI model for summary generation.
RAG_MODEL_2 mistralai/Mistral-Small-24B-Instruct-2501 Second TogetherAI model for summary generation.
RAG_NUM_PAPERS 10 Number of papers to retrieve per query.
ADMIN_PASSWORD (none) Password for the admin dashboard. When not set, the dashboard is open (no login required).
DATA_DIR data Base directory for users.json and results/. Set to a mounted volume path (e.g. /data) for persistent storage on Railway.
SECRET_KEY (default) Secret key for signing session cookies. Set a random string in production.

Local PostgreSQL Setup

The app requires a PostgreSQL database. The easiest way to run one locally is with Docker.

Option A: Docker (recommended)

# Start a Postgres container
docker run --name annotation-pg \
  -e POSTGRES_USER=dev \
  -e POSTGRES_PASSWORD=dev \
  -e POSTGRES_DB=annotation \
  -p 5432:5432 \
  -d postgres:16

Then add this to your .env:

DATABASE_URL=postgresql://dev:dev@localhost:5432/annotation

To stop and remove the container later:

docker stop annotation-pg && docker rm annotation-pg

Option B: Homebrew (macOS)

brew install postgresql@16
brew services start postgresql@16
createdb annotation

Then add this to your .env:

DATABASE_URL=postgresql://localhost:5432/annotation

Connecting to the database

To inspect the database directly:

# Docker
docker exec -it annotation-pg psql -U dev -d annotation

# Homebrew
psql annotation

Useful commands once connected: \dt (list tables), \d users (describe a table), SELECT * FROM users;.

Migrating from flat files

If you have existing data in data/users.json and data/results/, import it into the database:

python scripts/migrate_json_to_db.py

Running the Server

uvicorn app:app --reload --port 8001

Tables are automatically created on startup. Then visit http://localhost:8001.

You can verify the database connection at http://localhost:8001/healthz.

Managing Annotators

Each annotator receives a unique token (no passwords). Tokens and per-user configuration live in the users table in PostgreSQL. Users can be added dynamically through the admin dashboard or the CLI — no server restart needed.

Admin Dashboard

Visit /admin to access the dashboard. From here you can:

  • Monitor progress — see which annotators have completed their queries, which are in progress, and which haven't started
  • Add annotators — create new users on the fly with the "Add Annotator" button

When ADMIN_PASSWORD is set, the dashboard requires login at /admin/login. When unset, the dashboard is open (useful for local development).

CLI

List existing users:

python scripts/create_users.py list

Generate multiple users at once:

python scripts/create_users.py generate --count 10 --queries 3

Add a single user:

python scripts/create_users.py add --name "Jane Doe" --queries 5

Share the generated token with the annotator — they enter it on the login page to begin.

Annotation Workflow

  1. Login — Annotator enters their token at /.
  2. Background questions — Pre-configured questions (radio, text, textarea) defined per user in data/users.json.
  3. Query input — Annotator types a free-form search query.
  4. RAG summaries — The system returns 4 summaries (A, B, C, D) for the query.
  5. Tournament — Annotator compares summaries in a bracket:
    • Round 1A: A vs B → select winner
    • Round 1B: C vs D → select winner
    • Final: the two winners face off
  6. Repeat — Steps 3–5 repeat for the configured number of queries.
  7. Done — Results are saved and the annotator sees a completion page.

Progress is saved after each query, so annotators can close and resume later.

RAG Pipeline

The RAG system (rag.py) has two stages:

  1. Retrieval — Fetches paper abstracts from Semantic Scholar (default) or OpenAlex, controlled by RAG_RETRIEVAL_API.
  2. Generation — Produces 4 summaries using two TogetherAI models, each with two prompt variants:
    • Without user background info
    • With user background info (from pre-task questions)

Labels A–D are randomly shuffled per query so annotators cannot infer which model or variant produced a summary. Model identity, prompt variant, and paper IDs are stored as metadata alongside each annotation.

Data Storage

All data is stored in PostgreSQL across three tables:

Table Description
users Annotator tokens, names, query counts, and pre-task question definitions (JSONB)
results Per-user result records with pre-task answers (JSONB) and completion status
annotations Individual query annotations with summaries, metadata, rankings, and timestamps

The admin dashboard export (/admin/export) generates a ZIP containing users.json and per-token results/{token}.json and results/{token}.csv files, preserving the same format as the previous flat-file storage.

Customizing Pre-task Questions

Pre-task questions are stored as a JSONB array in the users table. They can be configured per user when adding annotators via the admin dashboard or CLI. Supported question types:

Type Description
radio Single-select from options list
text Short free-text input
textarea Multi-line free-text input

Example:

{
  "id": "expertise",
  "type": "radio",
  "text": "How would you rate your expertise?",
  "options": ["Novice", "Intermediate", "Expert"]
}

Project Structure

app.py                    FastAPI application
database.py               PostgreSQL data layer (SQLAlchemy Core)
rag.py                    RAG pipeline (retrieval + generation)
templates/
  base.html               Shared layout
  start.html              Token login page
  questions.html           Pre-task background questions
  query.html              Query input + tournament UI
  done.html               Completion page
  admin.html              Admin dashboard
  admin_login.html        Admin login page
static/
  style.css               Styles
scripts/
  create_users.py         CLI for managing annotator tokens
  migrate_json_to_db.py   Migrate flat-file data to PostgreSQL

Deployment (Railway)

The app is configured for deployment on Railway via GitHub integration.

  1. Push the repo to GitHub.
  2. Create a new project on Railway and connect the GitHub repo.
  3. Add a Postgres plugin to the project — Railway will automatically set DATABASE_URL.
  4. Add environment variables in the Railway dashboard:
    • TOGETHER_API_KEY (required)
    • SECRET_KEY (recommended — set a random string for production sessions)
    • ADMIN_PASSWORD (recommended — protects the admin dashboard)
    • See example.env for the full list of optional variables.
  5. Railway will auto-detect the Python app and deploy using railway.toml.

Database tables are created automatically on first startup. A health check at /healthz verifies the database connection.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages