Code base for "LLMs Alignment in Low Resource Settings" - A comprehensive framework for aligning large language models using reinforcement learning and preference learning techniques optimized for resource-constrained environments.
This repository implements an end-to-end pipeline for aligning large language models (LLMs) to human preferences through a structural curriculum learning approach. The system is specifically optimized for low-resource settings, using QLoRA (Quantized Low-Rank Adaptation).
- Iterative refinement of policy through multiple rounds
- Dynamic weighting of structural cues (distribution, length, overlap, variance)
- Automatic checkpoint selection based on training metrics
- State tracking for resumable training
- Rule-based generation guided by structural constraints
- Support for multiple structural types (lists, tables, code, math, dialogue, prose)
- Contrastive learning with margin-based validation
- Demonstration-based augmentation for improved quality (In latest iterations)
- Binary preference modeling on synthetic data
- QLoRA-based efficient fine-tuning
- Transfer learning from previous checkpoints
- Proximal Policy Optimization.
- reward model support
- Reference policy management
- Length bonus and token-level penalization
- RLVR support (Reinforcement Learning with Value Rewards) (still under testing)
- Ensemble reward modeling support (For experimentation with multiple reward models)
- Multiple evaluation datasets: Alpaca, HumanEval, IFEval, TruthfulQA
- Best-of-N sampling for improved performance
- Both greedy and sampling-based decoding
- Python 3.10+
- 16GB VRAM recommended (tested on a Tesla V100)
-
Clone the repository:
git clone <repository-url> cd Low_resource_llm_alignment
-
Install dependencies:
pip install -r requirements.txt
-
Set up Hugging Face token:
export HF_TOKEN="your_huggingface_token"
python Experiments_run.pyThis will:
- Generate synthetic preference data based on the current policy, current structural level.
- Post-process and balance the synthetic data
- Train a reward model
- Run PPO training
- Evaluate on multiple benchmarks
- Select the best checkpoint for the next iteration
bash scripts/synthetic_preference_data_generation.sh \
--outDir data/iteration_0 \
--promptsFile data/sampled_alpaca_with_ids&prose.json \
--checkpoint_dir checkpoints/policy_v1 \
--maxDistancePositive 0.6bash scripts/reward_model_training_script_run.sh \
--dataset_path data/preference_data.csv \
--output_dir reward_models/rm_v1 \
--per_device_train_batch_size 8bash scripts/ppo_training_run.sh \
--model_directory meta-llama/Llama-3.2-3b \
--reward_model_name_or_path reward_models/rm_v1 \
--dataset_path data/rl_training_data.json \
--output_dir PPO/policy_v1bash scripts/run_evaluation_generations.sh \
--dataset_type alpaca \
--checkpoint_dir PPO/policy_v1/adapter_model/lora_policy \
--outFile evaluations/results.jsonl{
"id": "unique_id",
"instruction": "What is machine learning?",
"input": "optional_context",
"output": "ground_truth_answer",
"structural_class": "Prose|Numbered-list|Table|..."
}instruction,input,chosen,rejected,preference,contrastive_status,structural_class,positive_distance,negative_distance,ground_truth- Alpaca: Instruction-following and helpfulness quality (gitHub: Bench)
- HumanEval: Code generation capability (gitHub: Bench)
- IFEval: Instruction following (gitHub: Bench)
- TruthfulQA: Factual accuracy.
- Reduce batch sizes
- Enable gradient checkpointing
- Use mixed precision training (
--fp16 True)
- Reduce learning rate (try 1e-6)
- Lower KL coefficient (
--kl_coef 0.01)
If you use this code for your research, please cite our paper:
This project is licensed under the Apache License 2.0. See LICENSE for details.
In our work, we relied on adapting several open-source codebases to our setup, including:
- RLCD as a building block for preference data generation (link: RLCD)
- AlpacaFarm and SALMON for Proximal Policy Optimization. (link: AlpacaFarm, SALMON)