Repository navigation
Enforcing token limits
Keeping Rust source files under 25,000 tokens is essential for effective AI coding assistance, as large files can exceed model context windows and degrade assistant performance. After extensive research into the Rust ecosystem, I've identified multiple practical approaches that can be integrated into your CI/CD pipeline to automatically enforce these limits.
The challenge is unique because traditional file size metrics (lines or bytes) don't accurately represent how AI models tokenize code. A 1MB file might contain fewer tokens than a 500KB file depending on code density and formatting. This research focuses on tools that count tokens the way AI models actually process them, along with automation strategies to prevent oversized files from entering your codebase.
The most critical component is accurately counting tokens as AI models see them. tiktoken-rs emerges as the recommended solution, providing a Rust implementation of OpenAI's tiktoken library. It supports multiple encoding schemes including cl100k_base (used by GPT-3.5/GPT-4) and o200k_base (GPT-4o), ensuring your token counts match what AI assistants actually process.
For command-line integration, tc (token-counter) provides the simplest path forward. Installing via cargo install token-counter, you can check files with tc src/*.rs and integrate it directly into CI/CD pipelines. It defaults to the cl100k_base encoding and supports any HuggingFace tokenizer model, making it versatile for different AI backends.
Here's a practical example of counting tokens in Rust code:
use tiktoken_rs::cl100k_base;
use std::fs;
fn count_rust_file_tokens(file_path: &str) -> Result<usize, Box<dyn std::error::Error>> {
let content = fs::read_to_string(file_path)?;
let bpe = cl100k_base()?;
let tokens = bpe.encode_with_special_tokens(&content);
Ok(tokens.len())
}For teams using HuggingFace models, the tokenizers crate offers blazing-fast performance, processing gigabytes of text in under 20 seconds. While it requires more setup, it provides flexibility for teams using non-OpenAI models.
Automating file size checks in your CI/CD pipeline ensures no oversized files slip through code review. The freenet-actions/check-file-size action provides comprehensive checking with automatic PR comments listing violations. While it checks file size in bytes rather than tokens, it's useful for catching obviously oversized files early.
For token-specific checking, here's a complete GitHub Actions workflow:
name: Token Limit Enforcement
on: [pull_request]
jobs:
token-check:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Install Rust toolchain
uses: dtolnay/rust-toolchain@stable
- name: Install token counter
run: cargo install token-counter
- name: Check token counts
run: |
echo "## Token Count Report" >> $GITHUB_STEP_SUMMARY
# Check each Rust file
for file in $(find src -name "*.rs"); do
TOKEN_COUNT=$(tc "$file" | grep -o '[0-9]\+' | tail -1)
echo "$file: $TOKEN_COUNT tokens" >> $GITHUB_STEP_SUMMARY
if [ $TOKEN_COUNT -gt 25000 ]; then
echo "::error file=$file::File exceeds 25,000 token limit ($TOKEN_COUNT tokens)"
exit 1
fi
done
- name: Analyze changed files specifically
run: |
# Get list of changed Rust files
CHANGED_FILES=$(git diff --name-only origin/main...HEAD | grep '\.rs$' || true)
if [ -n "$CHANGED_FILES" ]; then
echo "### Changed Files Analysis" >> $GITHUB_STEP_SUMMARY
for file in $CHANGED_FILES; do
if [ -f "$file" ]; then
TOKEN_COUNT=$(tc "$file" | grep -o '[0-9]\+' | tail -1)
echo "- $file: $TOKEN_COUNT tokens" >> $GITHUB_STEP_SUMMARY
fi
done
fiTo make these checks required for merging, navigate to your repository's branch protection rules and add the job name (e.g., "token-check") to the list of required status checks. This prevents any PR with oversized files from being merged.
While Clippy doesn't have built-in token counting lints, you can leverage existing complexity lints or create custom ones. The too-many-lines-threshold configuration in clippy.toml provides basic size limiting:
too-many-lines-threshold = 500
cognitive-complexity-threshold = 25For more sophisticated needs, Dylint offers a maintainable approach to custom lints without forking Clippy. After installing with cargo install cargo-dylint dylint-link, you can create project-specific lints that check file-level metrics including token counts. This approach is particularly valuable for teams wanting to enforce custom standards beyond simple size limits.
Catching issues before they reach CI saves time and frustration. The pre-commit framework with Rust-specific hooks provides automatic local checking:
# .pre-commit-config.yaml
repos:
- repo: https://github.com/pre-commit/pre-commit-hooks
rev: v4.4.0
hooks:
- id: check-added-large-files
args: ["--maxkb=1000"] # 1MB limit
- repo: local
hooks:
- id: token-count-check
name: Check token counts
entry: bash -c 'for f in $(git diff --staged --name-only | grep "\.rs$"); do if [ -f "$f" ]; then tokens=$(tc "$f" | grep -o "[0-9]\+" | tail -1); if [ "$tokens" -gt 25000 ]; then echo "$f exceeds 25,000 tokens ($tokens)"; exit 1; fi; fi; done'
language: system
files: '\.rs$'For teams preferring native Git hooks, a simple bash script in .git/hooks/pre-commit can provide similar protection:
#!/bin/bash
MAX_TOKENS=25000
for file in $(git diff --staged --name-only | grep '\.rs$'); do
if [ -f "$file" ]; then
TOKEN_COUNT=$(tc "$file" | grep -o '[0-9]\+' | tail -1)
if [ "$TOKEN_COUNT" -gt "$MAX_TOKENS" ]; then
echo "Error: $file has $TOKEN_COUNT tokens (limit: $MAX_TOKENS)"
exit 1
fi
fi
doneIntegrating checks into your build process ensures they're never skipped. A build.rs script can enforce limits at compile time:
// build.rs
use std::fs;
use std::process::Command;
fn main() {
// Only run in CI to avoid slowing local development
if std::env::var("CI").is_ok() {
let output = Command::new("tc")
.args(&["src", "--format", "json"])
.output()
.expect("Failed to run token counter");
// Parse output and check limits
// Panic if any file exceeds 25,000 tokens
}
println!("cargo:rerun-if-changed=src/");
}Task runners like cargo-make or just provide more flexible automation. Here's a justfile example:
# Check token counts before building
check-tokens:
#!/usr/bin/env bash
set -e
echo "Checking token counts..."
for file in $(find src -name "*.rs"); do
tokens=$(tc "$file" | grep -o '[0-9]\+' | tail -1)
if [ "$tokens" -gt 25000 ]; then
echo "ERROR: $file has $tokens tokens (exceeds 25,000 limit)"
exit 1
fi
done
echo "All files within token limits ✓"
# Build with pre-checks
build: check-tokens
cargo build --release
# Run all quality checks
qa: check-tokens
cargo fmt --check
cargo clippy -- -D warnings
cargo testResearch reveals a significant gap in the Rust ecosystem: no battle-tested tools specifically designed for AI coding agent file size enforcement exist today. While rust-code-analysis from Mozilla provides code metrics and cargo-bloat analyzes binary sizes, neither addresses source file token limits for AI consumption.
The community has identified related challenges, particularly with rust-analyzer's 2.56MB file size limit causing IDE features to fail on large files. This highlights the broader need for better large file handling in Rust tooling. However, these limits are based on bytes rather than tokens, missing the unique requirements of AI assistants.
For teams needing immediate solutions, combining existing tools provides effective enforcement. tiktoken-rs for accurate token counting, tc for CLI integration, GitHub Actions for CI/CD automation, and pre-commit hooks for local development create a comprehensive system that prevents oversized files from impacting AI assistant effectiveness.
Looking forward, the Rust community would benefit from purpose-built crates that combine token counting with code quality metrics, provide intelligent file splitting suggestions, and integrate seamlessly with both traditional development tools and AI assistants. Until such tools emerge, the approaches outlined here offer practical, implementable solutions for keeping your Rust codebase AI-friendly.