Skip to content

Enforcing token limits

Konstantin Vyatkin edited this page Jun 24, 2025 · 1 revision

Enforcing token limits for AI-friendly Rust development

Keeping Rust source files under 25,000 tokens is essential for effective AI coding assistance, as large files can exceed model context windows and degrade assistant performance. After extensive research into the Rust ecosystem, I've identified multiple practical approaches that can be integrated into your CI/CD pipeline to automatically enforce these limits.

The challenge is unique because traditional file size metrics (lines or bytes) don't accurately represent how AI models tokenize code. A 1MB file might contain fewer tokens than a 500KB file depending on code density and formatting. This research focuses on tools that count tokens the way AI models actually process them, along with automation strategies to prevent oversized files from entering your codebase.

Token counting tools built for AI models

The most critical component is accurately counting tokens as AI models see them. tiktoken-rs emerges as the recommended solution, providing a Rust implementation of OpenAI's tiktoken library. It supports multiple encoding schemes including cl100k_base (used by GPT-3.5/GPT-4) and o200k_base (GPT-4o), ensuring your token counts match what AI assistants actually process.

For command-line integration, tc (token-counter) provides the simplest path forward. Installing via cargo install token-counter, you can check files with tc src/*.rs and integrate it directly into CI/CD pipelines. It defaults to the cl100k_base encoding and supports any HuggingFace tokenizer model, making it versatile for different AI backends.

Here's a practical example of counting tokens in Rust code:

use tiktoken_rs::cl100k_base;
use std::fs;

fn count_rust_file_tokens(file_path: &str) -> Result<usize, Box<dyn std::error::Error>> {
    let content = fs::read_to_string(file_path)?;
    let bpe = cl100k_base()?;
    let tokens = bpe.encode_with_special_tokens(&content);
    Ok(tokens.len())
}

For teams using HuggingFace models, the tokenizers crate offers blazing-fast performance, processing gigabytes of text in under 20 seconds. While it requires more setup, it provides flexibility for teams using non-OpenAI models.

GitHub Actions workflows that catch oversized files

Automating file size checks in your CI/CD pipeline ensures no oversized files slip through code review. The freenet-actions/check-file-size action provides comprehensive checking with automatic PR comments listing violations. While it checks file size in bytes rather than tokens, it's useful for catching obviously oversized files early.

For token-specific checking, here's a complete GitHub Actions workflow:

name: Token Limit Enforcement
on: [pull_request]

jobs:
  token-check:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      
      - name: Install Rust toolchain
        uses: dtolnay/rust-toolchain@stable
      
      - name: Install token counter
        run: cargo install token-counter
      
      - name: Check token counts
        run: |
          echo "## Token Count Report" >> $GITHUB_STEP_SUMMARY
          
          # Check each Rust file
          for file in $(find src -name "*.rs"); do
            TOKEN_COUNT=$(tc "$file" | grep -o '[0-9]\+' | tail -1)
            echo "$file: $TOKEN_COUNT tokens" >> $GITHUB_STEP_SUMMARY
            
            if [ $TOKEN_COUNT -gt 25000 ]; then
              echo "::error file=$file::File exceeds 25,000 token limit ($TOKEN_COUNT tokens)"
              exit 1
            fi
          done
      
      - name: Analyze changed files specifically
        run: |
          # Get list of changed Rust files
          CHANGED_FILES=$(git diff --name-only origin/main...HEAD | grep '\.rs$' || true)
          
          if [ -n "$CHANGED_FILES" ]; then
            echo "### Changed Files Analysis" >> $GITHUB_STEP_SUMMARY
            for file in $CHANGED_FILES; do
              if [ -f "$file" ]; then
                TOKEN_COUNT=$(tc "$file" | grep -o '[0-9]\+' | tail -1)
                echo "- $file: $TOKEN_COUNT tokens" >> $GITHUB_STEP_SUMMARY
              fi
            done
          fi

To make these checks required for merging, navigate to your repository's branch protection rules and add the job name (e.g., "token-check") to the list of required status checks. This prevents any PR with oversized files from being merged.

Custom lints through Clippy and Dylint

While Clippy doesn't have built-in token counting lints, you can leverage existing complexity lints or create custom ones. The too-many-lines-threshold configuration in clippy.toml provides basic size limiting:

too-many-lines-threshold = 500
cognitive-complexity-threshold = 25

For more sophisticated needs, Dylint offers a maintainable approach to custom lints without forking Clippy. After installing with cargo install cargo-dylint dylint-link, you can create project-specific lints that check file-level metrics including token counts. This approach is particularly valuable for teams wanting to enforce custom standards beyond simple size limits.

Pre-commit hooks and local enforcement

Catching issues before they reach CI saves time and frustration. The pre-commit framework with Rust-specific hooks provides automatic local checking:

# .pre-commit-config.yaml
repos:
  - repo: https://github.com/pre-commit/pre-commit-hooks
    rev: v4.4.0
    hooks:
      - id: check-added-large-files
        args: ["--maxkb=1000"]  # 1MB limit
  
  - repo: local
    hooks:
      - id: token-count-check
        name: Check token counts
        entry: bash -c 'for f in $(git diff --staged --name-only | grep "\.rs$"); do if [ -f "$f" ]; then tokens=$(tc "$f" | grep -o "[0-9]\+" | tail -1); if [ "$tokens" -gt 25000 ]; then echo "$f exceeds 25,000 tokens ($tokens)"; exit 1; fi; fi; done'
        language: system
        files: '\.rs$'

For teams preferring native Git hooks, a simple bash script in .git/hooks/pre-commit can provide similar protection:

#!/bin/bash
MAX_TOKENS=25000

for file in $(git diff --staged --name-only | grep '\.rs$'); do
    if [ -f "$file" ]; then
        TOKEN_COUNT=$(tc "$file" | grep -o '[0-9]\+' | tail -1)
        if [ "$TOKEN_COUNT" -gt "$MAX_TOKENS" ]; then
            echo "Error: $file has $TOKEN_COUNT tokens (limit: $MAX_TOKENS)"
            exit 1
        fi
    fi
done

Build-time validation and task automation

Integrating checks into your build process ensures they're never skipped. A build.rs script can enforce limits at compile time:

// build.rs
use std::fs;
use std::process::Command;

fn main() {
    // Only run in CI to avoid slowing local development
    if std::env::var("CI").is_ok() {
        let output = Command::new("tc")
            .args(&["src", "--format", "json"])
            .output()
            .expect("Failed to run token counter");
        
        // Parse output and check limits
        // Panic if any file exceeds 25,000 tokens
    }
    
    println!("cargo:rerun-if-changed=src/");
}

Task runners like cargo-make or just provide more flexible automation. Here's a justfile example:

# Check token counts before building
check-tokens:
    #!/usr/bin/env bash
    set -e
    echo "Checking token counts..."
    
    for file in $(find src -name "*.rs"); do
        tokens=$(tc "$file" | grep -o '[0-9]\+' | tail -1)
        if [ "$tokens" -gt 25000 ]; then
            echo "ERROR: $file has $tokens tokens (exceeds 25,000 limit)"
            exit 1
        fi
    done
    
    echo "All files within token limits ✓"

# Build with pre-checks
build: check-tokens
    cargo build --release

# Run all quality checks
qa: check-tokens
    cargo fmt --check
    cargo clippy -- -D warnings
    cargo test

Addressing the ecosystem gap

Research reveals a significant gap in the Rust ecosystem: no battle-tested tools specifically designed for AI coding agent file size enforcement exist today. While rust-code-analysis from Mozilla provides code metrics and cargo-bloat analyzes binary sizes, neither addresses source file token limits for AI consumption.

The community has identified related challenges, particularly with rust-analyzer's 2.56MB file size limit causing IDE features to fail on large files. This highlights the broader need for better large file handling in Rust tooling. However, these limits are based on bytes rather than tokens, missing the unique requirements of AI assistants.

For teams needing immediate solutions, combining existing tools provides effective enforcement. tiktoken-rs for accurate token counting, tc for CLI integration, GitHub Actions for CI/CD automation, and pre-commit hooks for local development create a comprehensive system that prevents oversized files from impacting AI assistant effectiveness.

Looking forward, the Rust community would benefit from purpose-built crates that combine token counting with code quality metrics, provide intelligent file splitting suggestions, and integrate seamlessly with both traditional development tools and AI assistants. Until such tools emerge, the approaches outlined here offer practical, implementable solutions for keeping your Rust codebase AI-friendly.