Control EXACTLY what your agent knows when the context window is about to overflow!
ContextSieve offers multiple tools for transparent context window management on Gemini chats, offering a pin-based UI/UX system per prompt/response that anyone can understand. Define a context limit, and you'll get timely alerts and options to optimize the context stack as it fills up. Try it out now!
- Live Token Telemetry: Track token breakdowns across the system prompt, user inputs and assistant outputs.
- Near-Capacity Monitoring: Sends alerts before a user-defined context limit.
- AI Context Summarizer: Compress old turns into executive summaries.
- Sliding Window Pruning: Retain recent chats and preserve pinned facts.
- Active Context Toggling: Omit verbose response blocks individually from the context.
- Inline Editing: Modify context blocks directly inside the active window.
Just one round of summarization + trimming filler text, and ContextSieve saved 79.9% of the context window (see below)!
-
Clone repo:
git clone [https://github.com/zh-yng/ContextSieve.git](https://github.com/zh-yng/ContextSieve.git) cd ContextSieve -
Install deps:
npm install
-
Create
.envor.env.localand grab an API Key from Google AI Studio:GEMINI_API_KEY="your_api_key_here"
-
Run app:
npm run dev
-
Open browser:
http://localhost:3000

