Routing across free LLM tiers

A free public chat dies the first afternoon you pin it to one generous API key. Caps are real. The only honest design is a queue of providers and a local floor.

Most “free AI chat” landings are a wrapper around a single vendor. That works until the daily request budget trips, the model is deprecated, or the ToS forbids resale. Then the page 500s and the ad slot is the only thing still loading.

agime.ai’s public chat does not do that. Each reply walks an ordered list: Groq, Gemini, OpenRouter, then other ToS-safe free tiers. If a key is unset, that adapter is skipped. If a provider returns 429 or the day’s counter is spent, the next adapter runs. Only when the free list is dry do we touch the local GPU — one concurrent slot, so the paid LIZ units keep theirs.

Two rules make this survivable. First, the budget is persisted per provider and UTC day, not guessed in memory. Second, we log which tier served the reply and never log the message. The free page already tells you it is not private. We still refuse to ship your prompt into an analytics vendor.

This is the same routing idea as the rest of the platform, cut down for a page that has no account. If you want the private version — one process, one database, memory that stays yours — that is LIZ. If you just want to try a reply, open the chat.

All notes