The gateway for AI automations
Power every AI automation. One endpoint.
AI-Proxy is a local, self-hosted gateway built for automation engineers. Point n8n, Cursor, your agents and pipelines at one OpenAI-compatible endpoint — running on the Gemini, Claude and ChatGPT plans you already pay for. No per-token keys. No surprise bills.
What AI-Proxy unlocks
Run your AI automations on plans you own
n8n flows, Cursor, opencode, Buzz agents, cron jobs — point them all at one OpenAI-compatible endpoint and they run on the Gemini, Claude and ChatGPT subscriptions you already pay for. No per-run token bills on your busiest pipelines.
One-click quick configs
Pre-built quick-config buttons wire your favourite tools and automations to the right models in a single click — no fiddling with base URLs, keys or model names. Ship the workflow, not the setup.
One key, every model
Gemini, Claude, Codex and more behind a single endpoint. You never switch keys — you just switch the model name. Your whole automation stack shares one credential you control.
Native tool-calling for real agents
Agentic clients — the n8n Tools Agent, opencode, Cursor, your own agents — call functions natively through the one endpoint. tool_calls, streamed tool-call deltas over SSE, and the tool-result round-trip all pass through end-to-end, on the plans you already own. Not just chat: agents that take actions.
Beyond text — images, video, voice
Generate images through your Gemini subscription, and plug in HeyGen avatars, ElevenLabs voices, Luma video, or any OpenAI-compatible provider — so multimodal agents run through the same gateway.
Cloud + local in one list
Connect a local Ollama server with one click and your offline models join the same unified list — private, free, and available even with no internet. Mix cloud and local models in the same workflow.
Your machine, your keys
AI-Proxy runs locally and is fully self-hosted. OAuth tokens live on your Mac, never on our servers. It's designed to fail loudly when a quota runs out — never to silently bill per-token rates.
Built for automation engineers
You ship agents and workflows. Stop paying per token to run them.
n8n flows, Cursor, opencode, your own agents and cron jobs — every run burns API credits. Point them all at one OpenAI-compatible endpoint and they run on the Gemini, Claude and ChatGPT plans you already pay for. One key for the whole stack, no per-run token anxiety.
Real tools, not just chat
Function-calling works end-to-end — tool_calls, streamed deltas and the tool-result round-trip all pass through, so your agents actually take actions.
No per-run bills
High-volume pipelines don't rack up per-token charges — they ride the subscriptions you already have.
Cloud + local, one list
Mix frontier models with local Ollama in the same workflow — private and offline when you need it.
Never breaks on a dead model
A stale or retired model name self-heals down a candidate cascade to a live model instead of failing the run.
Install
Available for macOS via Homebrew. Linux support is on the roadmap.
# install via Homebrew
brew tap meta-thinking/tap
brew install --cask ai-proxy
AI-Proxy lives in your menu bar: sign in to your AI subscriptions from the app, then point your favourite tool at the local endpoint with the key it gives you.
Take the shortest route to innovation
Create an account and start free — the local gateway is yours forever, no card required. Upgrade to Pro for image, video & voice, or grab a one-time lifetime license.