Mission control for the fleet
powering our AI agents
AI Fleet Router turns our rack of GPUs into one endpoint that serves text, image, voice, video and music — routing every request to the fastest node, overflowing to the cloud only when it has to, and giving our team and our AI agents one place to generate.
Running our agents on our own GPUs
couldn't stay a black box
The moment we went past one model on one machine, we lost the plot. Here's the chaos AI Fleet Router was built to end.
No idea what's happening
Requests vanish into a cluster of boxes. Which node served it? How fast? Did it fall back to the cloud? We were flying blind.
Hot nodes, idle nodes
One Spark is pinned at 100% while three sit cold. Without a live view of GPU load, we couldn't balance the fleet we paid for.
Cloud bills we couldn't explain
Overflow quietly spills to paid cloud models. By month-end the invoice is a mystery and nobody can say which agent caused it.
No proof for the team
Clients and teammates want numbers — throughput, cost, uptime. Screenshotting terminals doesn't cut it. We needed real reports.
One control plane for the whole fleet
Generate, route, and observe — text, image, voice, video and music — across every node, local and cloud, from a single real-time control plane.
LLM inference
Serve chat and completion models to our apps and agents at full speed — one endpoint, tag-normalized model names, and a safe fallback for misconfigured clients.
// textImage generation
Generate images with FLUX on our own GPUs — plus Grok and MiniMax in the cloud. Text-to-image, image-to-image, and subject reference for a consistent character. Requests fan out to a free image pool — no queue, no cold start.
// imageVoice generation
Turn text into natural, expressive voice on the fleet — with emotion and pacing controls, same router, same private library. No audio ever leaves our hardware.
// voiceCapability-aware routing
The router scores every backend by model fit, current load, locality and client spread, then sends each request to the node that will answer it fastest.
// smart routingCloud overflow, capped
Local first. When the fleet saturates or a model isn't hosted, traffic overflows to cloud models — drawn from a fleet-wide pool so we never blow past our account limits.
// cost controlClient & agent portal
app.aifleetrouter.com: magic-link sign-in for people, scoped API keys for agents, strict per-tenant isolation, and a private library with 30-day retention.
// multi-tenantLive request feed
Watch generation stream in as it happens — model, node, client and tokens-per-second on every call, updating in real time.
// real-timePer-node fleet health
GPU, memory, CPU, disk, temperature and wattage for every backend — with loaded models and VRAM, and one-click drain for maintenance.
// gpu telemetryAnalytics & PDF reports
Per-model and per-client breakdowns of throughput, tokens and spend over 24h / 7d / 30d — exportable as a clean PDF for the team in seconds.
// reportingSubscription & quota tracking
Live "quota remaining" bars for every paid provider — Kimi, MiniMax, Grok, ChatGPT, OpenRouter and Nous Portal — with reset-window countdowns, per-client consumption, and a DM the moment any one drops below 20%.
// quotasFleet MCP — one server, every tool
The whole fleet as native agent tools over MCP — image, video, voice, voice cloning, music, transcription and RAG memory — for Claude, Cursor, Codex, OpenClaw or any MCP-capable agent. Prefer a raw API? The same fleet also speaks OpenAI- and Ollama-compatible endpoints.
// see it below ↓Uptime & incidents
Per-node and per-pool uptime with live incident detection and a 30-day health history across text, image and voice — so we catch a degraded backend before anyone notices.
// reliabilityVideo generation
Turn a still into motion — image-to-video via Grok or MiniMax, kicked off from a single prompt or an agent and logged in the same live feed as everything else.
// videoVoice cloning
Clone a voice from a short sample on the Fish-Speech engine — the router auto-transcribes the sample for you — then generate expressive speech in it, by name. The sample further down this page is a clone of Jeff.
// voice cloneMusic & transcription
Generate music on the fleet's subscription, and transcribe audio back to text on a dedicated speech-to-text pool — both one call away for any agent.
// music · sttEmbeddings & agent memory
Free local embeddings on the fleet (bge-m3 by default) power per-agent RAG memory — retrieval our agents can lean on, running on our own GPUs and immune to node drains.
// embeddings · ragVoice guardrails
A locked voice stays locked across a whole batch, an unknown reference is rejected instead of silently swapped, and a failure comes back as a structured error an agent can actually act on — not a guess.
// safetyBuilt like the terminal we live in
Real screenshots from the live dashboard — client names and IPs redacted. Dense, fast, and information-rich.
The whole fleet on one screen
Open the dashboard and we know exactly what the fleet is doing: live provider quotas, a streaming request feed across text, image and voice, and every node's health — all in one view.
- ▸Provider quota bars — Kimi, MiniMax, Ollama Cloud, video
- ▸Streaming live feed — text, images and audio
- ▸All-time requests, tokens, avg t/s and uptime up top
- ▸Every node's health at a glance
Every node, healthy or not — at a glance
The Fleet view gives each backend its own live card: utilization bars for GPU, memory, CPU and disk, plus temperature and power draw. Loaded text, image and voice models show their VRAM footprint, and a single Drain toggle pulls a node out of rotation cleanly.
- ▸Real-time GPU / MEM / CPU / DISK meters per node
- ▸Loaded models with warm/cold state and VRAM
- ▸Image + voice generation models, per node
- ▸Health status, temp, wattage and in-flight counts
Know which model is actually fast
The Models view ranks everything running on the fleet by throughput — text, image and voice. Average and max tokens-per-second, time-to-first-token, total tokens and request counts, so we can right-size which model runs where and spot the cloud models pulling their weight (or not).
- ▸Requests-vs-speed across every model on the fleet
- ▸Latency — average duration and TTFT per model
- ▸Local vs cloud models, side by side
- ▸24h / 7d / 30d windows
Local first. Cloud when it counts.
The router keeps work on our own silicon by default and overflows to six cloud providers only when the fleet is saturated or a request needs a model we don't host — each with its own live quota bar, reset countdown, and automatic fallback chain.
- ▸Live local-vs-cloud traffic split (71/29 all-time)
- ▸Per-provider quota, reset window, and today's usage
- ▸Falls back automatically when a subscription is exhausted
- ▸One-click drain for clean maintenance
A studio for our team, clients, and agents
app.aifleetrouter.com is a hardened, multi-tenant studio on top of the fleet — Images, Voice and a shared Library. People and clients sign in with a magic link; agents authenticate with a scoped key or MCP. Everything they create lands in their own private library, generated entirely on our hardware.
- ▸Magic-link sign-in, strict per-tenant isolation
- ▸Image + voice generation, routed across the fleet
- ▸Agent-native: one-command install or MCP — output lands in the library
- ▸Private library, 30-day retention — nothing leaves our box
Photorealistic — and perfectly on-brand
Images render on our own GPUs with FLUX — schnell for speed, dev for quality — and a saved character keeps a consistent likeness across every generation. Each result carries its model, seed and full prompt, ready to download.
- ▸FLUX schnell + dev, on local silicon
- ▸Reusable characters for a consistent likeness
- ▸“Put yourself in it” + one-tap prompt optimizer
- ▸Model, seed & full prompt saved with every image
One MCP server. Every tool on the fleet.
Point any MCP-capable agent — Claude, Cursor, Codex, OpenClaw — at fleet-mcp and it gets image, video, voice, voice cloning, music, transcription and RAG memory as native tools, running on our own hardware. No separate API keys per provider, no separate integration per modality.
- ▸7 tool families: image · video · voice · clone · music · transcribe · embed/RAG
- ▸Router mode for our own agents, portal mode with per-client keys for isolated access
- ▸Self-bootstrapping install — one command wires the whole mount
- ▸Structured errors + guardrails, so a tool call fails loud instead of guessing
"[excited] Big news — we just launched our AI Fleet Router with multi-modal features to produce voice, images, text and more. [emphasis] This is a voice clone of Jeff J Hunter!"
Not a demo — it runs our agents
AI Fleet Router serves the real agents and products we operate every day. Point any agent framework at one endpoint and it's on the fleet.

Why a marketer built a GPU router
AI Fleet Router didn't come from a lab. It came from needing to run a small army of AI agents — reliably, privately, and without a runaway cloud bill.
From AI Persona Method™ to a fleet of GPUs
Jeff has spent 11+ years scaling businesses with humans + automation — featured in Entrepreneur Magazine and Business Insider, creator of the AI Persona Method™, and the founder behind 1,000+ students building businesses that run without them.
As his team deployed 15+ AI Employees across messaging channels, the question stopped being "can AI do the work" and became "where does all this inference actually run?" Renting cloud tokens for every agent doesn't scale — so Jeff stood up a fleet of local GPU boxes to serve models privately. But a pile of Sparks with no visibility is just expensive guesswork. AI Fleet Router was the missing control plane — the dashboard that finally made the fleet observable, balanced, and accountable.
Build your own AI-powered income
AI Fleet Router is the kind of infrastructure that runs a business on AI. Learn the playbook behind it — 8 proven ways to make money with AI, live group calls, and 100+ guides — inside AI Money Group.