Live observability for self-hosted LLM fleets

Mission control for the fleet
powering our AI agents

AI Fleet Router turns our rack of GPUs into one endpoint that serves text, image, voice, video and music — routing every request to the fastest node, overflowing to the cloud only when it has to, and giving our team and our AI agents one place to generate.

// the infrastructure behind our AI employees, agents, and client work
AI FLEET ROUTER / overview
aifleetrouter.com
Requests
116.2k
24h · 3.2k
Tokens
45.4M
24h · 1.8M
Avg t/s
27.9
all time
Inf. time
17.2d
cumulative
Clients
59
active · 13
Local/Cloud
71/29
% split
Live Feed · streaming
gemma4:26bacer-2
client · agent-research
58.4 t/s10:07:46
ornith:35bacer-03
client · openclaw-ops
40.7 t/s10:07:54
gemma4:31b-cloudgx10-1
client · agent-support
31.6 t/s10:07:58
acer-2Acer GN100 Spark #2HEALTHY
GPU
78%
MEM
85%
TEMP
49°
gx10-1GX10 Spark 1HEALTHY
GPU
41%
MEM
76%
TEMP
39°
6 GPU nodes · DGX Sparks
6 cloud providers · Kimi · Grok · ChatGPT · MiniMax · OpenRouter · Nous
6 modalities · text · image · video · voice · music · stt
116.2k requests · routed
45.4M tokens · generated
96.4% uptime · 30-day
Why we built it

Running our agents on our own GPUs
couldn't stay a black box

The moment we went past one model on one machine, we lost the plot. Here's the chaos AI Fleet Router was built to end.

01

No idea what's happening

Requests vanish into a cluster of boxes. Which node served it? How fast? Did it fall back to the cloud? We were flying blind.

02

Hot nodes, idle nodes

One Spark is pinned at 100% while three sit cold. Without a live view of GPU load, we couldn't balance the fleet we paid for.

03

Cloud bills we couldn't explain

Overflow quietly spills to paid cloud models. By month-end the invoice is a mystery and nobody can say which agent caused it.

04

No proof for the team

Clients and teammates want numbers — throughput, cost, uptime. Screenshotting terminals doesn't cut it. We needed real reports.

What it does

One control plane for the whole fleet

Generate, route, and observe — text, image, voice, video and music — across every node, local and cloud, from a single real-time control plane.

💬

LLM inference

Serve chat and completion models to our apps and agents at full speed — one endpoint, tag-normalized model names, and a safe fallback for misconfigured clients.

// text
🎨

Image generation

Generate images with FLUX on our own GPUs — plus Grok and MiniMax in the cloud. Text-to-image, image-to-image, and subject reference for a consistent character. Requests fan out to a free image pool — no queue, no cold start.

// image
🎙️

Voice generation

Turn text into natural, expressive voice on the fleet — with emotion and pacing controls, same router, same private library. No audio ever leaves our hardware.

// voice
🧭

Capability-aware routing

The router scores every backend by model fit, current load, locality and client spread, then sends each request to the node that will answer it fastest.

// smart routing
☁️

Cloud overflow, capped

Local first. When the fleet saturates or a model isn't hosted, traffic overflows to cloud models — drawn from a fleet-wide pool so we never blow past our account limits.

// cost control
🔑

Client & agent portal

app.aifleetrouter.com: magic-link sign-in for people, scoped API keys for agents, strict per-tenant isolation, and a private library with 30-day retention.

// multi-tenant
📡

Live request feed

Watch generation stream in as it happens — model, node, client and tokens-per-second on every call, updating in real time.

// real-time
🖥️

Per-node fleet health

GPU, memory, CPU, disk, temperature and wattage for every backend — with loaded models and VRAM, and one-click drain for maintenance.

// gpu telemetry
📊

Analytics & PDF reports

Per-model and per-client breakdowns of throughput, tokens and spend over 24h / 7d / 30d — exportable as a clean PDF for the team in seconds.

// reporting
💳

Subscription & quota tracking

Live "quota remaining" bars for every paid provider — Kimi, MiniMax, Grok, ChatGPT, OpenRouter and Nous Portal — with reset-window countdowns, per-client consumption, and a DM the moment any one drops below 20%.

// quotas
🔌

Fleet MCP — one server, every tool

The whole fleet as native agent tools over MCP — image, video, voice, voice cloning, music, transcription and RAG memory — for Claude, Cursor, Codex, OpenClaw or any MCP-capable agent. Prefer a raw API? The same fleet also speaks OpenAI- and Ollama-compatible endpoints.

// see it below ↓
🩺

Uptime & incidents

Per-node and per-pool uptime with live incident detection and a 30-day health history across text, image and voice — so we catch a degraded backend before anyone notices.

// reliability
🎬

Video generation

Turn a still into motion — image-to-video via Grok or MiniMax, kicked off from a single prompt or an agent and logged in the same live feed as everything else.

// video
🗣️

Voice cloning

Clone a voice from a short sample on the Fish-Speech engine — the router auto-transcribes the sample for you — then generate expressive speech in it, by name. The sample further down this page is a clone of Jeff.

// voice clone
🎵

Music & transcription

Generate music on the fleet's subscription, and transcribe audio back to text on a dedicated speech-to-text pool — both one call away for any agent.

// music · stt
🧠

Embeddings & agent memory

Free local embeddings on the fleet (bge-m3 by default) power per-agent RAG memory — retrieval our agents can lean on, running on our own GPUs and immune to node drains.

// embeddings · rag
🛡️

Voice guardrails

A locked voice stays locked across a whole batch, an unknown reference is rejected instead of silently swapped, and a failure comes back as a structured error an agent can actually act on — not a guess.

// safety
Inside the platform

Built like the terminal we live in

Real screenshots from the live dashboard — client names and IPs redacted. Dense, fast, and information-rich.

The whole fleet on one screen

Open the dashboard and we know exactly what the fleet is doing: live provider quotas, a streaming request feed across text, image and voice, and every node's health — all in one view.

  • Provider quota bars — Kimi, MiniMax, Ollama Cloud, video
  • Streaming live feed — text, images and audio
  • All-time requests, tokens, avg t/s and uptime up top
  • Every node's health at a glance
Overview · live dashboardidentifiers redacted
AI Fleet Router dashboard overview — provider quotas, live request feed across text/image/voice, and per-node health

Every node, healthy or not — at a glance

The Fleet view gives each backend its own live card: utilization bars for GPU, memory, CPU and disk, plus temperature and power draw. Loaded text, image and voice models show their VRAM footprint, and a single Drain toggle pulls a node out of rotation cleanly.

  • Real-time GPU / MEM / CPU / DISK meters per node
  • Loaded models with warm/cold state and VRAM
  • Image + voice generation models, per node
  • Health status, temp, wattage and in-flight counts
Fleet · backendsIPs redacted
Fleet view — per-node GPU/MEM/CPU/DISK utilization, loaded models, and image/voice generation pools; node IPs redacted

Know which model is actually fast

The Models view ranks everything running on the fleet by throughput — text, image and voice. Average and max tokens-per-second, time-to-first-token, total tokens and request counts, so we can right-size which model runs where and spot the cloud models pulling their weight (or not).

  • Requests-vs-speed across every model on the fleet
  • Latency — average duration and TTFT per model
  • Local vs cloud models, side by side
  • 24h / 7d / 30d windows
Models · performance7d
Models view — requests vs speed and latency (duration + TTFT) ranked per model across the fleet

Local first. Cloud when it counts.

The router keeps work on our own silicon by default and overflows to six cloud providers only when the fleet is saturated or a request needs a model we don't host — each with its own live quota bar, reset countdown, and automatic fallback chain.

  • Live local-vs-cloud traffic split (71/29 all-time)
  • Per-provider quota, reset window, and today's usage
  • Falls back automatically when a subscription is exhausted
  • One-click drain for clean maintenance
Usage · provider quotaslive
Provider usage view — live quota bars, reset windows, and wired models for Kimi and Grok subscriptions

A studio for our team, clients, and agents

app.aifleetrouter.com is a hardened, multi-tenant studio on top of the fleet — Images, Voice and a shared Library. People and clients sign in with a magic link; agents authenticate with a scoped key or MCP. Everything they create lands in their own private library, generated entirely on our hardware.

  • Magic-link sign-in, strict per-tenant isolation
  • Image + voice generation, routed across the fleet
  • Agent-native: one-command install or MCP — output lands in the library
  • Private library, 30-day retention — nothing leaves our box
Studio · app.aifleetrouter.comLibrary
AI Fleet Studio Library - a grid of images generated on the fleet by the team and their agents

Photorealistic — and perfectly on-brand

Images render on our own GPUs with FLUX — schnell for speed, dev for quality — and a saved character keeps a consistent likeness across every generation. Each result carries its model, seed and full prompt, ready to download.

  • FLUX schnell + dev, on local silicon
  • Reusable characters for a consistent likeness
  • “Put yourself in it” + one-tap prompt optimizer
  • Model, seed & full prompt saved with every image
Studio · image detailflux1-dev
AI Fleet Studio image detail - a photorealistic portrait generated with FLUX, showing model, seed and full prompt

One MCP server. Every tool on the fleet.

Point any MCP-capable agent — Claude, Cursor, Codex, OpenClaw — at fleet-mcp and it gets image, video, voice, voice cloning, music, transcription and RAG memory as native tools, running on our own hardware. No separate API keys per provider, no separate integration per modality.

  • 7 tool families: image · video · voice · clone · music · transcribe · embed/RAG
  • Router mode for our own agents, portal mode with per-client keys for isolated access
  • Self-bootstrapping install — one command wires the whole mount
  • Structured errors + guardrails, so a tool call fails loud instead of guessing
fleet-mcpv1.12
{ "mcpServers": { "fleet": { "command": "npx", "args": ["-y", "fleet-mcp"], "env": { "FLEET_KEY": "flk_••••" } } } }
tools exposed
fleet_image fleet_video fleet_voice fleet_voice_clone fleet_music fleet_transcribe fleet_embed
router mode portal mode · per-client keys
🔊 Real voice, generated on the fleet

"[excited] Big news — we just launched our AI Fleet Router with multi-modal features to produce voice, images, text and more. [emphasis] This is a voice clone of Jeff J Hunter!"

voice: jeffjhunter1 · expressive · openaudio-s2-pro · no third-party voice API
▶ 11-second sample · inline [tags] shape the delivery
Running in production

Not a demo — it runs our agents

AI Fleet Router serves the real agents and products we operate every day. Point any agent framework at one endpoint and it's on the fleet.

Nous Research Hermes logo
Hermes
Nous Research's autonomous agent — pointed at the router with a single install command, requests streaming straight into the dashboard.
🦞
OpenClaw
The open-source agent gateway working across WhatsApp, Telegram, Discord and more — served entirely on the local fleet.
🤖
AI Employees
15+ VA Staffer AI Employees answering across channels 24/7 — private inference, no per-token cloud bill.
⚙️
Our own products
AI Persona OS and the internal apps, portals and creative projects we ship — all generating on the same fleet.
The story

Why a marketer built a GPU router

AI Fleet Router didn't come from a lab. It came from needing to run a small army of AI agents — reliably, privately, and without a runaway cloud bill.

JJH
Jeff J Hunter
Built by Jeff J Hunter · Founder, VA Staffer

From AI Persona Method™ to a fleet of GPUs

Jeff has spent 11+ years scaling businesses with humans + automation — featured in Entrepreneur Magazine and Business Insider, creator of the AI Persona Method™, and the founder behind 1,000+ students building businesses that run without them.

As his team deployed 15+ AI Employees across messaging channels, the question stopped being "can AI do the work" and became "where does all this inference actually run?" Renting cloud tokens for every agent doesn't scale — so Jeff stood up a fleet of local GPU boxes to serve models privately. But a pile of Sparks with no visibility is just expensive guesswork. AI Fleet Router was the missing control plane — the dashboard that finally made the fleet observable, balanced, and accountable.

🦞
Early OpenClaw contributor & advocate. Jeff is an active supporter of OpenClaw — the open-source gateway that lets AI agents work across WhatsApp, Telegram, Discord and dozens more channels. Running 15+ AI Employees on OpenClaw + AI Persona OS, he's helped shape its security best practices and real-world deployment workflows — and AI Fleet Router is what keeps those agents fed with fast, local inference.
🌐 jeffjhunter.com 🦞 OpenClaw project 📬 TheTip.ai newsletter
Want this for your business?

Build your own AI-powered income

AI Fleet Router is the kind of infrastructure that runs a business on AI. Learn the playbook behind it — 8 proven ways to make money with AI, live group calls, and 100+ guides — inside AI Money Group.

// learn how Jeff runs 15+ AI Employees on his own fleet