Open-model inference

Private inference for coding agents.

Fast multi-region endpoint for open-weight models. Zero data retention, zero training, and optimized for long-context workloads.

  • CodePrivate
  • ModelsOpen
  • SetupMinutes
Agent integrations
Zro
Claude Code
Codex
Cursor
Cline
Pi
OpenClaw
OpenCode
Hermes
  • API
  • AgentsCLI + IDE
  • InfraMulti-region
Multi-region infrastructure

Private inference across regions.

Zro serves coding-agent workloads across multiple regions, with zero request retention and no training on customer data.

Location

Multi-region

Retention

Zero by default

Training

Never

Zro is tuned for long-context, multi-turn coding sessions. Under the endpoint, MoonMath applies HyperQuant compression, custom kernels, and hardware-aware deployment.

Compression

HyperQuant

Kernels

Custom attention

Hardware

AMD · NVIDIA · TPU

Performance

Built for coding‑agent workloads.

Coding agents

$ zro launch claude, codex, opencode, hermes, openclaw, pi

Install the npm package, log in once, then launch supported coding tools with temporary session config.

@moonmath-ai/zro npm packageClaude Code · Codex CLI
npm install -g @moonmath-ai/zro
zro login
zro launch claude
zro launch codex
Supports Claude Code, Codex CLI, OpenCode, Hermes, OpenClaw, and Pi.
Coding models

Open coding models.
One endpoint.

Open-source models are becoming competitive with closed-source systems for coding tasks. iZro serves MiniMax M3, GLM-5.2, and Kimi K2.7 Code, with more open coding models coming soon.

Details

Built for private inference.

Zro is a private inference endpoint for coding agents. It serves open-weight coding models across multiple regions, with zero request retention, no training on customer data, and setup paths for the tools developers already use.

Yes. Zro exposes OpenAI-compatible access for chat completions, so existing clients and agent tools can point at the Zro base URL.

Yes. Zro also supports Anthropic-compatible Messages requests at /v1/messages for tools that expect that API shape.

Use the @moonmath-ai/zro npm package for launch-supported tools: run zro login once, then launch your harness with zro launch. The integrations page also covers manual setup for Cursor and Cline.

No. Prompt and completion bodies are not retained by default after inference is processed.

Yes. Zro is built for responsive, streaming inference, so developer tools, agents, and production apps do not have to trade speed for privacy.

MiniMax M3, GLM-5.2, and Kimi K2.7 Code are available now, with availability shown for each region when you create an API key.

No. Customer prompts and completions are never used for training, fine-tuning, evaluations, analytics, or dataset creation.

Zro runs on privacy-forward multi-region infrastructure, with zero request retention by default.

Plans start at $20/month for $60 of inference spend. Usage packs are available without a subscription and expire after 90 days. Plan spend resets monthly.