Skip to main content
These controls are optional. Start with a provider, access profile, and gateway key in Gateway, then open the profile’s Advanced sections when you need them.

Use one profile across model APIs

A gateway key is assigned to an access profile, not to a client API. An entitled application can call Forge with Chat Completions, Responses, or Anthropic Messages while Forge selects the configured provider operation and translates supported text and tool requests. Bedrock Converse and Gemini Generate Content are also supported client and provider operations. The same profile, identity, policy, and budget apply to translated requests. Unsupported fields or content are rejected rather than silently removed. Streaming responses use the client’s API format. For incomplete Chat tool histories, Forge records a bounded repair and supplies missing results before policy checks and the provider call. Orphan results are removed, and the last duplicate result within a tool exchange is kept. Results are never moved across conversation turns. For Claude models, Forge checks effort levels against the selected model before sending a request. Older models may support low through high but not xhigh or max; Haiku 4.5 does not support effort. Forge rejects per-message effort on known unsupported models and adds Anthropic’s required beta header when it is used. A translated continuation with signed Claude 4.5 thinking uses that generation’s manual thinking format and requires enough output tokens for its minimum budget. The Models endpoint lists models available to that key. Your application chooses a model name; Forge chooses the eligible destination. For provider-specific features, check the selected provider and model’s capabilities before relying on them. Under Access profiles → Advanced web search, turn on Intercept model web searches and enter a Tavily API key. Forge stores the key securely; the saved value is never displayed again. You can replace it by entering a new key. Tavily provides API keys in its dashboard. When a supported model requests web search, Forge obtains results from Tavily, applies the profile’s controls, and continues the model response. A request can make up to three searches. Chat Completions, Responses, and Anthropic Messages clients can use this setting, including Messages clients routed to Bedrock Converse. Search is off by default and uses the Tavily account associated with the key you supply. Forge records model tokens and spend separately from Tavily search credits. Search result content and the API key are not included in the search-operation record. Search-enabled streaming requests complete the search and model continuation before Forge emits the final streamed response. They can therefore take longer to produce the first event. Forge does not retry an uncertain search result. Enabling search adds tool guidance to supported model requests, including requests where the model chooses not to search. Those extra model input tokens count toward usage and spend.

Caching and failover

Advanced caching can keep a session on the same destination to improve provider prompt-cache reuse. Exact response caching stores eligible repeated text responses for a selected expiry; a hit uses no new provider tokens and shows estimated savings separately from spend. Current policy and budget checks still run for cache hits. Streaming, tool, and stateful requests do not use exact response caching. Advanced failover lets you choose which errors may try another eligible destination. Destinations at a later route priority are fallback tiers. Forge records every provider attempt and its result in Usage. A stream cannot move to another provider after its response has begun. After a streamed response is admitted, Forge sends keepalive comments during silence. It also times out a provider that sends heartbeats or metadata without model content for too long.

Experiments

Use Gateway → Experiments to sample requests from an access profile without changing the answer returned by the primary route. Mirror sends a copy to one candidate destination. Compare also evaluates the primary and candidate answers with a judge model. Set a sample rate and spend cap. Choose View results on an experiment to open its results panel, see usage for the live traffic and test model, and inspect recent samples. Sampling details contains the eligible, sampled, completed, and skipped counts. Results show whether the test model is cheaper or more expensive than live traffic. The percentage compares matched requests with two completed calls and known costs, excluding judge spend. The usage table retains all recorded charges, including known charges from failed requests. When live traffic has no cost, results show the dollar amounts without a percentage. The primary, candidate, and judge calls each have their own token and spend records. Candidate and judge costs count against the experiment cap and appear in Usage and Spend. Open a sample to see its request and linked Live session. Choose a candidate route that supports the provider operation Forge will execute, including Converse for a translated Bedrock request.

Find MCP tools

For an MCP client with many available tools, Forge can search the entitled tool catalog before the client loads individual tool schemas. Tool calls still use the actor’s existing access and policy checks.

Usage, Spend, and budgets

The same four organization summaries remain above every gateway tab: requests, tokens, spend, and active gateway keys. Traffic summaries cover the last 30 days. Usage has request-level filters for time, outcome, provider, model, endpoint, and gateway key. These filters apply to the request table. Open a request to see provider attempts, cache savings, or its Live session. The table is paginated for large histories. Spend shows cost breakdowns by provider, model, and subject with the same time and destination context. Gateway keys can have hard spend or token limits and a separate soft spend alert. A soft alert notifies administrators when recorded spend crosses its threshold; it does not block a request. Model cost comes from provider usage, configured route pricing, or Forge’s pricing catalog. When a provider reports a served service tier, Forge uses that tier’s rate, including priority pricing where available. Unknown costs remain unknown rather than being counted as zero. Provider cache savings and exact response-cache savings are estimates shown separately from billed spend.