> ## Documentation Index
> Fetch the complete documentation index at: https://docs.forge.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Advanced Gateway controls

> Configure search, caching, failover, and experiments for governed model traffic.

These controls are optional. Start with a provider, access profile, and gateway key in [Gateway](/secure/llm-gateway), then open the profile's **Advanced** sections when you need them.

## Use one profile across model APIs

A gateway key is assigned to an access profile, not to a client API. An entitled application can call Forge with Chat Completions, Responses, or Anthropic Messages while Forge selects the configured provider operation and translates supported text and tool requests. Bedrock Converse and Gemini Generate Content are also supported client and provider operations. The same profile, identity, policy, and budget apply to translated requests. Unsupported fields or content are rejected rather than silently removed. Streaming responses use the client's API format.

For incomplete Chat tool histories, Forge records a bounded repair and supplies missing results before policy checks and the provider call. Orphan results are removed, and the last duplicate result within a tool exchange is kept. Results are never moved across conversation turns.

For Claude models, Forge checks effort levels against the selected model before sending a request. Older models may support `low` through `high` but not `xhigh` or `max`; Haiku 4.5 does not support effort. Forge rejects per-message effort on known unsupported models and adds Anthropic's required beta header when it is used. A translated continuation with signed Claude 4.5 thinking uses that generation's manual thinking format and requires enough output tokens for its minimum budget.

The **Models** endpoint lists models available to that key. Your application chooses a model name; Forge chooses the eligible destination. For provider-specific features, check the selected provider and model's capabilities before relying on them.

## Web search

Under **Access profiles → Advanced web search**, turn on **Intercept model web searches** and enter a **Tavily API key**. Forge stores the key securely; the saved value is never displayed again. You can replace it by entering a new key. [Tavily provides API keys in its dashboard](https://help.tavily.com/articles/9170796666-how-can-i-create-an-api-key).

When a supported model requests web search, Forge obtains results from Tavily, applies the profile's controls, and continues the model response. A request can make up to three searches. Chat Completions, Responses, and Anthropic Messages clients can use this setting, including Messages clients routed to Bedrock Converse. Search is off by default and uses the Tavily account associated with the key you supply. Forge records model tokens and spend separately from Tavily search credits. Search result content and the API key are not included in the search-operation record.

Search-enabled streaming requests complete the search and model continuation before Forge emits the final streamed response. They can therefore take longer to produce the first event. Forge does not retry an uncertain search result.

Enabling search adds tool guidance to supported model requests, including requests where the model chooses not to search. Those extra model input tokens count toward usage and spend.

## Caching and failover

**Advanced caching** can keep a session on the same destination to improve provider prompt-cache reuse. Exact response caching stores eligible repeated text responses for a selected expiry; a hit uses no new provider tokens and shows estimated savings separately from spend. Current policy and budget checks still run for cache hits. Streaming, tool, and stateful requests do not use exact response caching.

**Advanced failover** lets you choose which errors may try another eligible destination. Destinations at a later route priority are fallback tiers. Forge records every provider attempt and its result in **Usage**. A stream cannot move to another provider after its response has begun.

After a streamed response is admitted, Forge sends keepalive comments during silence. It also times out a provider that sends heartbeats or metadata without model content for too long.

## Experiments

Use **Gateway → Experiments** to sample requests from an access profile without changing the answer returned by the primary route. **Mirror** sends a copy to one candidate destination. **Compare** also evaluates the primary and candidate answers with a judge model. Set a sample rate and spend cap. Choose **View results** on an experiment to open its results panel, see usage for the live traffic and test model, and inspect recent samples. **Sampling details** contains the eligible, sampled, completed, and skipped counts.

Results show whether the test model is cheaper or more expensive than live traffic. The percentage compares matched requests with two completed calls and known costs, excluding judge spend. The usage table retains all recorded charges, including known charges from failed requests. When live traffic has no cost, results show the dollar amounts without a percentage.

The primary, candidate, and judge calls each have their own token and spend records. Candidate and judge costs count against the experiment cap and appear in **Usage** and **Spend**. Open a sample to see its request and linked Live session. Choose a candidate route that supports the provider operation Forge will execute, including Converse for a translated Bedrock request.

## Find MCP tools

For an MCP client with many available tools, Forge can [search the entitled tool catalog](/developer/tool-reference#search-a-large-tool-catalog) before the client loads individual tool schemas. Tool calls still use the actor's existing access and policy checks.

## Usage, Spend, and budgets

The same four organization summaries remain above every gateway tab: requests, tokens, spend, and active gateway keys. Traffic summaries cover the last 30 days.

**Usage** has request-level filters for time, outcome, provider, model, endpoint, and gateway key. These filters apply to the request table. Open a request to see provider attempts, cache savings, or its Live session. The table is paginated for large histories. **Spend** shows cost breakdowns by provider, model, and subject with the same time and destination context.

Gateway keys can have hard spend or token limits and a separate soft spend alert. A soft alert notifies administrators when recorded spend crosses its threshold; it does not block a request. Model cost comes from provider usage, configured route pricing, or Forge's pricing catalog. When a provider reports a served service tier, Forge uses that tier's rate, including priority pricing where available. Unknown costs remain unknown rather than being counted as zero. Provider cache savings and exact response-cache savings are estimates shown separately from billed spend.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.