Skip to main content
Codex sends model requests through the Responses API. In Forge, create an access profile with a model the user or service account may call, then issue a gateway key for that identity. Copy the gateway endpoint from Gateway. The endpoint already includes /llm-gateway/v1.

Configure inference

Add the following to your user-level Codex config.toml on macOS or Linux at ~/.codex/config.toml, or on Windows at %USERPROFILE%\.codex\config.toml. Replace the model and URL with those shown in Forge. If Codex does not recognize the model, ask your administrator for a matching Codex model catalog. Set model_catalog_json to that file’s absolute path before the first TOML table.
Supply FORGE_GATEWAY_KEY to the process that starts Codex using your normal secret delivery method. Do not put the key in config.toml. Restart the desktop app after changing the file. The CLI and desktop app read the same local configuration. For machine-wide rollout and custom model catalogs, follow the official Codex gateway guide. Start a new Codex task and ask it to reply with a short fixed phrase. Check that the reply arrives and the task finishes. In Forge, open Gateway → Usage and check the model, result, token counts, and the linked Live session. Then ask Codex to read a file and use its contents in its reply. Send a follow-up question about the same file. These steps check a completed stream, a tool call, its result, and conversation continuation with your model.

Connect Forge MCP

Inference and MCP have separate credentials. To give Codex access to Forge’s authorized inventory, investigation, and governance tools, add the Forge MCP server and complete its OAuth sign-in:
For an approved upstream MCP server in the Forge registry, use forge mcp install SERVER --client codex instead. See Forge MCP for the two endpoints, permissions, and installation flow. The Codex MCP guide describes the shared CLI and desktop configuration.