milpa / ai-gateway
Dual-provider LLM gateway for the Milpa PHP framework: OpenAI + Anthropic chat-completions client with tool-call format translation, an agentic tool-use loop, and a ToolRegistry facade for LLM/MCP exposure.
Requires
- php: >=8.3
- guzzlehttp/guzzle: ^7.10
- guzzlehttp/psr7: ^2.7
- milpa/core: ^0.6.2 || ^0.7
- milpa/tool-runtime: ^0.9 || ^0.10
- psr/http-client: ^1.0
- psr/http-factory: ^1.1
- psr/http-message: ^2.0
- psr/log: ^3
Requires (Dev)
- friendsofphp/php-cs-fixer: ^3.65
- phpstan/phpdoc-parser: ^2.3
- phpstan/phpstan: ^2.1
- phpunit/phpunit: ^11.5
README
Milpa AI Gateway
A dual-provider LLM gateway for the Milpa PHP framework — one client for OpenAI and Anthropic chat completions, translating each provider's tool-call wire format to and from a single shape, plus an agentic tool-use loop that drives a
milpa/tool-runtimeToolRegistry(resolve → validate → authorize → execute → audit) until the model is done.
milpa/ai-gateway is the LLM tier of Milpa: the piece that turns a milpa/tool-runtime
ToolRegistry into something a model can actually drive. LlmService implements
milpa/core's LlmServiceInterface seam against two concrete providers — OpenAI and
Anthropic — so callers write one message shape and one tool-call shape regardless of which
provider answers. AgentOrchestrator runs the loop every agent needs: ask the model, execute
whatever tools it asks for through the registry pipeline, feed the results back, repeat until
the model returns a final answer or a step budget runs out. No product coupling, no
Telegram/HTTP-specific code — those live in your host application.
Install
composer require milpa/ai-gateway
Quick example
Register a tool on a ToolRegistry (from milpa/tool-runtime), wrap it in McpClientService,
and hand both to AgentOrchestrator along with an LlmService:
use Milpa\AiGateway\AgentOrchestrator; use Milpa\AiGateway\LlmService; use Milpa\AiGateway\McpClientService; use Milpa\ToolRuntime\ToolRegistry; use Psr\Log\NullLogger; $registry = new ToolRegistry(new NullLogger()); $registry->register( 'get_time', 'Get the current time', [], fn () => ['time' => '12:00 PM'], ); $mcpClient = new McpClientService($registry); $llm = new LlmService(apiKey: getenv('OPENAI_API_KEY'), model: 'gpt-4o', provider: 'openai'); // LlmService talks HTTP through PSR-18 (Psr\Http\Client\ClientInterface), defaulting to a // Guzzle client when none is injected — see "Bringing your own HTTP client" below. $orchestrator = new AgentOrchestrator($llm, $mcpClient); echo $orchestrator->run('What time is it?'); // -> asks the model, the model requests `get_time`, AgentOrchestrator executes it through // the registry, feeds the result back, and returns the model's final answer.
Swap provider: 'anthropic' and a Claude model name (or let a claude model name in
$model select it automatically) to point the same call at Anthropic instead — LlmService
translates the tool list, the message history, and the tool-call response to and from
Anthropic's shape internally, so AgentOrchestrator and McpClientService never see a
provider-specific format.
The agent loop
AgentOrchestrator::run() alternates between two calls until the model is done or
$maxSteps (default 20) is reached:
- Ask —
LlmService::generateResponse()sends the running message history plus the registry's tool summaries (McpClientService::getToolSummaries()) to the provider and returns a single OpenAI-shaped assistant message. - Act — if that message carries
tool_calls, each one is executed viaMcpClientService::callTool(), which runs it through the fullToolRegistrypipeline (validate → authorize → execute → audit) under whateverToolContextwas set withsetToolContext(). The result — rendered through aRendererRegistrywhen one is configured, JSON otherwise — is fed back into the message history as atoolmessage, and the loop repeats.
If a tool result requires confirmation or is blocked by policy, the loop stops immediately and returns that outcome instead of continuing — the caller (a chat handler, a CLI, a bot) is responsible for the confirm/cancel round trip on the next user turn.
Stopping the loop before it acts: ToolCallGate
The loop above executes what the model asked for. ToolCallGate is the seam that lets somebody
decide before it does:
interface ToolCallGate { // The reason this call does not proceed, or null if it does. public function refuse(string $tool, array $arguments): ?string; public function record(string $tool, array $arguments, array $result): void; }
refuse() runs before the call and record() after it. The two halves are not symmetric on
purpose: the gate sees the intention, and the record sees the outcome, and resuming a long
session needs both — with only the intention, an agent picking up where it left off knows what its
past self was going to try but not whether it worked, so it repeats work already done or work that
already failed.
The gate is an interface here and nothing else: this package brings no policy of its own. Who may
run what, and whether a human is asked first, belongs to whoever holds the session — milpa/agent
implements exactly this seam with SessionPolicy, per-session permissions and human questions that
survive the process.
A gate that refuses returns the reason as a string, not a boolean. Whoever receives the refusal needs to know why in order to do something about it, and that information was already there.
The table: OptionTable
Since 0.5 the loop can operate against a table — the set of options the agent currently has in front of it — through one port with two questions that look alike and are not:
interface OptionTable { public function remove(string $option, string $code, ?string $message = null): void; public function removed(): array; // the PROJECTION: what the model sees public function wasRemoved(string $option): bool; // the FACT: did an authority declare it gone? }
The catalogue the model sees is re-derived every step from the registry minus removed() — it
is a projection, never a snapshot. And a refusal that removed the option no longer ends the run:
there is nothing left to route around, so the reason goes back to the model and the loop continues
with a different table. Both came out of measurement: a frozen catalogue made "did the agent re-read
the world?" unanswerable, and a run-ending refusal turned every denial into a shutdown (0 of 32
runs recovered).
SecondOpinionGate also stopped being quiet anywhere: an empty answer, a verdict-less answer and a
failed judge each say so with their own cause, a denial leaves a witness on stderr outside the
stream it writes to, and an ALLOW logs nothing — approving silently is correct; failing silently
was being counted as approval.
Provider translation
LlmService speaks one shape to its callers — OpenAI's messages / tool_calls — and
translates both directions for Anthropic:
- Outbound:
systemmessages become Anthropic's top-levelsystemparameter;toolrole messages becomeusermessages carrying atool_resultcontent block; an assistant message withtool_callsbecomestool_usecontent blocks. Tool summaries are reshaped from{name, description, inputSchema}to Anthropic's{name, description, input_schema}, with an emptypropertiesobject substituted where a tool declares none (Anthropic requires a non-empty schema object, not an empty array). - Inbound: Anthropic's
content: [{type: text, ...}, {type: tool_use, ...}]array is flattened back into a single OpenAI-shaped assistant message (content+tool_calls), soAgentOrchestratorruns identical logic regardless of provider.
Bringing your own HTTP client (PSR-18)
LlmService's constructor accepts a PSR-18 ClientInterface (plus PSR-17 request/stream
factories) — inject your own for connection pooling, retry/circuit-breaker middleware, or
tests that assert on the outgoing request without touching the network:
use Milpa\AiGateway\LlmService; $llm = new LlmService( apiKey: getenv('ANTHROPIC_API_KEY'), model: 'claude-3-5-sonnet-20241022', provider: 'anthropic', logger: $logger, // Psr\Log\LoggerInterface; optional httpClient: $yourPsr18Client, // Psr\Http\Client\ClientInterface; omit for Guzzle requestFactory: $yourFactory, // Psr\Http\Message\RequestFactoryInterface; optional streamFactory: $yourFactory, // Psr\Http\Message\StreamFactoryInterface; optional );
When httpClient is omitted, LlmService builds a Guzzle client with a 600s timeout
shared by both providers. That number used to be OpenAI-only-60s / Anthropic-only-600s (a
per-request Guzzle option on the Anthropic call, since Claude tool-use responses can run
long) — PSR-18's sendRequest() takes only a RequestInterface, with no per-call options
bag, so a per-provider timeout has no seam to hang off anymore. The default now simply
covers the slower case for both. Inject your own ClientInterface if you need the tighter
OpenAI-side timeout back.
What lives where
| Layer | Package | Owns |
|---|---|---|
| Contracts | milpa/tool-runtime |
LlmServiceInterface — the seam LlmService implements. |
| Tool execution | milpa/tool-runtime |
ToolRegistry, ToolContext, ToolResult, channel rendering — the pipeline McpClientService and AgentOrchestrator drive. |
| Gateway | milpa/ai-gateway (this package) |
The concrete LlmService (OpenAI + Anthropic, format translation both ways), McpClientService (registry facade), and AgentOrchestrator (the ask-act loop). |
| Your app | your host / plugins | API keys and secrets management, the PSR-3 logger you wire in, and any channel-specific glue (Telegram, web chat, CLI) around AgentOrchestrator::run(). |
Requirements
- PHP ≥ 8.3
milpa/core^0.6milpa/tool-runtime^0.5guzzlehttp/guzzle^7.10 — the default PSR-18 implementationLlmServicefalls back to when noClientInterfaceis injected (also bringsguzzlehttp/psr7, used as the default PSR-17 factory)psr/http-client,psr/http-factory,psr/http-message— the interfacesLlmService's constructor is typed againstpsr/log^3
Security note
LlmService can log provider request/response detail at debug level, including a slice of
the raw LLM response body — never enable that logging in production. See
SECURITY.md for the specifics.
Documentation
Full API reference: getmilpa.github.io/ai-gateway — generated straight from the source DocBlocks and dressed with the Milpa design system.
Contributing
Contributions are welcome — see CONTRIBUTING.md. Please report security issues via SECURITY.md, and note that this project follows a Code of Conduct.
License
Apache-2.0 © Rodrigo Vicente - TeamX Agency.
Milpa is designed, built, and maintained by Rodrigo Vicente - TeamX Agency.