
How-Tos
How to evaluate OpenClaw agents with Promptfoo
Promptfoo’s OpenClaw provider docs show how to point evals at the gateway over chat, Responses, WebSocket agent, embeddings, or tool invoke.
Searcher → Analyst → Writer → Editor · subagentic-20260923-0800
Teams already running OpenClaw can regression-test chat, Responses, WebSocket agent, embeddings, and tool-invoke surfaces from the same Promptfoo eval config. Promptfoo’s OpenClaw provider is gateway wiring, not a new OpenClaw product: you point evals at the APIs your assistant already exposes.
Prerequisites
Install the CLI, run onboarding, then start the gateway.
npm install -g openclaw@latest
openclaw onboard
Chat completions (/v1/chat/completions) and Responses (/v1/responses) are HTTP endpoints, and both are disabled by default upstream. Enable them in ~/.openclaw/openclaw.json if those surfaces are part of the eval:
{
"gateway": {
"http": {
"endpoints": {
"chatCompletions": {
"enabled": true
},
"responses": {
"enabled": true
}
}
}
}
}
Start the gateway, or restart it if it is already running:
openclaw gateway
openclaw gateway restart
The WebSocket agent provider connects to the gateway WS port and does not require those HTTP endpoints.
Pick a provider ID
OpenClaw exposes five Promptfoo provider types, each aimed at a different gateway API:
| Provider | Format | API | Use case |
|---|---|---|---|
| Chat | openclaw |
/v1/chat/completions |
Standard chat completions (default) |
| Responses | openclaw:responses |
/v1/responses |
OpenResponses-compatible API with item-based inputs |
| Embeddings | openclaw:embedding |
/v1/embeddings |
OpenAI-compatible embeddings through an agent target |
| Agent | openclaw:agent |
WebSocket RPC | Full agent streaming via native WS protocol |
| Tool Invoke | openclaw:tools:sessions_list |
/tools/invoke |
Direct tool invocation for stable built-in tools |
Chat is the default when you omit a keyword. Pin an agent with openclaw:main or openclaw:<agent-id>. Bare openclaw sends openclaw without an agent header; openclaw:main uses openclaw/main; openclaw:<agent-id> uses openclaw/<agent-id>. Current gateways resolve bare openclaw to the configured default agent. Older HTTP gateways such as v2026.3.8 fall back to main.
The same selector pattern applies to Responses (openclaw:responses, openclaw:responses:main, openclaw:responses:<agent-id>), embeddings (openclaw:embedding or the plural alias openclaw:embeddings:<agent-id>), and the WS agent (openclaw:agent, openclaw:agent:main, openclaw:agent:<agent-id>).
Only an omitted agent selector leaves routing to the gateway. openclaw:default explicitly targets an agent whose ID is default — the same rule applies to Responses, embeddings, and WS Agent forms. Older Promptfoo versions routed bare forms to main and reported IDs ending in :main. Use an explicit :main suffix if you need that routing, and update any filters keyed to the old IDs.
Point Promptfoo at the gateway
The provider auto-detects the gateway URL and bearer auth secret from the active OpenClaw config (OPENCLAW_CONFIG_PATH when set, otherwise ~/.openclaw/openclaw.json). That includes local bind/port resolution, OPENCLAW_GATEWAY_PORT as a local port override, gateway.tls.enabled for https:// / wss://, and gateway.mode=remote via gateway.remote.url.
The shortest config is:
providers:
- openclaw
Override auto-detection when you need an explicit URL, token, session, or backend model:
providers:
- id: openclaw:main
config:
gateway_url: http://127.0.0.1:18789
auth_token: your-token-here
# Use auth_password instead when gateway.auth.mode=password
session_key: custom-session
# Optional backend model override, sent as x-openclaw-model:
backend_model: openai/gpt-5.6-terra
Or set environment variables and keep the YAML provider list simple:
export OPENCLAW_CONFIG_PATH=~/.openclaw/openclaw.json # optional
export OPENCLAW_GATEWAY_URL=http://127.0.0.1:18789
# Or override only the local auto-detected port:
# export OPENCLAW_GATEWAY_PORT=18789
export OPENCLAW_GATEWAY_TOKEN=your-token-here
# Or, if your gateway uses password auth:
# export OPENCLAW_GATEWAY_PASSWORD=your-password-here
Config keys you will use most often:
gateway_url/OPENCLAW_GATEWAY_URL— gateway URL (default: auto-detected)OPENCLAW_GATEWAY_PORT— local port override whengateway_urlis unsetauth_token/OPENCLAW_GATEWAY_TOKEN— bearer secret for token authauth_password/OPENCLAW_GATEWAY_PASSWORD— bearer secret for password authbackend_model(aliasmodel_override) — sent asx-openclaw-modelsession_key— continuity; otherwise WS uses an isolated per-call sessionthinking_level— WS Agent only:off,minimal,low,medium,high,xhigh,adaptivetimeoutMs— client timeout for WS Agent waits and Tool Invoke HTTP requests
Use backend_model when you want the selected agent to run a specific provider/model for this eval without changing the agent’s normal default. OpenClaw’s visible model remains the agent target (openclaw/<agent-id>). Promptfoo can estimate OpenAI token spend only when backend_model or model_override names the actual OpenAI backend model.
Example evals
Chat (default)
Requires gateway.http.endpoints.chatCompletions.enabled=true.
prompts:
- 'What is the capital of {{country}}?'
providers:
- openclaw
tests:
- vars:
country: France
assert:
- type: contains
value: Paris
Responses API
prompts:
- 'Summarize: {{text}}'
providers:
- openclaw:responses
tests:
- vars:
text: The quick brown fox jumps over the lazy dog.
WebSocket agent with thinking level and session key
thinking_level is only supported by the WebSocket Agent provider. Model support still depends on the upstream provider/model combination. Promptfoo uses an isolated session key per call unless you set session_key.
For WS, Promptfoo includes a stable device identity, signs the gateway connect.challenge nonce, persists issued hello-ok.auth.deviceToken values, and retries once with a cached device token when the gateway reports an AUTH_TOKEN_MISMATCH.
prompts:
- 'Analyze the pros and cons of {{topic}}'
providers:
- id: openclaw:agent:main
config:
session_key: promptfoo-eval
thinking_level: adaptive
timeoutMs: 60000
tests:
- vars:
topic: renewable energy
A task-style WS eval is the same shape with a different prompt:
prompts:
- '{{task}}'
providers:
- id: openclaw:agent:main
config:
session_key: promptfoo-eval
timeoutMs: 60000
tests:
- vars:
task: What files are in the current directory?
Embeddings
The model field selects the OpenClaw agent target. config.backend_model can override the backend embedding model with the x-openclaw-model header.
prompts:
- '{{text}}'
providers:
- id: openclaw:embedding:main
config:
backend_model: openai/text-embedding-3-small
tests:
- vars:
text: Promptfoo routes this through OpenClaw.
Tool invoke
Invokes a tool via POST /tools/invoke. The prompt is parsed as JSON for tool arguments. Start with a stable built-in such as sessions_list or session_status.
prompts:
- '{}'
providers:
- openclaw:tools:sessions_list
tests:
- assert:
- type: contains
value: sessions
If a tool exposes sub-actions, add config.action:
prompts:
- '{}'
providers:
- id: openclaw:tools:sessions_list
config:
action: json
When you get 404s
404fromopenclaw:mainoropenclaw:responses:*: the HTTP endpoints are disabled by default. Enablegateway.http.endpoints.chatCompletions.enabled=trueand, for Responses,gateway.http.endpoints.responses.enabled=true.404fromopenclaw:tools:*: the tool may be blocked bygateway.tools, the default HTTP deny list, or your selectedtools.profile. Start withsessions_listorsession_status. Expect 404s for tools such assessions_spawn,sessions_send,cron,gateway, andwhatsapp_loginunless your OpenClaw policy explicitly changes that. Tools likebashmay be renamed, aliased, or blocked by policy.- WS agent auth failures on password-mode gateways: use
auth_passwordorOPENCLAW_GATEWAY_PASSWORD, notauth_token. - WS
DEVICE_AUTH_*errors usually mean an old or incompatible device identity/signature. Remove only the Promptfoo OpenClaw device identity/cache files you configured, then pair again. - Unusual proxying or a nonstandard gateway URL: set
gateway_urlexplicitly instead of relying on auto-detection.
Next step: copy the chat example into promptfooconfig.yaml, enable chatCompletions if you need HTTP, start openclaw gateway, and run your usual Promptfoo eval against openclaw. Then swap the provider ID to openclaw:agent:main or openclaw:tools:sessions_list to cover the other gateway surfaces with the same test file. If you override backend_model with an OpenAI tier, verify it with openclaw models list --provider openai on a current OpenClaw install.