Agent Configuration
All the settings available when creating or updating your Omniagent
When you create an Omniagent, you define everything about how it behaves — its identity, resources, voice, language, and model settings. These apply consistently across every channel the agent is deployed to.
This page is the complete reference for all agent-level settings.
Setting and updating configuration
Create and update agent configuration
Identity
Name, companion, and language
Mode
Conversation, or puppeteer mode where the agent speaks your lines
Knowledge
Attach document collections and FAQs
Tools
Actions your agent can take during conversations
Tags
Label your agent for filtering and analytics
Idle timeout
Keep sessions open indefinitely
Voice
Choose the voice for your agent's speech output
Web search
Let the agent search the web during conversations
Provider settings
Temperature, turn detection, and noise reduction
Cascade architecture
How agents on a Cascade key behave differently
Channel overrides
Override settings for a specific channel
Setting and updating configuration
Every field on this page can be set when you create the agent or updated later. Only the fields you include are changed — everything else stays the same.
Create an agent with your initial configuration:
curl -X POST https://companion-api.napster.com/public/agents \
-H "X-Api-Key: $NAPSTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"companionId": "comp_abc123",
"name": "Support Agent",
"voiceId": "alloy",
"providerSettings": {
"temperature": 0.7
}
}'Update any field later with a PATCH request — only include what you want to change:
curl -X PATCH https://companion-api.napster.com/public/agents/agent_abc123 \
-H "X-Api-Key: $NAPSTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"language": "es",
"voiceId": "shimmer"
}'Identity
These fields define who your agent is.
Name
A label to identify the agent in your project — for example, "Support Agent", "Sales Bot", or "Onboarding Assistant". Required when you create an agent.
Companion
The companionId is required and determines the agent's visual appearance, personality, and behavior. See Companion for how to pick or create one.
Language
When you set language, the agent strictly follows that language for the entire session — it will not switch even if the user speaks in a different language. Use ISO language codes — en, es, fr, de, etc.
If you omit language, the agent defaults to English but can switch to a different language when the user asks.
{
"companionId": "comp_abc123",
"name": "Support Agent",
"language": "es"
}Supported languages (OpenAI Realtime)
On the Realtime architecture (Azure OpenAI or OpenAI), the agent supports the following languages. Cascade, Humain, and Microsoft Foundry keys support different languages. A digital twin speaks 32 of these languages — see Digital Twins — Languages.
| Code | Language | Code | Language | Code | Language |
|---|---|---|---|---|---|
af | Afrikaans | he | Hebrew | no | Norwegian |
ar | Arabic | hi | Hindi | fa | Persian |
hy | Armenian | hu | Hungarian | pl | Polish |
az | Azerbaijani | is | Icelandic | pt | Portuguese |
be | Belarusian | id | Indonesian | ro | Romanian |
bs | Bosnian | it | Italian | ru | Russian |
bg | Bulgarian | ja | Japanese | sr | Serbian |
ca | Catalan | kn | Kannada | sk | Slovak |
zh | Chinese | kk | Kazakh | sl | Slovenian |
hr | Croatian | ko | Korean | es | Spanish |
cs | Czech | lv | Latvian | sw | Swahili |
da | Danish | lt | Lithuanian | sv | Swedish |
nl | Dutch | mk | Macedonian | tl | Tagalog |
en | English | ms | Malay | ta | Tamil |
et | Estonian | mr | Marathi | th | Thai |
fi | Finnish | mi | Maori | tr | Turkish |
fr | French | ne | Nepali | uk | Ukrainian |
gl | Galician | ur | Urdu | ||
de | German | vi | Vietnamese | ||
el | Greek | cy | Welsh |
Mode
The mode field sets how the agent speaks:
| Mode | Behavior |
|---|---|
conversation | The default. The agent listens to the user and answers out loud. |
puppeteer | The agent speaks only the lines your client sends with the talk command. Use it for scripted presentations, guided tours, or when your own system decides every line. |
{
"mode": "puppeteer"
}Set mode on the agent, or per session on POST /public/connections and POST /public/ws-connections. Sessions started from an agent use the agent's mode.
Puppeteer mode requires an API key that uses the Cascade architecture. With any other key, opening a session in puppeteer mode returns 400 UnsupportedSessionMode. A Cascade key with only text-to-speech is enough. If the key also has speech recognition and a language model, the agent still listens and replies, but its replies arrive as text in message_received instead of being spoken.
Agents in puppeteer mode can't use phone channels. Adding a SIP or VoIP channel config returns 400 TelephonyChannelNotAllowed.
Knowledge
Knowledge gives your agent domain expertise through two types of resources.
Documents
The knowledgeBaseId field attaches a knowledge collection to the agent. The agent retrieves relevant documents from this collection during conversations to ground its responses in your content. See Files for details.
FAQs
The faqCollections field takes an array with the ID of one FAQ collection. When a user asks a question that matches one of your FAQs, the agent responds with the answer you defined rather than generating one. See FAQs for details.
{
"knowledgeBaseId": "kb_product_docs",
"faqCollections": ["faq_support"]
}You can attach at most one FAQ collection. Passing more than one ID returns a 400 validation error.
Tools
The functions field accepts an array of tool IDs. These define what your agent can do during conversations — look up data, trigger workflows, generate content. You can attach as many tools as you need.
Tools must be created before attaching them to the agent. See Tools for details.
{
"functions": ["fn_order_status", "fn_returns", "fn_book_appointment"]
}Tags
The tags field accepts string-to-string key-value pairs for labeling the agent. Use tags to categorize agents by source, environment, team, or any dimension relevant to your application. Tags are stored on the agent and returned on all sessions created from it.
{
"tags": {
"env": "production",
"team": "support",
"source": "ios-app"
}
}Tag keys are strings with a maximum length of 64 characters. Values are strings with a maximum length of 512 characters.
Idle timeout
By default, a session closes after 3 minutes without audio or messages, with closeReason: "idle_timeout". The client is warned 60, 30, and 10 seconds before — see Session errors and shutdown. Set disableIdleTimeout to true on the agent to keep all sessions open indefinitely, even if no audio or messages are exchanged.
{
"disableIdleTimeout": true
}Voice
The voiceId field sets the voice for the agent's speech output. It is required when using a catalog or custom companion. When using a digital twin, voiceId is not required — the voice comes from the voiceUrl provided during digital twin creation.
| Provider | Supported voices |
|---|---|
| Azure OpenAI | alloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar |
{
"voiceId": "alloy"
}Web search
When useWebSearch is true, the agent can search the web during conversations to look up real-time information beyond its knowledge base — current events, product details, public data, and more.
{
"useWebSearch": false
}Web search is enabled by default for agents created through the API (agents created in the dashboard start with it off). Pass "useWebSearch": false when creating or updating an agent to turn it off. Without web search, the agent still answers from its model's general knowledge as well as its knowledge base and FAQs. You can also override the setting per channel via channel overrides.
Provider settings
The providerSettings object controls model behavior and audio processing.
The provider settings below apply to keys on the Realtime architecture (the Azure OpenAI and OpenAI providers). On a Cascade key, only some of them apply — see Cascade architecture. If your key uses the Other architecture (HUMAIN or Microsoft Foundry), contact us for platform-specific configuration details. For Foundry, see Connect a Microsoft Foundry agent.
| Field | Type | Description |
|---|---|---|
temperature | float | Controls randomness in the agent's responses. Lower values (e.g. 0.3) produce more focused output; higher values (e.g. 0.9) increase variety. |
instructions | string | Overrides the companion's system prompt. Use this to replace the default prompt with custom instructions without modifying the companion itself. |
turnDetection | object | Voice Activity Detection (VAD) configuration. Controls how the system detects when the user starts and stops speaking. |
noiseReduction | object | Noise reduction configuration. Reduces background noise from the user's microphone input. |
{
"providerSettings": {
"temperature": 0.7,
"turnDetection": {
"threshold": 0.9,
"prefix_padding_ms": 400,
"silence_duration_ms": 500
},
"noiseReduction": {
"type": "nearField"
}
}
}Temperature
The temperature field controls how varied vs. consistent the agent's responses are. Lower values keep the agent on-script; higher values let it improvise.
| Range | Behavior | Best for |
|---|---|---|
0.0 – 0.3 | Strict, deterministic. The agent says the same thing the same way. | Classifiers, scripted flows, agents that should never deviate |
0.3 – 0.6 | Focused but natural. Small variation in phrasing, same substance. | Customer support, technical Q&A, documentation assistants |
0.6 – 0.9 | Warm and varied. Different phrasing every time, on-message. | Sales, onboarding, conversational hosts |
1.0 – 1.2 | Loose and creative. Noticeable variety, occasionally off-script. | Brainstorming partners, casual chat |
1.5+ | Highly creative. Surprising, sometimes unpredictable. | Creative writing, ideation, experimentation |
For production voice agents, 0.5 to 0.7 is the safe range. Higher values produce more interesting language but also more drift — answers vary between sessions, occasionally hallucinate, occasionally drop instructions.
Turn detection
The turnDetection object configures Voice Activity Detection (VAD) — how the system decides when the user has started and stopped speaking. This directly affects interruption behavior and response timing.
| Field | Type | Description |
|---|---|---|
threshold | float | Activation threshold for detecting speech. Higher values (e.g. 0.8) require louder, clearer speech to trigger; lower values (e.g. 0.3) are more sensitive. |
prefix_padding_ms | integer | Milliseconds of audio to include before detected speech begins. Prevents clipping the start of an utterance. |
silence_duration_ms | integer | Milliseconds of silence required before the system considers the user done speaking. Lower values (e.g. 200) make the agent respond faster; higher values (e.g. 800) wait longer for the user to continue. |
Noise reduction
The noiseReduction object accepts a single type field:
| Value | Description |
|---|---|
nearField | Optimized for close-range microphones — laptops, headsets, phones held to the ear. |
farField | Optimized for distant microphones — speakerphones, smart speakers, conference room setups. |
If you omit noiseReduction, no noise processing is applied. Choose the mode that matches your user's typical microphone distance for best results.
Cascade architecture
With a Cascade API key, the agent runs on separate speech-recognition, language, and text-to-speech models instead of one realtime model. Tools, knowledge, FAQs, and digital twins work the same way. These features behave differently:
- MCP isn't supported. Opening a session with MCP servers or connectors attached returns
400 UnsupportedProvider. - Only some provider settings apply.
instructionsworks as usual, andturnDetectionapplies when the session starts.temperatureandnoiseReductionare ignored. - Mid-session
set_settingsonly changesinstructionsand tools. - User speech events arrive after the transcript.
speech_startedand the related events arrive together once the user's speech is transcribed, so you can't usespeech_startedto detect the user interrupting. Interrupted responses reportreason: "interrupted"instead of"turn_detected". - Puppeteer mode is available. The agent speaks only the lines your client sends.
Channel overrides
All of the settings above apply across every channel by default. If you need different behavior on a specific channel, you can override providerSettings, functions, faqCollections, and knowledgeBaseId per channel using the channel config.
For example, to use different turn detection settings and attach an additional tool on SIP calls:
curl -X PUT https://companion-api.napster.com/public/agents/agent_abc123/channels/sip \
-H "X-Api-Key: $NAPSTER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"functions": ["fn_order_status", "fn_returns", "fn_transfer_call"],
"providerSettings": {
"turnDetection": {
"threshold": 0.7,
"silence_duration_ms": 800
}
}
}'Channel overrides apply to every session on that channel. If omitted, the agent's base configuration is used. See the WebRTC, WebSockets, VoIP, and SIP guides for channel-specific setup.