Build Your Omniagent

Agent Configuration

All the settings available when creating or updating your Omniagent

When you create an Omniagent, you define everything about how it behaves — its identity, resources, voice, language, and model settings. These apply consistently across every channel the agent is deployed to.

This page is the complete reference for all agent-level settings.

Setting and updating configuration

Every field on this page can be set when you create the agent or updated later. Only the fields you include are changed — everything else stays the same.

Create an agent with your initial configuration:

curl -X POST https://companion-api.napster.com/public/agents \
  -H "X-Api-Key: $NAPSTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "companionId": "comp_abc123",
    "name": "Support Agent",
    "voiceId": "alloy",
    "providerSettings": {
      "temperature": 0.7
    }
  }'

Update any field later with a PATCH request — only include what you want to change:

curl -X PATCH https://companion-api.napster.com/public/agents/agent_abc123 \
  -H "X-Api-Key: $NAPSTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "language": "es",
    "voiceId": "shimmer"
  }'

Identity

These fields define who your agent is.

Name

A label to identify the agent in your project — for example, "Support Agent", "Sales Bot", or "Onboarding Assistant". Required when you create an agent.

Companion

The companionId is required and determines the agent's visual appearance, personality, and behavior. See Companion for how to pick or create one.

Language

When you set language, the agent strictly follows that language for the entire session — it will not switch even if the user speaks in a different language. Use ISO language codes — en, es, fr, de, etc.

If you omit language, the agent defaults to English but can switch to a different language when the user asks.

{
  "companionId": "comp_abc123",
  "name": "Support Agent",
  "language": "es"
}

Supported languages (OpenAI Realtime)

On the Realtime architecture (Azure OpenAI or OpenAI), the agent supports the following languages. Cascade, Humain, and Microsoft Foundry keys support different languages. A digital twin speaks 32 of these languages — see Digital Twins — Languages.

CodeLanguageCodeLanguageCodeLanguage
afAfrikaansheHebrewnoNorwegian
arArabichiHindifaPersian
hyArmenianhuHungarianplPolish
azAzerbaijaniisIcelandicptPortuguese
beBelarusianidIndonesianroRomanian
bsBosnianitItalianruRussian
bgBulgarianjaJapanesesrSerbian
caCatalanknKannadaskSlovak
zhChinesekkKazakhslSlovenian
hrCroatiankoKoreanesSpanish
csCzechlvLatvianswSwahili
daDanishltLithuaniansvSwedish
nlDutchmkMacedoniantlTagalog
enEnglishmsMalaytaTamil
etEstonianmrMarathithThai
fiFinnishmiMaoritrTurkish
frFrenchneNepaliukUkrainian
glGalicianurUrdu
deGermanviVietnamese
elGreekcyWelsh

Mode

The mode field sets how the agent speaks:

ModeBehavior
conversationThe default. The agent listens to the user and answers out loud.
puppeteerThe agent speaks only the lines your client sends with the talk command. Use it for scripted presentations, guided tours, or when your own system decides every line.
{
  "mode": "puppeteer"
}

Set mode on the agent, or per session on POST /public/connections and POST /public/ws-connections. Sessions started from an agent use the agent's mode.

Puppeteer mode requires an API key that uses the Cascade architecture. With any other key, opening a session in puppeteer mode returns 400 UnsupportedSessionMode. A Cascade key with only text-to-speech is enough. If the key also has speech recognition and a language model, the agent still listens and replies, but its replies arrive as text in message_received instead of being spoken.

Agents in puppeteer mode can't use phone channels. Adding a SIP or VoIP channel config returns 400 TelephonyChannelNotAllowed.


Knowledge

Knowledge gives your agent domain expertise through two types of resources.

Documents

The knowledgeBaseId field attaches a knowledge collection to the agent. The agent retrieves relevant documents from this collection during conversations to ground its responses in your content. See Files for details.

FAQs

The faqCollections field takes an array with the ID of one FAQ collection. When a user asks a question that matches one of your FAQs, the agent responds with the answer you defined rather than generating one. See FAQs for details.

{
  "knowledgeBaseId": "kb_product_docs",
  "faqCollections": ["faq_support"]
}

You can attach at most one FAQ collection. Passing more than one ID returns a 400 validation error.


Tools

The functions field accepts an array of tool IDs. These define what your agent can do during conversations — look up data, trigger workflows, generate content. You can attach as many tools as you need.

Tools must be created before attaching them to the agent. See Tools for details.

{
  "functions": ["fn_order_status", "fn_returns", "fn_book_appointment"]
}

Tags

The tags field accepts string-to-string key-value pairs for labeling the agent. Use tags to categorize agents by source, environment, team, or any dimension relevant to your application. Tags are stored on the agent and returned on all sessions created from it.

{
  "tags": {
    "env": "production",
    "team": "support",
    "source": "ios-app"
  }
}

Tag keys are strings with a maximum length of 64 characters. Values are strings with a maximum length of 512 characters.


Idle timeout

By default, a session closes after 3 minutes without audio or messages, with closeReason: "idle_timeout". The client is warned 60, 30, and 10 seconds before — see Session errors and shutdown. Set disableIdleTimeout to true on the agent to keep all sessions open indefinitely, even if no audio or messages are exchanged.

{
  "disableIdleTimeout": true
}

Voice

The voiceId field sets the voice for the agent's speech output. It is required when using a catalog or custom companion. When using a digital twin, voiceId is not required — the voice comes from the voiceUrl provided during digital twin creation.

ProviderSupported voices
Azure OpenAIalloy, ash, ballad, coral, echo, sage, shimmer, verse, marin, cedar
{
  "voiceId": "alloy"
}

When useWebSearch is true, the agent can search the web during conversations to look up real-time information beyond its knowledge base — current events, product details, public data, and more.

{
  "useWebSearch": false
}

Web search is enabled by default for agents created through the API (agents created in the dashboard start with it off). Pass "useWebSearch": false when creating or updating an agent to turn it off. Without web search, the agent still answers from its model's general knowledge as well as its knowledge base and FAQs. You can also override the setting per channel via channel overrides.


Provider settings

The providerSettings object controls model behavior and audio processing.

The provider settings below apply to keys on the Realtime architecture (the Azure OpenAI and OpenAI providers). On a Cascade key, only some of them apply — see Cascade architecture. If your key uses the Other architecture (HUMAIN or Microsoft Foundry), contact us for platform-specific configuration details. For Foundry, see Connect a Microsoft Foundry agent.

FieldTypeDescription
temperaturefloatControls randomness in the agent's responses. Lower values (e.g. 0.3) produce more focused output; higher values (e.g. 0.9) increase variety.
instructionsstringOverrides the companion's system prompt. Use this to replace the default prompt with custom instructions without modifying the companion itself.
turnDetectionobjectVoice Activity Detection (VAD) configuration. Controls how the system detects when the user starts and stops speaking.
noiseReductionobjectNoise reduction configuration. Reduces background noise from the user's microphone input.
{
  "providerSettings": {
    "temperature": 0.7,
    "turnDetection": {
      "threshold": 0.9,
      "prefix_padding_ms": 400,
      "silence_duration_ms": 500
    },
    "noiseReduction": {
      "type": "nearField"
    }
  }
}

Temperature

The temperature field controls how varied vs. consistent the agent's responses are. Lower values keep the agent on-script; higher values let it improvise.

RangeBehaviorBest for
0.0 – 0.3Strict, deterministic. The agent says the same thing the same way.Classifiers, scripted flows, agents that should never deviate
0.3 – 0.6Focused but natural. Small variation in phrasing, same substance.Customer support, technical Q&A, documentation assistants
0.6 – 0.9Warm and varied. Different phrasing every time, on-message.Sales, onboarding, conversational hosts
1.0 – 1.2Loose and creative. Noticeable variety, occasionally off-script.Brainstorming partners, casual chat
1.5+Highly creative. Surprising, sometimes unpredictable.Creative writing, ideation, experimentation

For production voice agents, 0.5 to 0.7 is the safe range. Higher values produce more interesting language but also more drift — answers vary between sessions, occasionally hallucinate, occasionally drop instructions.

Turn detection

The turnDetection object configures Voice Activity Detection (VAD) — how the system decides when the user has started and stopped speaking. This directly affects interruption behavior and response timing.

FieldTypeDescription
thresholdfloatActivation threshold for detecting speech. Higher values (e.g. 0.8) require louder, clearer speech to trigger; lower values (e.g. 0.3) are more sensitive.
prefix_padding_msintegerMilliseconds of audio to include before detected speech begins. Prevents clipping the start of an utterance.
silence_duration_msintegerMilliseconds of silence required before the system considers the user done speaking. Lower values (e.g. 200) make the agent respond faster; higher values (e.g. 800) wait longer for the user to continue.

Noise reduction

The noiseReduction object accepts a single type field:

ValueDescription
nearFieldOptimized for close-range microphones — laptops, headsets, phones held to the ear.
farFieldOptimized for distant microphones — speakerphones, smart speakers, conference room setups.

If you omit noiseReduction, no noise processing is applied. Choose the mode that matches your user's typical microphone distance for best results.

Cascade architecture

With a Cascade API key, the agent runs on separate speech-recognition, language, and text-to-speech models instead of one realtime model. Tools, knowledge, FAQs, and digital twins work the same way. These features behave differently:

  • MCP isn't supported. Opening a session with MCP servers or connectors attached returns 400 UnsupportedProvider.
  • Only some provider settings apply. instructions works as usual, and turnDetection applies when the session starts. temperature and noiseReduction are ignored.
  • Mid-session set_settings only changes instructions and tools.
  • User speech events arrive after the transcript. speech_started and the related events arrive together once the user's speech is transcribed, so you can't use speech_started to detect the user interrupting. Interrupted responses report reason: "interrupted" instead of "turn_detected".
  • Puppeteer mode is available. The agent speaks only the lines your client sends.

Channel overrides

All of the settings above apply across every channel by default. If you need different behavior on a specific channel, you can override providerSettings, functions, faqCollections, and knowledgeBaseId per channel using the channel config.

For example, to use different turn detection settings and attach an additional tool on SIP calls:

curl -X PUT https://companion-api.napster.com/public/agents/agent_abc123/channels/sip \
  -H "X-Api-Key: $NAPSTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "functions": ["fn_order_status", "fn_returns", "fn_transfer_call"],
    "providerSettings": {
      "turnDetection": {
        "threshold": 0.7,
        "silence_duration_ms": 800
      }
    }
  }'

Channel overrides apply to every session on that channel. If omitted, the agent's base configuration is used. See the WebRTC, WebSockets, VoIP, and SIP guides for channel-specific setup.

On this page