AI4IA keeps chat, agents, documents, memory, tools, and voice in one workspace. The most useful distinction is between what you ask for, what context the model receives, and what actions it is allowed to take. The Conversation Inspector makes those choices visible; the API enforces them.
- Open your environment's web app and sign in with Microsoft Entra ID. Local development may instead use a configured development identity.
- Start a conversation with New chat, or reopen one from the sidebar. The sidebar groups conversations by how recently they changed and offers a search field once the list is long. Its panel button collapses it to a slim rail.
- Open the Conversation Inspector with the settings button in the conversation header, or select the model name there to go straight to its controls. Setup controls the model, instructions, agent, tools, and voice; Context controls documents and memory; Usage explains the recorded consumption.
- Describe the outcome you need and attach or select only the relevant sources.
Document library, Photo avatars, Agents & workflows and Settings open as pages in the same workspace. The browser's back button returns to the conversation, and a live voice session keeps running in a small player while a page is open. Settings holds appearance and accessibility options, deletion status, and links to help. On a phone, the menu button opens the sidebar and the inspector opens over the conversation.
Features vary by deployment. A hidden control can mean the operator disabled the capability or the selected model cannot use it; it is not a permission you can grant by changing a browser setting.
Give the assistant a goal, constraints, and an expected output. Use attachments for one-off material and the library for sources you expect to reuse. Check citations against the original source before relying on an answer.
Model controls reflect the server's catalog. Context size, output limits, reasoning effort, sampling, input modalities, and tool support vary by model. Plain chat only models can answer ordinary questions but cannot run agents or workflows that need tools. A larger context window is a capacity limit, not a promise that every document will be included or every fact recalled.
A selected agent is the standing persona. A leading @agent mention overrides
it for one turn, for example @coder explain this function; a mention later in
the message is ordinary text. Type @ at the start for available agents, or
use /agents. The internal @conversation badge means conversation-attached
tools without a selected agent, not another agent to invoke.
The inspector shows inherited instructions and tools alongside conversation overrides. Saved server values, not an unsaved control or a model's claim about its abilities, determine the next turn.
| Use | Best fit | Important boundary |
|---|---|---|
| Plain conversation | A question or exploratory task | Context and model limits still apply |
| Agent | A reusable persona, model, and tool bundle | Attach only the tools it needs |
| Workflow | An ordered, repeatable set of steps | Each step has its own effective capabilities |
| Durable workflow | Work that should survive an API restart | Requires deployment support and an explicit per-run choice |
In the workflow editor, Build defines the steps; Run & test runs them and keeps their results visible. The result distinguishes completed, failed, and unstarted steps. Open in chat opens the run's conversation.
Tools for this step adds capabilities to the selected agent's tools. Read the effective capability list: chat-only capabilities are not promised in a workflow. In particular, a step cannot upload or analyze a new library document; document-review templates work on sources already uploaded and ready.
Memory tools are explicit per step. Enable Save memory when a step must store a fact, and look for the tool's result rather than trusting text such as "I've remembered that." A deduplicated fact can correctly report that nothing new was stored. Under Documents, a non-empty workflow selection restricts the run; selecting none allows the run to read your ready library documents.
Keep running if the app restarts uses Azure Durable Task Scheduler. The page may stop waiting after two minutes without cancelling the run; its result can still arrive in the run's chat. Without this option, a replica restart can interrupt the request. Durability does not grant additional tools, remove approval requirements, or undo a tool's external effects.
The chat Run workflow tool is deliberately narrower than the workflow editor. It advertises only enabled workflows whose resolved steps are safe, read-only, non-recursive, and compatible with that execution path.
When the operator enables publishing, saved builder definitions can be submitted for independent review. Save edits first: the submission uses the saved revision, not an unsaved form. Shared names explicit recipients or group IDs; Tenant-visible is visible only within the application's authenticated tenant. Neither setting makes the asset public on the Internet.
Submission explicitly permits an authorized reviewer to inspect that frozen version. Optional operator review is a separate choice. The author cannot self-approve. After approval, the author activates the exact reviewed version. Changed source, audience, models, tools or resource metadata needs another review. Personal MCP connections and unreviewed private dependencies are not copied.
Published catalog entries have stable handles and exact source-version references. Your own permissions, model constraints and tool approvals still apply. A withdrawn, superseded or changed selection is refused rather than silently replaced. Required tools cannot be removed; supported optional narrowing is recorded. A reviewed Exclude skills profile does not load skills and cannot remove a required skill. Otherwise, publication requires explicitly versioned official skill resources rather than mutable defaults.
Receipts distinguish the approved profile from the actual offered subset. Published voice requires an owned conversation and retains a bounded server session-end receipt; provider usage/parameters that the relay does not observe remain unknown. Expiry blocks new work without discarding accepted work's accounting or cleanup. Group-dependent unattended execution cannot use a stored token-claim snapshot as current authorization.
Activity shows observable work: searching, reading, invoking tools, being blocked, or failing. A completed turn retains that bounded activity history.
An Execution receipt gives more detail: the effective redacted prompt, model/region, admitted and displaced context, source versions, offered and invoked tools, bounded arguments/results, approvals, usage, and safety coverage. Loaded skills include their source URI, version resolution, hash, and truncation. Long workflows also retain independently bounded step receipts.
Under Runtime, Application-effective model parameters shows the controls actually sent for each model call after the application adapted them for the chosen model. An omitted control says its provider default is unknown. Workflow steps and delegated agents show their own defaults rather than inheriting the parent turn's sliders.
Estimated model cost uses the token rates and price version recorded during execution. Known subtotal means some calls could be priced but the total is unknown; Unknown is not zero. These are model-token estimates, not a bill, and exclude tool, media, and other service charges. Changing session settings or updating prices does not change a saved receipt. Older receipts show parameters and cost as not recorded, rather than reconstructing them from current settings.
These are not chain-of-thought. They show what the application supplied and executed, not the model's private reasoning or proof of which source caused an answer. Shortened payloads retain their original redacted size and digest.
The single Attach control accepts the types and limits advertised by the server. Uploads run sequentially with visible progress and retry/dismiss actions. A library upload becomes selected conversation context only after association succeeds. Navigation is temporarily blocked while an upload is active so it cannot land in a different conversation.
| Path | Use it for | Lifetime and scope |
|---|---|---|
| Session attachment | One-off material for this chat | Bounded, session-scoped context; not a reusable library entry |
| Document library | Reusable documents, images, audio, or video | Owner-scoped source bytes, analysis, manifest, and retrieval index |
Only ready library documents participate in retrieval, sharing, media deep-links, memory saves, and document tools. Upload acceptance alone does not mean analysis and indexing have completed.
- Automatic - Content Understanding chooses the modality-appropriate Azure analyzer and is the normal default.
- Mistral Document AI / Mistral OCR 4 are explicit PDF/image alternatives, limited to 30 pages and 30 MB per request.
- When enabled, Content Understanding Read / Layout are preview, synchronous options for small files: 10 MB and the first five PDF pages.
The analyzer is part of deduplication: the same bytes analyzed two ways produce separately attributable results. Ready Content Understanding documents can expose Evidence with structured fields, confidence, grounding, and provider details. Confidence needs workload-specific interpretation; it is not a correctness guarantee. See document and multimodal understanding.
You do not upload separately to Search. Library ingestion extracts content, chunks it, obtains embeddings, and indexes the chunks. Retrieval combines keyword and vector search, with semantic reranking when configured.
Deployments can use per-user or shared indexes; access filtering still applies to every query. An explicit conversation selection is an allowlist. Clearing it to an empty selection disables library context; older sessions without a selection retain the all-accessible behavior. Revoked sharing is rechecked, so a stale selected id cannot restore access.
Search is derived state. Cosmos owns the manifest and Blob owns the source bytes and parsed artifacts. Without a Search endpoint, the library uses an in-memory chunk store even outside local development. That index is replica-local and lost on restart; stored summaries and parsed-document reads remain available. It is not equivalent to shared, persistent retrieval on a scaled deployment.
Owner-scoped maintenance endpoints under /api/library separate retrieval from
the original document:
| Action | Endpoint | Additional model work |
|---|---|---|
| Inspect one | GET /documents/{id}/index |
None |
| Rebuild one | POST /documents/{id}/reindex |
Embeddings |
| Rebuild all your ready documents | POST /documents/reindex |
Embeddings |
| Remove one from retrieval | DELETE /documents/{id}/chunks |
None |
Reindexing reuses saved extraction, not another analyzer run. The saved
chunks.jsonl sidecar preserves boundaries and media grounding; older documents
without it fall back to parsed Markdown and may lose time grounding.
Removing chunks leaves the document and analysis intact. These are maintenance
operations, not agent tools; rebuilding is metered and entitlement-gated.
Private means owner-only, shared grants read access by email, and public means readable by authenticated users of the configured tenant. Public does not create an anonymous internet link.
Sharing revocation affects subsequent reads. It does not erase snippets already saved in another conversation or its historical receipt.
The orange microphone starts and stops Voice Live in the current conversation. Finalized spoken turns are saved to the normal transcript. Play on an assistant message is separate text-to-speech and does not require a live socket.
Open Setup > Voice for provider, model, voice, locale, and supported audio options. Azure OpenAI uses catalogued realtime deployments. Optional Azure Speech uses a curated managed-model catalog in East US 2; it does not accept arbitrary model names, custom endpoints, or personal voices.
Azure Speech also offers MAI voices, marked preview in the Voice list, and MAI Transcribe 2 (preview) under Transcription in place of the model's default. Preview options have no service-level agreement. If Azure refuses one, Voice Live shows Azure's error instead of switching to another model. A turn it can't transcribe is flagged in the call bar while the session continues.
The Azure Speech Voice list groups its US English voices by family: Dragon HD, Multilingual, Neural and MAI (preview). Speaking rate (0.5 to 1.5) makes any Speech voice slower or faster; Microsoft doesn't document it for the preview MAI voices. Voice variation (HD voices) (0 to 1) sets how much a Dragon HD voice varies its intonation. It is disabled for other voices and keeps your value for the next HD voice. It is separate from Temperature, which shapes the reply itself. Leave either empty for the voice's own default.
The GPT-5.x Azure Speech models ignore Temperature, so it is disabled for them (your saved value still applies to the other models). GPT-5.2, GPT-5.4 and the GPT-5.6 models also answer at the lowest reasoning effort, none, so spoken replies start quickly. If Azure refuses a reply, the call bar shows Azure's reason while the session stays connected; try another speech model or voice in Setup > Voice.
With Azure Speech, Turn detection sets how Azure tells that you've finished speaking: Semantic (English) or Semantic (multilingual). Stop the reply when I start talking lets you cut in by speaking, and Let Azure trim interrupted replies keeps only the part of an interrupted reply you heard in the conversation.
Settings apply to the next connection without silently reconnecting the current one. The API supplies the selected agent persona or saved conversation instructions; voice has no competing instructions field. Every live session also gets short spoken-delivery guidance from the API, after any persona: answer the way people talk, usually in a few sentences, offer more instead of listing, and never read out lists, markdown, links or emoji. The persona's own instructions still apply.
If AI4IA is updated while your tab is open, a banner says A new version of AI4IA is available. Choose Reload when convenient; it waits while a voice session, transcript save, reply or upload is in progress, and asks before clearing a message you haven't sent. Until you reload, Voice Live asks you to reload before starting a new session, because voice behavior lives in the page.
You can also type while connected. The composer then shows Send to: with the live session selected (the default), a typed line goes to it, joins the voice transcript and is answered out loud. A line typed while the assistant is replying, or while you are speaking, waits and is sent as soon as it can be. Choose Text chat to send to the conversation's text model instead; that reply is not spoken, and it enters the live session's context on its next connection. A failed microphone permission or connection attempt does not create an empty conversation. If saving finalized turns fails, use Retry or Discard; stopping still releases the microphone and socket. A lost/muted microphone or unrecoverable audio context closes the connection rather than leaving a misleading "live" indicator.
Voice connects directly to the API's WebSocket ingress, through its separately scoped APIM route. It does not bypass authentication, Origin validation, entitlements, metering, or tool governance.
Image, video, processing, and export tools create durable artifacts when enabled. Downloads use authenticated API routes, not anonymous Blob links.
For images, open Setup > Agent & tools > Image generation in a saved
conversation. Select one to three models and a size/quality they share, then
save. Start image in chat or Start comparison in chat adds
/generate_image without discarding your draft. Each request snapshots the setup;
comparison output stays in selection order and records its model and deployment
provenance.
When image editing is enabled, choose Edit under an image in the
conversation, or the ✏️ button on one of your own PNG or JPEG library images
selected for the conversation. Describe the change and keep the recommended
model (GPT Image 2.5 Sunburst) or pick another editing model. Whole image
edits everything; Selected region marks a rectangle to change: drag across
the preview or set its edges with the sliders. The result arrives as a new image
in the conversation; the source is never changed. You can also type
/edit_image <what to change> to edit the latest image, or attach the
edit_image tool to an agent. Region edits are unavailable for rotated phone
photos; edit the whole image instead.
Video generation is asynchronous and slower than a text reply. Supported clip lengths are 4, 8, or 12 seconds, with 4 seconds as the default.
Video generation is offered only while a video model is enabled. Sora 2, the
only available video model, retires on October 15, 2026 with no replacement, so
this deployment no longer offers /generate_video or the agent tool. Agents
that list it keep their other tools. Clips you already generated remain in their
conversations and stay viewable.
Cost estimate unavailable is not free. Published estimates can differ from Azure billing, especially for provider-specific media meters.
Photo avatars are off by default. An operator can turn them on only after Microsoft approves the deployment's Limited Access registration for custom avatars; until then the sidebar shows no Photo avatars entry. See custom photo avatars for the design and its boundaries.
When they are on, Photo avatars in the sidebar opens your gallery. Describe a fictional adult and give the avatar a name. Style, age, gender, and ethnicity are optional and start unspecified. Before creating, you confirm that the character is fictional, an adult, and not modeled on a real or identifiable person. Each creation is billed and counts toward your avatar and 24-hour limits, which the form shows along with the estimated cost.
Generation typically takes under a minute, and the gallery shows its progress while it is open. Every preview is labelled AI-generated. An avatar marked Re-verifying… can't be used in a live session until the avatar service confirms it again; its preview and Delete still work. Report records a problem with an avatar and links to Microsoft's abuse report form. Delete asks first, then removes the avatar and its preview. Right after a creation it may ask you to wait a few seconds. An avatar created under an earlier avatar configuration can't be deleted from the gallery; ask an operator to remove it.
In Photo avatars, choose Use in Voice Live on a ready avatar. This selects Azure Speech and returns to the conversation, where the avatar's portrait stands on a stage beside the transcript (above it on narrow or tall screens) with Start talking. Selection alone does not open the microphone or start billing. Choose Start talking (or the chat microphone) to speak with it using the current conversation's instructions and agent. Its live video speaks the replies, labelled AI-generated for the whole session.
Speak, or type: while the session is connected, Send to is set to the avatar, so a typed line is answered out loud. Choose Text chat to send a line to the conversation's text model instead; that reply is not spoken.
The stage follows your screen. On a wide screen, the Conversation Inspector steps aside for the session; reopen it from the header at any time. Focus view gives the avatar the whole conversation area with captions, Fill frame crops the video to fill the stage, and Full screen enlarges the stage, label included. On a phone, the stage's buttons become a slim column beside the video. Before a session, the minimize button shrinks the stage to a slim bar above the composer; that choice is remembered, and choosing an avatar again brings the stage back. Opening another page keeps the session in a small player with Return to conversation and End session.
Choose avatar opens the gallery; Voice only removes the avatar without changing the Speech voice. You can also use Setup > Voice > Choose avatar, or pick a ready avatar from the Avatar list with Azure Speech selected. Unavailable avatars and failed availability checks are explained explicitly; they never silently start an audio-only session.
Avatar time is billed per second while the session is connected, even when nobody is talking, so a session ends on its own after a stretch of silence (a countdown warns you first) or at its time limit. End session stops it at once. If your browser can't play the avatar video, live avatar use is disabled with an explanation. Choose Voice only to continue without video.
On speakers, your microphone can pick up the avatar's own voice. So by default (While the avatar talks: Pause my microphone, in Setup > Voice under the avatar), your microphone sends silence while the avatar is speaking, and the stage and call bar read Speaking · mic paused. To cut in, choose Interrupt: it stops the reply and opens your microphone again. With headphones, choose Keep listening instead and interrupt simply by talking. Either way, the browser is also asked to cancel the avatar's voice from your microphone. The choice applies to your next session.
To interrupt by talking while using speakers, choose Keep listening with precise echo cancellation (preview). This page then also sends Azure what it plays, so Azure can remove the avatar's voice from your microphone. Your microphone opens once Azure confirms the session and stays open while the avatar talks. Interrupt still works.
This mode is a preview, offered only where your deployment's catalog has it. It removes only this page's own sound: other tabs or apps playing on your speakers still reach the microphone. If Azure doesn't accept the mode, the session ends with Azure's message. Switch back to Pause my microphone and start again.
If the connection stops responding, the client stops its microphone rather than queueing increasingly stale audio. A keepalive timeout is a connection failure, not a successful avatar response. Start a new session only after the connection has recovered; accepted audio is not replayed automatically.
Memory can carry personal context between conversations. In Context > Memory, Automatic memory is on by default. Turn it off to stop automatic recall, saving, and model memory tools across chats, agents, and workflows, including delayed/resumed work. This is a capability switch, not a consent ceremony or tool approval. It does not enable a backend disabled by the operator.
You can still create, list, edit, or delete your own records while it is off. User-created or edited memories are protected from automatic consolidation. The control keeps a change pending even if you switch conversations or close and reopen the inspector. After the request settles it reloads your current server setting. If confirmation fails, the last confirmed setting is restored visually but marked unconfirmed; reload it before trying again. Changing accounts never applies one owner's pending result to another. Turning memory back on makes retained records available to new work.
/forget removes this conversation's memories; /forget me removes all active
memories for your profile. Deleting a document also fences and removes memory
derived from it. A stale edit produces a conflict rather than overwriting a
newer version.
Expand Memories supplied below an answer for the bounded memory context recorded in its execution receipt, links to your inspector items, and an entry point to the full receipt. It describes supplied context, not which memory influenced a sentence or hidden reasoning. Memory tool returns are shown separately because a return alone does not prove later model delivery. Old answers without provenance stay unrecorded, and truncated or missing references are not reconstructed from your current records.
Disabling or deleting memory does not rewrite old answers or historical receipts. Already-sent prompts cannot be withdrawn, and existing transcript text can still be sent as conversation history. Provider backups retain their own retention window. See memory architecture for the deletion boundary.
Custom MCP servers and WebIQ require deployment support. MCP connection secrets live in Key Vault outside local development. Remote endpoints are checked before discovery and again when invoked. Neither a remote server nor a retrieved page can grant itself permission.
The MCP server builder's MCP protocol setting defaults to
2025-06-18 (legacy). 2025-11-25 (stateful) is a separate explicit choice,
not a switch to stateless. Choose it or 2026-07-28 (stateless) only after
verifying that specific server supports the selected version, then
Save & reconnect. The curated Foundry Toolbox selects November independently;
that does not change your servers' defaults. The API also accepts
an explicit protocolVersion on create/update; older update clients that omit it
preserve the saved selection. Protocol and routing-schema changes require renewed
tool consent. A rejected request never automatically switches protocol, drops
authentication or replays a tool call; change the saved selection explicitly if
you need to return to legacy. Official/Foundry protocol selection is catalog-owned,
not editable through the BYO server builder.
Ask for live information in chat, or use /research <query>. WebIQ is a tool
provider, not an @webiq agent. The available model tools are:
| Tool | Purpose |
|---|---|
web_search |
Web results and source content |
news_search |
News, publisher information, and timestamps |
video_search |
Videos, playlists, summaries, and timestamped moments |
image_search |
Existing images and source-page metadata |
browse_url |
Public HTTPS page content and returned links |
classic_search |
Structured answers such as weather, finance, places, and events |
finance_search |
Instrument prices and available as-of metadata |
places_search |
Places, businesses, hours, and available contact data |
sports_search |
Schedules, scores, and event data |
sonic_search |
Blended web/news/finance search |
web_autosuggest |
Query suggestions; beta entitlement, not an answer source |
These are tools, not eleven slash commands. Filters vary by endpoint. Classic search supports 30 categories but returns at most six answer types per call. An omitted type is not evidence that no information exists.
Safe search stays strict; output is bounded, redacted, and treated as untrusted.
Returned links are not followed automatically. A crawl may report pending
instead of content; there is no implied background polling. Credentials do not
prove entitlement to every vertical or beta endpoint.
Under the default policy, browsing and sandbox computation require approval because the model chooses a destination or program. Other outbound first-party tools, such as WebIQ search, prompt when the turn also carries untrusted document, memory, or tool context.
The card identifies the tool, destination, and bounded argument preview. Warnings identify hidden or omitted arguments. Approve only if the action matches your intent: a source can contain instructions designed to steer the assistant. Each approval is bound to one call's exact arguments, expires after ten minutes, and cannot authorize a different call or conversation.
When available, Setup > Agent & tools lets you opt in for the current saved conversation. Run & test has a separate per-run workflow choice that resets for the next invocation. Consent is not active until the server confirms it.
Consent covers only the recorded enabled-tool contracts, lasts at most eight hours, and can be revoked. New tools or changed contracts need renewed consent. Session consent does not authorize a workflow invocation. A workflow step that requires approval but has no run consent fails visibly rather than running with unattended authority.
The tradeoff is real: hostile source content can influence later calls while per-call prompts are skipped. Ownership, scopes, destination checks, and usage limits still apply, and activity/receipts retain approval provenance. Revocation stops subsequent dispatch; it cannot undo an external request already in flight. Revoke auto-approval & stop run preserves completed/partial workflow evidence.
Usage reports known token, image, page, and estimated-cost subtotals with coverage. Missing billing dimensions remain Unknown, not zero. Prompt pressure describes the latest token-metered turn and may be unavailable after switching models or when the provider omits prompt usage.
Application quotas are soft preflight checks, not hard spending caps. Concurrent requests can overshoot, missing provider prices/usage can undercount, and ledger-check failures can allow work. Azure budget notifications are also alerts rather than a mechanism that stops spending.
Admin access is enforced by the API. Admins can inspect usage by model, user, agent, date, deployment, and request outcome, plus Azure resource metrics and fixed-query operations panels. Capped scans are labelled truncated; unavailable, partial, and stale sources remain distinguishable. Telemetry does not provide complete proxy queue/fairness or provider-quota forecasting.
Admins are also unrestricted on their own usage. Anyone the deployment makes an
admin (AI4IA_ADMIN_SUBJECTS or the admin app role) skips per-user limits:
entitlement caps and an account disable, the photo avatar count and
daily-creation limits, the live avatar minute cap, the document library cap, and
the limits on how many agents, workflows and custom MCP servers they can save.
The admin dashboard shows Your usage limits: Unlimited (admin), and the photo
avatar gallery shows Unlimited (admin). Their usage is still metered and
appears in the admin views. Approvals and security checks, feature switches,
provider terms and safety systems, technical size bounds, Azure quotas and
throttling, group-policy restrictions and hard quota still apply, and a live
avatar session still ends after the idle timeout.
Cosmos is canonical for conversations, usage, agents/workflows, document manifests, and memory text and vectors. Blob holds source documents and generated artifacts. Search indexes and document chunks are rebuildable; deleting canonical data is a different operation.
To delete a conversation, choose the bin icon on its sidebar row and confirm on the row. The icon is always shown on the open conversation; on other rows, hover over it or tab to it. The conversation leaves your list and the app cleans up its messages and attachments right away, then tells you when it's done. If cleanup doesn't finish, the notice says so and offers Finish cleanup; Deletion status in the sidebar lists unfinished deletions. If the app can't confirm what happened, the row says so and offers Try again, which checks the same request and won't delete twice. Conversations created before resumable deletion was turned on are deleted the older, best-effort way.
Conversation deletion is not transactional erasure across all stores: an already-authorized concurrent write can leave an orphaned child record, which is most likely for those older conversations because they have no write fences. Library documents, memories, generated media and backups follow their own retention; do not treat a successful delete as a physical-erasure guarantee.
A model's region or data-zone selection concerns inference routing, not where your conversation and documents are stored. Global deployments are not region-resident just because their account has a regional name. Read the region and capability map before using sensitive material with a residency requirement.
Provider safety assessments are observations, not proof of safety. AI4IA's recorded policy is non-blocking assessment visibility; provider-native refusals still apply and modality coverage remains incomplete.
| Symptom | First thing to check |
|---|---|
| A control is missing | Deployment availability and the selected model's capabilities |
| A document is absent from context | Ready state, access, explicit selection, and the turn's context budget |
| A memory edit conflicts | Reload the latest record before retrying |
| A conversation won't delete | Read the message on its sidebar row. For an unconfirmed result, choose Try again; if cleanup didn't finish, choose Finish cleanup or open Deletion status |
| Voice fails before connecting | Microphone permission, sign-in, API URL, and allowed Origin |
| Voice settings seem unchanged | Stop and reconnect; settings affect the next connection |
| Voice won't start and a new version is available | Choose Reload in the banner, then start Voice Live again |
| Speech is not offered | The operator's provider allowlist and Speech feature gate |
| Photo avatars are missing | They are off by default and need Microsoft's Limited Access approval |
| Can't use an avatar | Choose Use in Voice Live in the gallery. The avatar must be ready and verified, Azure Speech Voice Live must be available, and the browser must support avatar video. End an active session or finish saving its transcript before choosing another avatar |
| An avatar session ended by itself | It ends after a stretch of silence or at its time limit; start Voice Live again |
| A typed line isn't spoken | While connected, check that Send to is set to the live session or avatar, not Text chat |
| The avatar looks small | Use Focus view or Full screen, or minimize other panels; the stage grows with the space it has |
| The avatar interrupts itself or answers its own words | On speakers, keep While the avatar talks on Pause my microphone, or use headphones |
| The avatar doesn't stop when I talk | While it speaks your microphone is paused: choose Interrupt, use headphones with Keep listening, or try Keep listening with precise echo cancellation (preview) on speakers |
| An avatar session with precise echo cancellation ends at once | The preview isn't accepted there: choose Pause my microphone and start again |
| Voice or the avatar hears you but never answers | The call bar shows Azure's reason for a refused reply. Try another speech model or voice in Setup > Voice |
| Search or another tool fails | Its visible error/approval state; an enabled gate does not prove upstream entitlement |
| An admin panel is unavailable | Resource wiring, API identity permissions, and source freshness |
Report the time, selected model/provider, and safe correlation/error code. Do not paste access tokens, prompts, audio, documents, or tool payloads into general operational logs.