Custom AI avatar development cost depends on the application behind the character: what it must answer, which systems it can use and where it will run. Budget for development + provider usage + hosting and support + any installation hardware. A presenter-video subscription covers a different purchase from an integrated assistant.
As a simple planning example, 400 hours at an assumed $100 per hour equals $40,000 in development. Those are illustrative inputs, not a Virtual Verse quote or an observed market average. The worksheet below shows how to replace them with a scope you can compare between suppliers.
If you are preparing a brief for custom AI avatar development, start by choosing the job the avatar needs to do. “A digital human for our website” leaves most of the cost undecided.
Define which kind of avatar you are buying
A recorded presenter, a live conversational avatar and a spatial character are different deliverables. Decide which behaviour matters to the user before comparing prices.
| Deliverable | What the user can do | What the estimate must include |
|---|---|---|
| Recorded presenter | Watch prepared material | Script, character or presenter, voice, editing and output rights |
| Live conversational avatar | Ask questions and receive responses during a session | Conversation design, speech, approved knowledge, session handling and a user interface |
| Avatar connected to business systems | Retrieve account details, check availability or request an action | Authentication, backend integration, action validation and failure handling |
| Spatial installation | Interact through position, movement or a physical setting | Sensing, character behaviour, target hardware, calibration and on-site testing |
These categories can overlap. A live video avatar can answer questions without being a rigged 3D character in a Unity scene. Microsoft's avatar overview documents both batch and real-time output. Confirm the actual capabilities of the selected product instead of budgeting from its marketing label.
A useful first brief reads: “A web assistant that explains five approved products in English and hands qualified enquiries to our team.” That is much easier to estimate than an assistant expected to answer anything across every channel.
The workstreams that determine the development quote
Ask for separate estimates for the character, the conversation and the surrounding application. Otherwise, a low headline price can hide work the buyer assumes is included.
Discovery and conversation design. Define the audience, successful task, permitted answers and human handoff. The supplier should identify information the avatar needs but your team has not yet prepared.
Character and animation. A stock character, a custom 3D character and a trained likeness require different production work. Include idle states, listening, speaking and transitions in the scope. Approving a still portrait does not approve how the character looks while talking.
Knowledge and response testing. Source selection, document preparation, retrieval and answer evaluation take work even when a provider supplies the underlying model. Decide who approves answers and how changed product information reaches production.
Application and integrations. A public information assistant is simpler to scope than an authenticated assistant that reads customer records or changes bookings. Each integration needs agreed permissions, validation, timeouts and a useful failure response.
Deployment and acceptance. Include target browsers or devices, microphone permissions, session cleanup, monitoring and a handover. A physical installation also needs an equipment list and an operator procedure.
For a bespoke likeness, resolve provider access and production dependencies before promising a launch date. Microsoft's custom avatar documentation describes access requirements, training material and consent. Custom avatar and custom voice are separate capabilities; do not assume that buying one includes the other.
An illustrative AI avatar budget worksheet
Use hours and deliverables to interrogate a quote, then replace the assumptions with the supplier's estimate. The following allocation is a fictional planning scenario, not a claim about how long your project will take.
Assume one web deployment, one language, an existing character, a small approved knowledge set and one simple enquiry handoff. Exclude account access, custom likeness production and physical sensors.
| Workstream | Illustrative hours | Cost at an assumed $100/hour |
|---|---|---|
| Discovery and conversation flows | 40 | $4,000 |
| Character setup and interaction states | 60 | $6,000 |
| Conversation and approved knowledge integration | 100 | $10,000 |
| Web application and enquiry handoff | 100 | $10,000 |
| Testing, deployment and handover | 100 | $10,000 |
| Development total | 400 | $40,000 |
At an assumed $75 per hour, the same 400 hours would be $30,000; at $150 per hour, $60,000. This arithmetic shows why an hourly rate alone is not a useful comparison. The hours, exclusions and acceptance criteria must also match.
These figures exclude provider charges, licensing, hardware, taxes, maintenance and contingency. They also assume your team supplies usable source content and timely approvals. Request change-control terms for additional languages, new integrations or a revised character after approval.
For a fixed-fee proposal, ask what concrete result each milestone purchases. “Integration complete” should identify the actual workflow and its tests, not just a successful API connection.
Calculate running costs from sessions and provider units
Recurring AI avatar cost follows actual usage and the chosen architecture. Do not multiply visitor minutes by an arbitrary universal rate: different providers and components use different billing units.
Start with demand. In a hypothetical month with 2,000 sessions averaging three minutes, the application handles 6,000 session-minutes. That figure is not automatically the number of billed speech minutes or model tokens. Users spend some time listening, some speaking and some waiting.
Have the supplier map demand to each billable component:
- Avatar streaming or rendering usage, including any minimum charge or idle-session rules.
- Speech recognition and generated speech, where billed separately.
- Model input and output, including conversation history and tool calls.
- Application hosting, storage, monitoring and support.
OpenAI's voice latency and cost guide explains cost factors in its voice stack. Microsoft's avatar documentation also distinguishes avatar output from speech-related charges. Use the selected providers' current pricing and record the date of the estimate.
Ask for low, expected and peak usage scenarios. Concurrent sessions determine whether the experience can serve a rush of visitors; they are a different number from monthly sessions. Agree how to cap spending and what users see when capacity is exhausted.
Keep the pilot narrow enough to answer a buying decision
A useful pilot proves that the avatar helps one audience complete one task within acceptable quality and running costs. It should also test the failure cases that could make the full rollout unsuitable.
Use these acceptance questions in the procurement brief:
- Can a first-time user start, interrupt and end a session without assistance?
- Does the avatar answer the agreed test set from approved material and acknowledge missing information?
- Does the application confirm backend actions before announcing success?
- Are response times acceptable on the intended devices and network, including slower sessions?
- Can the team update content, review failures and recover from an outage?
- Is there a usable alternative when the microphone or avatar service is unavailable?
For technical planning, our interactive AI avatar architecture checklist covers the response path and deployment tests in more detail.
An attractive character is useful evidence about presentation. It is not evidence that account access, answer reliability or operating cost has been solved. Tie the next investment to the task results and the remaining engineering work.
Compare suppliers using relevant evidence
Ask suppliers to identify which part of a reference project resembles your brief. A physical character installation can demonstrate sensing and interaction work without proving a transactional customer-support assistant.
Our Meet Eva Here project involved a Unity installation with computer vision and gesture/body tracking for artist Shavonne Wong at ArtScience Museum. That is relevant to spatial interaction. We do not use it here as a published conversational latency benchmark, a price benchmark or evidence of lead conversion.
For your proposal, request a demonstration of the intended interaction, a list of included dependencies and a clear handover. Clarify ownership of character assets, application code, source content and provider accounts. Also establish who maintains the experience if your original supplier is no longer involved.
Send a brief that produces a usable estimate
A useful enquiry contains the audience, one task, deployment environment, languages and desired character. Attach examples of the approved content and describe every system the avatar must read or update. Add expected session volume, peak concurrency and the deadline or event date if those are known.
If budget is still undecided, ask for two proposals: a constrained pilot and a production phase conditional on its results. Require each to show development, recurring costs, exclusions and acceptance criteria separately.
Send Virtual Verse your AI avatar brief. Include what the user should accomplish and where the experience will run, so we can scope the character, application and integration work together.

