AvtoUM Case Study: Architecture and Development of a Digital Employee
Some projects begin with a spec. Others begin with an observation. AvtoUM was born from the second kind: a business already pays for its website, for ads, for traffic — and yet a gap yawns between the visitor's first question and the actual lead..

Article contents17×
Some projects begin with a spec. Others begin with an observation. AvtoUM was born from the second kind: a business already pays for its website, for ads, for traffic — and yet a gap yawns between the visitor's first question and the actual lead. The manager answers late. The client arrives in the evening. Product information is scattered across different places. Conversation history gets lost between the website, Telegram, VK, and the internal CRM. A regular chat answers a single question and stops there. Let's break down how this project became a digital employee and what stands behind its architecture.
Starting point: the silent storefront
A website is the most expensive storefront a business owns. Years of work, money, advertising, search rankings are invested in it. The main flow of clients converges there. And this storefront has a flaw: it's mute. It can show a product, help find what's needed, and accept payment. It doesn't see the living person browsing pages right now, about to leave.
A visitor comes in, looks around, hesitates, and closes the tab. Silently. Maybe they had questions, but nobody answered. They didn't call, didn't write — just left for whoever was clearer and faster. And this happens constantly. During the day, when managers can't keep up with everyone. In the evening, when there's nobody left to answer.
By Russian market estimates, up to sixty percent of inquiries come outside working hours. The peak is between five and nine in the evening. During the day, people are busy at work — no time to choose and place an order. They get to their own affairs in the evening, when they return home. And at that very hour, managers have already gone home. A paradox emerges: the client is ready to buy precisely when there's nobody left to answer.
This quiet client leak is the most expensive loss a business has. The one that appears in no report. Take ten million rubles in revenue and separate out the share that comes outside working hours. By industry estimates, that's up to forty percent. More than three million that pass through evenings and weekends every year. And every evening, while the website stays silent, part of those millions goes to the competitor who answered first.
What was actually designed
AvtoUM was built as a digital employee, not a chatbot. The difference is fundamental. A chatbot answers a single question. A digital employee knows the company's context, understands the page the visitor came from, sustains the conversation, helps qualify the inquiry, saves it to the CRM, and hands the dialogue to a human when needed.
For the business, the point isn't adding one more window to the website. The point is turning an inquiry into a managed process: from the first message to a request, a call, a booking, or a sale. The formula is simple: not just an answer, but a managed customer journey.
The engineering proof of this idea is the chain Visitor → Dialogue → Contact → Request → Message. Every step of that chain must be implemented, tested, and recorded in the CRM. If even one link is missing, the system becomes a pretty toy.
First user scenario: the website widget
It all starts with the most visible layer — communication on the website. A company receives a widget token and connects it to its resource. The visitor doesn't need to install an app or go to a separate service.
When a dialogue opens, the system receives more than the message text. It sees the permitted page context: URL, title, section, and with the relevant integration — product, price, and availability. So the answer can account for what the person is looking at right now.
The connection runs over WebSocket. It's a two-way realtime channel: the visitor sees answers without page reloads, and a manager can join an already-open conversation. For a returning visitor, a pseudonymous identifier and a brief history are preserved — the conversation continues meaningfully instead of starting from scratch each time.
Managed behavior: layers of context
The digital employee's behavior is formed from several managed layers. The owner sets the role, name, communication style, company details, product knowledge, and custom instructions. But user configuration is only one layer. At the platform level, there are global rules and constraints. They exist so that one unfortunate phrase in a company's settings doesn't break the overall behavioral logic.
In the current implementation, the system context is assembled sequentially: platform rules, company persona, niche context, knowledge base, current page details, individual instructions, and closing constraints.
This is an important architectural decision. The prompt here is a managed model of priorities, not one big text block. Thanks to this, the system can be developed, tested, and adapted to different companies without copying the entire logic.
Unified channel and CRM contour
AvtoUM goes beyond the window on the website. The architecture provides a unified message contour for the website, VK, Telegram, and MAX. Each channel has its own adapter, but the core business logic isn't copied four times.
After normalization, the event enters a unified messaging pipeline. The system checks the company and plan, restores the contact, creates or finds a request, saves the message, checks the schedule and limits, and then chooses the response method.
If an exact managed quick answer exists, the system can return it without calling a large language model. That's faster, more predictable, and saves the AI budget. If a substantive answer is needed, context is assembled and the selected model is called.
As a result, the business owner sees not scattered chats but a connected CRM chain: company, contact, request, and message history.
Human control and realtime
Automation excludes humans where responsibility is required, and here the principle works: a live manager can take over an active dialogue. At the moment of takeover, the AI stops sending answers. The employee's message returns to the visitor's active WebSocket session via Redis Pub/Sub. After the conversation ends, control can be returned to the automated contour.
Redis is used here as the primary business database. It stores short-term history, dialogue state, presence, and realtime events. Long-lived entities — contacts, requests, and messages — are stored in PostgreSQL. This separation helps avoid mixing temporary state with transactional data and allows realtime and CRM to evolve independently.
The key idea worth keeping in mind: human control is a system state, not an emergency workaround.
AI contour and model selection
Now for the part that distinguishes a production system from a demo prototype.
The primary model in AvtoUM is YandexGPT. This is a deliberate choice, not a technical accident. For Russian business, it's critical that personal data processing happens within the legal framework of the Russian Federation. YandexGPT runs on servers located in Russia and complies with Federal Law No. 152 "On Personal Data." For corporate clients, this isn't a detail — it's a condition without which the system cannot be admitted to handle real inquiries.
Additional models are provided in the architecture as alternative contours. Different AI tasks have different complexity and different allowable costs: answering a client, extracting contacts, determining topic, analyzing intent, and business recommendations. The primary load falls on YandexGPT, while certain auxiliary operations can run on other models.
This delivers what I call tokenomics — a managed token economy. Every call is recorded by task type, number of input and output tokens, and cost. Cost is calculated in rubles, accounting for the company's plan and platform limits. Some operations run through fast managed answers, some through lightweight models, and only substantive requests go to the primary one. At both company and platform level, limits and soft degradation of expensive features are provided.
The key principle: client personal data is processed in a protected contour on servers in Russia, and the routing architecture lets us manage quality, cost, and regulatory compliance. These are three things that, in an AI product for business, are designed together.
Working with knowledge: precise wording
Company knowledge enters the digital employee's managed context along with conversation history and current page data. For most typical companies, this is enough to launch a controlled first contour without heavy retrieval infrastructure.
Here it's important to use precise terms. In the currently confirmed code, knowledge is connected as a structured context block. A full vector RAG with embeddings, a separate encoder, hybrid search, and a reranker is the next architectural level for large and constantly updated bases.
The system's boundaries and the connection point for such a retrieval pipeline are provided: it can be added when document volume and retrieval requirements truly justify it. This is an honest position: the current state is described as it is, and the next level is marked as a possibility, not as an implemented feature.
Analytics and lead intelligence
After the dialogue ends, the system's work finishes. AvtoUM extracts permitted contact details, determines the inquiry topic, evaluates emotional signals and lead temperature. These attributes are stored next to the request and help the manager grasp the context faster.
At the analytics level, daily metrics, dialogue topics, funnel, revenue, and AI expenses are collected. On this basis, a Business Snapshot is formed.
The AI advisor receives actual company metrics and deterministic signals. Its instructions explicitly forbid inventing numbers. If data is insufficient, the system reports this and suggests which measurements need to be enabled.
For me this is fundamental: a generative model complements accounting. It helps interpret data that the system has already collected. First facts, then AI interpretation.
Billing and product economics
For an AI product to exist as a business, it needs service economics.
AvtoUM implements plans, monthly dialogue limits, a trial period, add-on packages, renewal, plan changes, and proration. The payment contour is connected to YooKassa.
An important security rule — an incoming webhook cannot be trusted automatically. Before changing access, the system re-verifies the payment status with the provider. Operations use idempotency so a repeated event doesn't create a duplicate charge.
AI call costs are accounted for separately. Thanks to this, you can see not only revenue but the real cost of the intelligence contour.
Architecture passport
For maintaining such a project, one diagram isn't enough. So AvtoUM's architecture is broken into separate views.
The system map shows components and connections. Each node has a passport: purpose, responsibility, inputs, outputs, technologies, controls, confirmation level, and a link to the source in the project.
An ER data model is formed separately. It shows how companies, users, contacts, requests, messages, payments, employees, bookings, and analytics entities are connected.
The API registry records real routes and access levels. Sequence diagrams reveal the path of a message through the widget and messengers, as well as the payment scenario. The role matrix answers who has the right to perform an action and who bears responsibility.
There's also a failure register: what happens when the model, Redis, external CRM, or payment provider is unavailable; how the system detects the problem, what safe fallback it applies, and how it recovers.
For the client, this is a way to agree on product boundaries before expensive implementation. For the development team, it's a source of project constraints and a base for the Agent Pack.
Technology stack and production
The stack was chosen by roles within the system.
The interface is built on React 18, Vite, and Tailwind CSS. React Flow is used for interactive maps, Recharts for analytical views.
The backend runs on Python 3.12 and FastAPI. Asynchronous data work is built on SQLAlchemy and asyncpg, with migrations handled by Alembic. FastAPI dependencies are used for authorization, tenant context, and access control.
PostgreSQL stores the product's transactional data. Redis handles short-term state, history, presence, and Pub/Sub. APScheduler runs background processes: aggregations, plan events, classification, and notifications.
Production is assembled in Docker Compose. The backend and frontend containers are published only on loopback ports, and external traffic passes through the system Nginx with HTTPS.
At the current stage, Kubernetes and Kafka or RabbitMQ are absent. For the existing load, Docker Compose and Redis provide a simpler operational contour. If load and the number of independent handlers grow, container orchestration and a separate event bus become a justified next step, not a technology added for the resume.
Development method
In this project I was responsible for one isolated module. I formed the product concept and designed the frontend, backend, CRM, realtime contour, prompt model, model routing, billing, AI economics, administration, and production infrastructure.
In my work I use an AI-native approach and agentic tools, including Claude Code and Codex. This accelerates research, implementation, testing, and maintenance of large codebases. Architectural decisions, security boundaries, product priorities, and final verification remain my responsibility.
For repeatable operations, the project records rules, sources of truth, deployment order, and verification criteria. This approach matters more than the speed of code generation: the agent must work within the architecture instead of reinventing the project each time.
The honest boundary of the current implementation
I deliberately separate what's already confirmed by code and production configuration from what belongs to the next level of development.
Currently confirmed: async FastAPI, WebSocket, PostgreSQL, Redis, multi-layer prompts, LLM routing, CRM, human takeover, analytics, billing, AI cost accounting, and containerized production.
Full vector RAG with encoder and reranker, tool-driven ReAct, a formal Evals and LLM-as-a-Judge contour, Kafka or RabbitMQ, Kubernetes, and automated CI/CD are added under specific load and quality criteria. They are separate engineering subsystems with their own cost.
Author's column
Climax: what this case shows
AvtoUM is an example of how one architect can connect product, data, AI, backend, frontend, and operations into one managed contour. That's rare in a market where most AI projects are assembled from pieces written by different teams, then suffer from inconsistency for years.
The project shows something else too. An AI product for business exists as a business: with plans, limits, billing, cost accounting, and clear economics. A pretty interface and good model answers are part of the story. The rest is discipline, data, and engineering.
And the third thing worth keeping in mind. A digital employee for business isn't about replacing people. It's about expanding what a company can do when its website already brings clients and there's nobody to lead them to a deal. An AI employee works around the clock, answers in seconds, saves everything to the CRM, and shows the owner the numbers. Meanwhile the team focuses on live, hot clients.
Your website attracts visitors. AvtoUM turns them into clients. That phrase contains the whole essence of the project.
Glossary of terms
- Digital employee — an AI system that conducts dialogue with clients on the website and in channels, saves data to the CRM, and hands the dialogue to a human when needed.
- Widget — an embeddable website component through which a visitor opens a dialogue with the digital employee.
- WebSocket — a two-way realtime communication protocol between browser and server.
- Realtime — an operating mode where data is transmitted instantly, without delay or reloads.
- Presence — an indicator of a dialogue participant's current availability in the system.
- Rate limit — a cap on the number of requests or actions within a given period.
- Persona — the configurable role and communication style of the digital employee.
- Guardrails — constraints that define the boundaries of an AI system's behavior.
- Messaging pipeline — an automated chain of message processing from arrival to response.
- Channel adapter — a module that brings messages from a specific channel into a unified internal format.
- Event normalization — bringing messages from different sources into a single structure.
- Human takeover — a live manager taking over a dialogue with the AI stopping its responses.
- Redis Pub/Sub — an event exchange mechanism between components via Redis.
- LLM Router — a component selecting a model for a specific task by type, cost, and plan.
- Fallback — a backup scenario when the primary component is unavailable.
- Prompt caching — caching parts of a prompt to speed up and reduce call cost.
- Usage accounting — tracking AI resource consumption: tokens, calls, cost.
- Vector RAG — an approach where the model answers based on documents retrieved via vector search.
- Embeddings — numerical representations of text for comparing semantic similarity.
- Hybrid search — a combination of semantic and keyword search.
- Reranker — a model that reorders search results by relevance.
- Business Snapshot — a summary of key company metrics for decision-making.
- YooKassa — a Russian payment provider.
- Webhook — a notification from an external system that an event has occurred.
- Idempotency — the property of an operation producing the same result when repeated.
- ER model — a schema of entities and their relationships.
- API registry — a list of available routes and access levels.
- Sequence diagram — a diagram of the interaction sequence between components.
- RACI — a matrix of role and responsibility distribution.
- Failure scenarios — descriptions of system behavior under failures.
- Agent Pack — a set of project materials for AI agents to work within the architecture.
- FastAPI — a Python web framework for asynchronous APIs.
- Asyncpg — an asynchronous PostgreSQL driver for Python.
- SQLAlchemy — a Python library for working with databases.
- Alembic — a migration tool for SQLAlchemy.
- APScheduler — a background task scheduler for Python.
- Docker Compose — a tool for running multiple containers as one system.
- Nginx — a web server and reverse proxy.
- Loopback port — a port accessible only from within the machine.
- Kubernetes — a container orchestration system.
- Kafka, RabbitMQ — message brokers for event-driven architecture.
- CI/CD — the practice of continuous integration and delivery.
- Evals — a formal quality evaluation contour for an AI system.
- LLM-as-a-Judge — an approach where answer quality is assessed by another language model.
- AI-native development — an approach where AI tools are embedded in the development process from the start.
- YandexGPT — a large language model by Yandex, running on servers in Russia.
- FZ-152 (Federal Law No. 152) — Russian law "On Personal Data," regulating personal data processing.
- Tokenomics — the managed economics of token consumption: routing, limits, and cost control.
Respectfully,
Yuri Eliseev
AI Systems Architect · Full-Stack Product Engineer
Need a project of any complexity?
Let’s discuss an idea, product, AI system or technical challenge and define a realistic first step.
