Appearance
AI assistant and providers
Desorix's AI assistant does two jobs:
- Drafts for agents: a Draft reply button in the inbox writes a suggested reply. The agent edits it and sends it. Nothing is sent by itself.
- Auto-reply: answers customers automatically when no automation rule or chatbot flow did. When it isn't sure, it hands the conversation to your team.
Both answer only from the business description and documents you give it, and ask for a person otherwise. Meta's WhatsApp Business terms (from 15 January 2026) don't allow general-purpose AI chatbots, and a business-scoped assistant like this is what they allow.
AI is bring your own key: Desorix ships no API keys. Each call is billed by the provider to whichever key made it.
Keys
| Where | Who | Used for |
|---|---|---|
| AI → API keys (each workspace) | owners and admins | that workspace only |
| Administration → AI → Platform AI keys | the platform administrator | workspaces without a key of their own |
A workspace's own key always wins over the platform key. Keys are stored encrypted and never shown again; the list shows only sk-…1234. Test sends a tiny request and shows Works or the provider's error.
Two providers are supported:
- Anthropic (Claude): an API key from the Anthropic Console (
sk-ant-…). Desorix uses the official Anthropic PHP SDK. - OpenAI: an OpenAI key (
sk-…). You can also set an OpenAI-compatible base URL (https://…/v1) to use Azure OpenAI, OpenRouter, a local server, or any service with the same API.
Anthropic has no embeddings API, so meaning-based document search needs an OpenAI key. Without one, documents are searched by keywords (see below).
Models
The model is chosen per feature, because the two features have opposite needs:
| Feature | Default (Anthropic) | Default (OpenAI) | Why |
|---|---|---|---|
| Auto-reply | claude-haiku-4-5 | gpt-6-luna | Runs on every unanswered message, so speed and cost matter. |
| Drafts | claude-sonnet-5 | gpt-6-sol | Agent-initiated, low volume and reviewed by a person, so quality matters. |
| Document embeddings | none (Anthropic has none) | text-embedding-3-small, 512 dimensions | Small vectors keep search fast on shared hosting. |
- A workspace can pick another provider or model for each feature under AI.
- The platform administrator sets the defaults under Administration → AI → Default models.
- Drafts use adaptive thinking at medium effort; auto-replies use low effort on models that support it (Claude Haiku 4.5 doesn't take an effort setting).
Prices and cost
Every call is costed from Administration → AI → AI prices (USD per million tokens: input, output, cache write, cache read). The release ships the providers' standard list prices, last verified against their pricing pages on 26 September 2026. The page shows that date.
These are a snapshot, not live prices. Providers change prices and run promotions, so check the pricing pages and edit a row when they change. Desorix deliberately has no live price feed: the providers don't publish one, and scraping their pages would break silently.
| Model | Input | Output | Cache write | Cache read |
|---|---|---|---|---|
| claude-haiku-4-5 | 1.00 | 5.00 | 1.25 | 0.10 |
| claude-sonnet-5 | 2.00 | 10.00 | 2.50 | 0.20 |
| claude-opus-5 | 5.00 | 25.00 | 6.25 | 0.50 |
| gpt-6-luna | 0.10 | 0.50 | 0.125 | 0.01 |
| gpt-6-sol | 2.00 | 10.00 | 2.50 | 0.20 |
| gpt-6-astra | 10.00 | 50.00 | 12.50 | 1.00 |
| gpt-5-mini | 0.25 | 2.00 | none | 0.025 |
| text-embedding-3-small | 0.02 | none | none | none |
| text-embedding-3-large | 0.13 | none | none | none |
- Standard rates for short prompts. OpenAI charges more for long-context requests (about double) and for Fast mode. Desorix's prompts are short and use standard processing.
- Cache writes: Anthropic and GPT-5.6-or-later OpenAI models charge 1.25× the input price to write a prompt to the cache, and reads cost a tenth of it. Both providers report these tokens separately, and Desorix prices them separately.
- Edit a row when a provider changes its prices, or add a model with its own price. Edited rows are marked edited, and Restore bundled prices puts the shipped values back.
- A model without a price is logged at cost 0 with a note saying so.
- Every call, including failed ones, is recorded with its tokens, cost, model, key (workspace or platform) and outcome. The cost dashboard shows the totals.
Business description and documents
AI → About your business: what you sell, where, opening hours, delivery, tone of voice, and anything the AI must never promise. Auto-reply can't be switched on without it.
AI → Documents: FAQs, price lists, policies.
- File types: PDF, Word (.docx), text, Markdown or CSV up to 10 MB, or pasted text. Scanned PDFs (images of text) can't be read; paste their text instead.
- Indexing: documents are split into passages of about 300 tokens and indexed in the background. The list refreshes itself: Queued (waiting for the background worker), then Indexing…, then Ready. On cron-only hosting a job normally starts within a minute. Jobs run in order, so earlier jobs can delay it. After 3 minutes in the queue, the page says so and points to Administration → System. While a document is re-indexed, its previous index is still used.
- With an OpenAI key: each passage gets a vector, and the passages closest in meaning to the customer's message are given to the AI. The vector records the model and dimension count it was made with. After an administrator changes either, old documents are marked needs re-indexing and searched by keywords until you click Re-index, so vectors of different sizes are never compared.
- Without one: keyword search (MySQL/MariaDB FULLTEXT).
- Memory: search streams passages 200 at a time, so memory stays small (well under 2 MB) whatever the size of the knowledge base. Above 3,000 passages, a keyword pre-filter picks the 1,000 best candidates first.
Auto-reply
AI → Auto-reply: switch it on, choose the numbers (none ticked = all), the model, the confidence threshold (default 0.75), the maximum replies per conversation (default 5), and an optional handover message.
What happens to a customer message:
- Automation runs first: STOP/START, a paused conversation, a chatbot flow waiting for an answer, then keyword, welcome and away rules (see Automation).
- If none of them answered, the AI gets the business description, the matching document passages and the last messages of the conversation. It returns an answer, a confidence from 0 to 1, and whether a person is needed.
- If the answer is confident enough and no person is needed, it is sent, labelled AI assistant in the inbox.
- Otherwise the conversation is handed to the team, and the bot pauses for the agent pause time (24 h by default):
- the handover message is sent, if you set one (for every reason: needs a person, no answer, not confident enough, reply limit or error; not when the spend limit stops it);
- the conversation moves to Open;
- the thread shows why: a person is needed, no answer, not sure enough (with the confidence and the threshold, e.g. 0.40 against 0.75), reply limit, spend limit or error, followed by the AI's own reason, as in Try it;
- the Chatbot panel says the AI handed the conversation over, so it's clear who paused the bot.
Guards:
- Monthly spend limit (see below).
- Reply limit: at most N AI replies per conversation since a person last replied.
- Burst guard: three AI replies within two minutes usually means another bot is answering, so a person takes over.
- Errors and timeouts: a provider error, or a timeout after 30 seconds, hands over rather than leaving the customer waiting.
Replies are generated while the incoming message is processed, which on shared hosting happens right after Meta's webhook. Customers usually get an answer within a few seconds.
Try it on the AI page asks a question as a customer would. It shows the answer, the confidence, whether it would be sent, the cost and the passages used. Nothing is sent to WhatsApp.
What an auto-reply costs
Each auto-reply has two costs: the AI tokens, and a WhatsApp service message. From 1 October 2026 Meta charges service messages after a free allowance of 1,000 delivered service messages per number per month (see Costs). The AI page shows the last 30 days' cost per reply, and the dashboard's AI auto-reply row shows both costs together.
Monthly spend limit
AI → Monthly spend limit is a hard ceiling in the pricing currency.
- What it counts: Auto-reply (AI tokens plus the WhatsApp cost of AI replies), or All AI use, which also pauses drafts when the limit is reached.
- Before each auto-reply Desorix checks this month's spend plus the typical cost of one more reply. When that reaches the limit:
- auto-reply pauses and the conversation goes to your team;
- owners and admins get an email;
- the AI page shows a banner, and each stopped reply is logged with the reason.
- Lifting the pause: it lifts at the start of the next month, or now if an owner raises the limit and clicks Resume.
- Other currencies: when the pricing currency isn't USD, a limit needs the administrator's USD exchange rate (Administration → Pricing).
Privacy
- Messages, the business description and document passages are sent to the chosen provider.
- Provider data-retention terms apply; check them for your region and customers.
- Delete personal data (GDPR) on a contact removes their messages from Desorix, but not from the provider's logs.