Most personalization platforms let you measure lift however you like, which in practice means most customers measure it badly. Auxia builds global holdout groups into the decisioning engine itself. That one decision is why this review scores it where it does.

The measurement problem it solves

The standard failure in this category is comparing personalized visitors against a previous period. Conversion goes up, personalization gets the credit, and nobody separates the effect from seasonality, a pricing change, or the campaign that launched the same week. Every vendor’s case study numbers are vulnerable to this, which is why they are so consistently impressive.

A global holdout is the fix: a permanent slice of traffic that never receives personalization, so you always have a concurrent control. It is not technically difficult, and almost every platform leaves it as an option the customer must remember to configure and then defend internally when someone asks why a portion of traffic is deliberately unoptimized.

Building it into the engine changes who bears that discipline. That is worth more than several features.

What the product actually is

Two halves. The Decisioning Engine makes real-time per-customer decisions, ranking the next-best experience against available signals, backed by an ML feature store, a model library and multi-model experimentation. That last piece matters because it lets competing modelling approaches be compared rather than assuming the vendor’s default is right.

Agent Studio is the orchestration half: analysing campaigns, identifying gaps, building briefs, pulling content from a CMS, routing decisions for approval through Slack, and executing recurring playbooks across tools including Braze and Salesforce Marketing Cloud.

Notably, it orchestrates across systems you already run rather than asking you to replace them. In enterprise marketing stacks that is the difference between a six-week integration and a two-year migration.

Where the language outruns the product

The agentic framing is doing more work than the described behaviour supports. What is documented is analysis, generation and approval routing with humans in the loop. That is good workflow automation with a model attached, and calling it an agent that runs the work across your stack invites an expectation of autonomy the approval step contradicts.

This is not deception so much as category fashion, and it is worth mentally translating before you brief a stakeholder on what you are buying.

What to establish before committing

Pricing is not published, and billing is decision-based, so cost scales with how many decisions the engine makes rather than with seats. Count decisions before the first call: monthly active users multiplied by realistic touchpoints. That number, not your headcount, is what the quote will be built from.

The other open question is statistical. Embedded holdouts are excellent; no public statement describes how significance is determined from them. Ask what test is applied and what threshold triggers a result being called, because a holdout read impatiently reproduces the problem it was meant to solve.