Nosto is a mature commerce personalization platform with a customer list that suggests it works: around 1,500 brands including Marc Jacobs, Muji, Vuori and Kylie Cosmetics.
Its strongest argument is the surface most competitors ignore.
Search is the interesting part
Most personalization platforms concentrate on recommendations and content blocks: what to show on the homepage, what to put beside a product, which banner a returning visitor sees. Those matter, and they operate on visitors whose intent you are guessing at.
Site search is different. A visitor who types a query has stated what they want. Ecommerce search is also frequently bad, failing on plurals, synonyms and near-miss spellings, sending people to empty results pages when the product exists. The baseline is low and the intent is high, which is the best combination for improvement.
Nosto covering personalized search alongside recommendations, category merchandising, email, bundles and post-purchase means one system spans the journey rather than three vendors disagreeing about who a customer is.
Category merchandising driven by performance metrics is the other quietly valuable piece. Manual category ordering decays: someone sets it, products sell out or stop converting, and nobody revisits it for a quarter.
The measurement gap
A/B testing is offered. Holdouts and incrementality are not mentioned anywhere.
For a platform whose entire proposition is lift, that is the gap that matters. The default way commerce teams evaluate personalization is to compare revenue before and after switching it on, which conflates the platform with seasonality, promotions, product launches and every other change in the same period. It reliably produces impressive numbers and tells you very little.
A permanent global holdout, a slice of traffic that never receives personalized recommendations, is the fix. Auxia, reviewed elsewhere in this category, builds it into the engine. Most platforms leave it as something the customer must configure and then defend internally when someone asks why a portion of traffic is deliberately unoptimized.
Ask whether Nosto supports a permanent holdout natively. Ask, too, how the lift figures in their case studies were calculated, because the answer to the second question usually reveals the answer to the first.
The absence of published statistical methodology behind the A/B testing is the same gap in a smaller form, and it runs through this whole category rather than being specific to Nosto.
Scale is a prerequisite
Recommendation engines learn from co-viewing and co-purchasing behaviour. A catalogue of 300 products and modest traffic does not generate enough of either for a model to beat a well-chosen bestseller list.
No guidance on minimum viable scale is published, which is common and unhelpful. Before evaluating, ask for a reference customer at roughly your traffic and catalogue size, and what the measured difference was against a non-personalized baseline. If the answer is a before-and-after comparison, you have learned something about how this vendor measures.
