Technology & SaaSPixis

AI infrastructure for the AI marketing platform.

Model serving, evaluation, and the attribution loop underneath Pixis's AI marketing platform — with the infrastructure discipline to sustain a fast-growth SaaS trajectory.

Marketing analytics dashboard on a laptop

Client

Pixis

Sector

Technology & SaaS

Duration

10 months

Team

7 (AI engineers, an SRE embed, a platform architect, data engineers)

The client

About Pixis.

Pixis is a codeless AI-native marketing platform used by growth teams at consumer brands to plan, launch, and optimise paid media across channels. Between Series A and Series C the company grew from a handful of pilot customers to a global roster spanning D2C, FMCG, and category-defining consumer brands. The AI platform that had been prototyped by the founding data science team was now being asked to run at SaaS scale — with the cost discipline, evaluation rigour, and serving posture that trajectory demands.

Context

Where Pixis was when we started.

Pixis is an AI-native marketing platform serving growth teams at consumer brands. As the company scaled from Series A to Series C, the internal AI infrastructure needed to move from 'a few models in a notebook' to production infrastructure with the evaluation, cost discipline, and serving profile a SaaS company requires.

Challenge

The problem, unvarnished.

  • Model serving cost was scaling faster than revenue as customers grew — the unit economics that worked at Series A did not work at Series C.
  • Evaluation was manual — every new model release depended on human review by the DS team, making the DS team the release bottleneck.
  • Attribution model outputs varied between customers in ways the team couldn't debug at speed, and customer-success conversations were being lost to 'why did the model do that' questions.
  • The AI programme was operating without the SRE and cost-discipline patterns SaaS scale demands — on-call was ad-hoc, incident review was informal, capacity planning was reactive.
Approach

How we scoped and sequenced the work.

01

Cost-aware model serving with tiered routing

Migrated model serving to a tiered infrastructure — hot path on GPU, warm path on quantised CPU, cold path on batch — with routing based on latency budget and customer plan. Serving cost per query fell without touching the product.

02

Evaluation as a continuous product surface

Built the eval platform: golden datasets per model family, continuous eval on every deploy, regression alerts on drift. Every new model has to pass its eval to ship. The DS team stopped being the release blocker.

03

Attribution consistency across customers

Standardised the attribution model surface across customer segments. Where per-customer variation was necessary, it was made explicit and auditable — not a mystery in the config.

04

SRE embed for AI infrastructure

Embedded an SRE with the AI platform team for six months. On-call practices, incident review, capacity planning, and cost dashboards — the boring operational discipline that AI infra needs at scale.

Timeline

How the engagement unfolded.

01

Months 1–2

Diagnostic & unit-economics audit

Audited the model-serving path, evaluation practice, and unit economics. Mapped serving cost per query by customer segment; identified the model families driving 80% of the cost. Aligned CTO and CFO on the target unit economics.

02

Months 2–6

Serving migration & tiering

Migrated serving to a tiered infrastructure. GPU hot path for latency-critical calls; quantised CPU warm path for the majority; batch cold path for backfills. Routing decisions moved to config. Cost per query fell 62%.

03

Months 4–8

Evaluation platform

Built the eval platform. Golden datasets per model family, continuous eval on every deploy, regression alerts on drift. Release gate moved from DS review to eval pass. Release cadence increased.

04

Months 5–10

SRE embed and platform-team enablement

SRE embedded with the AI platform team. On-call practices, incident review ritual, cost dashboards, capacity planning ritual. AI platform team took ownership of the operational surface.

Solution

What we shipped.

A production AI infrastructure stack underneath Pixis's marketing platform, with tiered serving, continuous evaluation, and the SRE discipline to sustain SaaS growth. Attribution model surface standardised across customer segments.

Architecture & decisions

The choices behind the build.

01

Tiered serving by latency budget and customer plan

Serving isn't one path — it's three (hot/warm/cold), with routing decided by the customer's plan, the surface's latency budget, and model criticality. Cost per query is a routing decision, not a model decision.

02

Quantised CPU is the majority path

Most of Pixis's inference load is not latency-critical. Quantised CPU serving handles it; GPU is reserved for the surfaces where p99 latency matters. This choice alone drove most of the cost reduction.

03

Golden set per model family, not per model

Every model family (attribution, creative recommendation, budget allocation) has its own golden set curated by product SMEs. New models pass or fail the golden set before they ship. DS team is not the release gate.

04

Cost dashboards owned by the AI platform team

Cost visibility is a first-class dashboard, not a monthly finance report. The platform team can see per-customer, per-model, per-endpoint cost in real time — and act on it.

Rollout & adoption

How it landed inside the organisation.

The serving migration was invisible to customers by design. We ran the new tiered path in shadow for four weeks, comparing outputs against production, before flipping the routing. The evaluation platform rolled out model-family by model-family, starting with the surfaces where the DS team's manual review was hurting the most. SRE practices landed through embedding, not documentation — the SRE sat with the platform team, ran the first incidents, and left when the platform team could run the third one on their own.

Outcomes

The numbers that matter.

62%

Reduction in serving cost per query

Achieved through tiered serving and model quantisation without measurable degradation in customer-facing latency. Held over two quarters post-migration.

0

Human release-review gates

Continuous evaluation replaced manual DS review as the release gate for model updates. Release cadence increased from bi-weekly to multiple times per week without incident regression.

14

Attribution channels standardised

Consistent attribution surface across customer segments with explicit, auditable per-customer variation where required. Customer-success 'why did the model do that' questions dropped to near-zero.

Tech stack

What we built it on.

PythonPyTorchTensorRTKubernetesAWS SageMakerDatabricksSnowflakeTerraformDatadogPagerDuty
Reflection

What we learned.

The infrastructure work is the visible part; the operational-culture shift is the durable part. What actually kept Pixis's unit economics healthy through the growth wasn't the tiered serving stack — it was the AI platform team learning to think in cost-per-query and eval-pass terms every day. That's the shift that survives a growth round; the stack is replaceable, the operating habit is not.

What's next

The engagement today.

The engagement moved from build to advisory in month ten. Ongoing collaboration focuses on the next-generation model-serving infrastructure and the eval-platform roadmap. Attribution surface is being extended into a shared attribution API that Pixis's customers can plug into their own dashboards.

The move from 'a few models in a notebook' to production infrastructure with real SRE discipline was the operational shift we needed to sustain the trajectory. Headify stayed until our platform team owned it end to end.

VP, AI platform · Pixis (name withheld under engagement confidentiality)
Delivered by

Practices involved.

Products on top

Productised offerings involved.

Talk to the team

Bring us your technology & saas brief.

Book a 30-minute discovery call. A senior practitioner from the same practice that shipped this engagement will scope yours.