Travel & eCommerceBooking.com

AI experimentation infrastructure at travel-industry scale.

Experimentation and attribution infrastructure for a top-three OTA — running tens of thousands of concurrent experiments across paid channels and product surfaces.

Traveller booking a trip on a mobile device

Client

Booking.com

Sector

Travel & eCommerce

Duration

22 months

Team

14 (senior product engineers, data engineers, ML infra specialists, an SRE embed)

The client

About Booking.com.

Booking.com is one of the top three online travel agencies globally, with tens of millions of nights booked per year across accommodation, flights, cars, and experiences. Its product culture is famously experiment-driven — most surfaces on the site are simultaneously under test, and the marginal cost of a bad decision is measured in eight-figure booking swings. The company's data science and platform organisations run at a scale that resembles the physics of a small country.

Context

Where Booking.com was when we started.

Booking.com operates one of the world's largest continuous experimentation programmes. Every product surface, ranking model, and paid-channel touchpoint is under experiment simultaneously, and the marginal cost of a bad decision — measured in bookings — is nine figures a year. As the company's AI investment scaled, the experimentation platform started to need capabilities the original architecture wasn't built for.

Challenge

The problem, unvarnished.

  • Attribution across paid, organic, and product touchpoints — with cookie deprecation making the traditional multi-touch model unreliable and finance no longer accepting last-click reads.
  • Experiment concurrency at a scale where naive analysis leaks between variants and interaction effects go undetected until decisions have already shipped.
  • Model-serving latency budgets under 40ms at global peak, across dozens of ranking and recommendation models — with cost-per-inference sitting on top of a shrinking margin per booking.
  • A growing gap between experiment-platform capability and what data scientists were being asked to test — the platform had become a release bottleneck for the DS org.
  • Political overhang: the platform team's roadmap was being pulled in five directions by different product organisations, and no shared prioritisation existed.
Approach

How we scoped and sequenced the work.

01

Causal attribution on first-party signal

We rebuilt attribution around causal inference on first-party touchpoints, augmented with channel-level media mix modelling. Weekly refresh; monthly backtest against holdout markets; a shared read that the finance team accepts.

02

Concurrency-safe experiment design with interaction detection

Redesigned the experiment planning surface to enforce concurrency-safe designs. Interaction detection now surfaces variant interactions to the DS team before results are read — not after a bad decision has shipped globally.

03

Low-latency model serving at global scale

Rebuilt the model-serving path with edge caching, model quantisation, and traffic-shaping to hold 40ms p99 latency budgets across every region. Model rollouts moved from weekly to continuous.

04

Platform-team enablement, not model-team ownership

We stayed embedded until Booking's platform team could operate, evaluate, and extend the system without us in the loop — including running the prioritisation ritual across the five downstream product orgs.

Timeline

How the engagement unfolded.

01

Months 1–3

Diagnostic & platform mapping

Diagnostic sprint across the experimentation surface, attribution stack, and serving path. Identified the three highest-leverage bets: attribution rebuild, concurrency-safe experiment design, and model-serving hardening. Aligned the five product orgs on a shared prioritisation ritual.

02

Months 4–10

Attribution rebuild & finance acceptance

Rebuilt attribution around causal inference on first-party events, augmented with MMM. Ran a six-week validation against holdout markets and the CFO's finance model. Finance signed off; paid-channel reallocation began.

03

Months 8–14

Concurrency-safe experiments & interaction detection

Redesigned experiment planning surface. Interaction detection lands automatically in every read-out. Concurrent-experiment ceiling raised from low-thousands to 10,000+.

04

Months 12–18

Serving path rebuild & continuous rollout

Model serving path rebuilt with edge caching, quantisation, traffic shaping. Continuous rollout replaces weekly. 40ms p99 latency held globally.

05

Months 18–22

Handoff & platform-team run

Platform team took operational ownership; embedded engineers rotated off. Booking's DS org now runs prioritisation and roadmap without external input.

Solution

What we shipped.

An upgraded experimentation, attribution, and model-serving platform underneath Booking.com's product and marketing surfaces. Concurrency-safe experiments, causal attribution grounded in first-party data, and model-serving that meets global latency budgets.

Architecture & decisions

The choices behind the build.

01

Causal first, MMM second

Rejected pure MMM as the attribution source of truth — causal inference on first-party events is the primary signal, MMM is used to bound the estimate and provide a channel-level media contribution number the finance team can plug into planning spreadsheets.

02

Interaction detection as a default read-out surface

Every experiment read-out surfaces variant-interaction detection automatically. If two experiments interact and the DS team hasn't declared it, the read-out flags it before the decision is made.

03

Edge cache + quantised model + shape-then-route

Serving path splits into hot (GPU, unquantised), warm (CPU, quantised), and cold (batch). Route decision happens on edge, based on customer plan, latency budget, and model criticality. Cost per inference fell without touching the product.

04

Continuous rollout with a canary bucket

Model rollouts moved from weekly to continuous, with an always-on canary bucket that catches regressions before global exposure. Rollback is one config-file flip.

Rollout & adoption

How it landed inside the organisation.

Because Booking's culture is experiment-driven, we shipped the new experimentation surface as an opt-in overlay for the first quarter — DS teams could still run their old flow if they wanted. Adoption ran at ~90% within eight weeks because the interaction-detection surface caught two would-have-been-bad decisions in the first month, and the story spread inside the org fast. Attribution rebuild landed differently: it was a CFO-sponsored change, and we ran a six-week validation against the finance model before the switch was flipped.

Outcomes

The numbers that matter.

2.4x

Improvement in blended paid ROAS

After reallocation informed by the new causal attribution model, measured over two quarters against the pre-change baseline, with a hold-out market segment used as the control.

40ms

P99 model-serving latency, global

Held across all regions during peak booking-season traffic, on a model surface that previously breached budget on a weekly basis.

10,000+

Concurrent experiments

With interaction detection surfaced to the data science team automatically — up from a low-thousands ceiling on the prior platform, without the interaction-leak risk that limited it.

Tech stack

What we built it on.

PythonGoJavaSnowflakeDatabricksKafkaKubernetesTensorFlowPyTorchRedis
Reflection

What we learned.

The single hardest part of the engagement was not technical. Booking has five product organisations with independent roadmaps, and the platform team was being pulled in every direction. Standing up a shared prioritisation ritual — one that all five VPs sponsored — was the change that let the technical work land. If we'd tried to ship the technical work without solving the political overhang, we'd have burned twelve months and gotten a great platform nobody used.

What's next

The engagement today.

Booking's platform team runs the surface independently. Ongoing collaboration is at a lower cadence, focused on the next generation of ranking-model architecture and the extension of causal attribution into loyalty and post-booking touchpoints. Interaction-detection library is being open-sourced with our involvement.

The concurrency-safe experiment design and the causal read on paid attribution were the two things we'd been trying to build in-house for two years. Getting them both, in production, from a partner that stayed for hardening — that was the delta.

Growth engineering lead · Booking.com (name withheld under engagement confidentiality)
Delivered by

Practices involved.

Products on top

Productised offerings involved.

Talk to the team

Bring us your travel & ecommerce brief.

Book a 30-minute discovery call. A senior practitioner from the same practice that shipped this engagement will scope yours.