top of page

How Companies Are Using Experiments to Find Winning Growth Strategies

2 days ago
9 min read

Industry & Competitive Context

The online travel industry became one of the most data-intensive consumer sectors in the world over the past two decades, with firms competing on real-time pricing, inventory breadth, and user-experience design rather than on traditional brand advertising alone. Booking Holdings (formerly Priceline Group), the parent of Booking.com, Priceline, Agoda, Kayak, and OpenTable, grew into the industry's largest player by gross travel bookings, which reached $150.6 billion in 2023, up 24% from $121.3 billion in 2022, according to the company's SEC filings and year-end earnings release. Total revenue for 2023 was $21.4 billion, a 25% increase over the prior year, with net income rising 40% to $4.3 billion. The company also reported surpassing one billion room nights booked in a single year for the first time in 2023.

This scale was built in a sector where consumer choice is driven by marginal differences in page design, search ranking, pricing display, and checkout friction, making the online travel agency (OTA) business one of the earliest and most intensive adopters of large-scale digital experimentation. Harvard Business School professor Stefan Thomke, who has studied the company extensively and co-authored a Harvard Business School case on Booking.com with Daniela Beyersdorfer, has documented that Booking.com runs approximately 25,000 tests a year and, at peak capacity, can run roughly 1,000 concurrent experiments across 75 countries and 43 languages. Thomke's research, published in Harvard Business Review ("Building a Culture of Experimentation," March–April 2020) and in his book Experimentation Works: The Surprising Power of Business Experiments (Harvard Business Review Press, 2020), positions Booking.com alongside Microsoft, Google, Amazon, Netflix, and Facebook as one of the handful of firms that run more than 10,000 controlled online experiments annually.

Infographic on companies using experiments for growth, showing strategy, A/B tests, data analysis, and winning strategies.

Brand Situation Prior to the Experimentation-Led Approach

Booking.com was founded in 1996 by a Dutch university student and grew slowly for close to a decade before being acquired by Priceline in 2005, according to Thomke's HBR-published research and the Harvard Business School case. The company's transformation into the world's largest accommodation platform is explicitly attributed by Thomke to its adoption of systematic, democratized experimentation rather than to a single marketing campaign or product launch. Prior to this cultural shift, decisions about site design, content, and features were made in the conventional way documented across most large organizations: by senior designers, product leaders, and marketing executives relying on professional judgment, competitive benchmarking, and internal opinion rather than controlled testing at scale.

A pivotal, well-documented moment illustrating both the opportunity and the risk of this approach occurred in December 2017, when Booking.com's director of design proposed testing a radically simplified homepage just before the peak holiday travel season. The existing homepage, which had been optimized incrementally over years, featured extensive content, images, and promotional messaging. The proposed alternative stripped nearly all of this away, leaving only a minimal search window asking for destination, dates, and party size. This episode, documented in Thomke's HBR article and in the HBS "Cold Call" podcast featuring Thomke, is used as the entry point for explaining how the company's testing culture functioned under real commercial pressure.


Strategic Objective

Based on the publicly documented material, Booking.com's strategic objective was not a single campaign goal but an organization-wide commitment to replacing opinion-driven decision-making with evidence-driven decision-making at scale, across design, pricing, merchandising, and product features. Thomke's account states explicitly that the company sought to build a culture in which "curiosity is nurtured, data trumps opinions, any employee can launch tests, all experiments are ethical, and a new, more democratic model of leadership prevails." The intent documented in the source material was to make experimentation a default operating mechanism for growth decisions, rather than a specialized function confined to a research team.


Campaign Architecture & Execution

The architecture of Booking.com's experimentation system, as documented by Thomke, rests on radical decentralization of testing authority. According to his HBR article, any employee at the company can launch a live experiment on millions of customers without requiring managerial sign-off, and approximately 75% of the company's roughly 1,800 technology and product staff actively use the internal experimentation platform. Any employee can also halt a running experiment at any time if it appears to be harming the customer experience. Thomke's research describes this as a deliberate departure from traditional top-down approval processes, built on the premise that the volume of good ideas generated by a large, empowered workforce exceeds what a centralized team could originate or evaluate on its own.

The execution model documented in the sources follows the standard architecture of a controlled online experiment: a hypothesis is formed, traffic is split between a control version and one or more test variants, and statistically significant outcomes determine whether a change is rolled out, modified, or discarded. Thomke's account of the December 2017 homepage test illustrates the discipline applied to execution rather than replacing the existing homepage outright, the company tested the radically simplified version against the incumbent design with a defined customer segment, measured the results, and used the data, not internal preference, to decide the outcome. Thomke also documents the operating principle that validated test results are implemented with few exceptions, quoting an internal director's description of the mindset: "If the test tells you that the header of the website should be pink, then it should be pink."

The company's scale of execution is substantial relative to peers. Thomke's reporting places Booking.com's testing volume at approximately 25,000 experiments annually and notes the technical capacity to run around 1,000 experiments simultaneously across dozens of countries and languages, which required significant investment in standardized experimentation infrastructure and continuous data-quality monitoring, as noted in third-party analysis of Thomke's book Experimentation Works.


Positioning & Consumer Insight

The public source material does not document Booking.com's experimentation program as a brand positioning or advertising campaign in the traditional marketing sense; it functioned instead as an internal operating philosophy applied to product and user-experience decisions. The consumer insight embedded in this approach, as articulated by Thomke, is that customer behavior measured through actual clicks, bookings, and engagement is a more reliable guide to what drives conversion and retention than internal opinion, however senior or experienced the source. This is reflected in the internal maxim, documented by Thomke, that the company resists implementing even senior executives' design preferences unless those preferences are validated by test data.

Comparable publicly documented examples from other large technology companies reinforce this insight at an industry level rather than at the level of a single firm's campaign. Netflix's engineering and data science teams have published, through the company's own technology blog, that the company evaluates product and user-interface changes including artwork and thumbnail selection, row ordering, and recommendation algorithms through controlled A/B tests before any change becomes the default experience for subscribers. Netflix's technology blog states that this empirical approach exists specifically so that "product changes are not driven by the most opinionated and vocal Netflix employees, but instead by actual data, allowing our members themselves to guide us toward the experiences they love." This same underlying insight that aggregated customer behavior, not internal conviction, should arbitrate product and growth decisions is the common thread Thomke identifies across Microsoft, Google, Amazon, Netflix, and Booking.com in his HBR research.


Media & Channel Strategy

No verified public information is available describing a discrete media or advertising channel strategy tied specifically to Booking.com's experimentation program, because the publicly documented material characterizes experimentation as a product-and-design discipline applied primarily to the company's own website and app, not as a paid media campaign. The channels involved, as described in the source material, are Booking.com's owned digital properties its website and mobile application across which test variants are served to segmented user traffic.

Where channel-level experimentation has been publicly documented in the broader industry, it relates to measurement methodology rather than creative strategy. Public commentary on Netflix's approach to marketing measurement, drawing on the company's own published engineering and data science material, describes the use of geographic experiments randomizing media markets rather than individual users to measure the effect of channels such as television advertising that cannot be tested through conventional user-level A/B testing. This is documented as an extension of the same experimentation discipline into channels where individual randomization is not feasible, rather than as a separate campaign strategy.


Business & Brand Outcomes

The only outcomes that can be stated with direct evidentiary support are the company-level financial results disclosed in Booking Holdings' SEC filings and earnings releases, and the scale metrics of the experimentation program itself as reported by Thomke. These should be read as parallel, publicly documented facts about the same company and period rather than as a demonstrated causal chain, since none of the source material provided a quantified, audited link between a specific experiment and a specific financial result.

On the experimentation program itself: Thomke's HBR-published research states that by running approximately 25,000 tests a year, Booking.com "transformed itself from a small start-up to the world's largest accommodation platform." This is a direct, attributable claim from a credible, named academic source published in Harvard Business Review, though it is presented by Thomke as a broad characterization of the company's trajectory rather than as an isolated, quantified return on any individual test.

On company financial performance, Booking Holdings' own disclosures show total revenue grew from $17.0 billion in 2022 to $21.4 billion in 2023, a 25% increase, while gross travel bookings grew 24% to $150.6 billion and room nights booked surpassed one billion for the first time, growing 17% year-over-year. Full-year 2023 net income rose 40% to $4.3 billion, and adjusted EBITDA increased 34% to $7.1 billion. CEO Glenn Fogel's statement accompanying the 2023 results described these as "record levels of gross bookings, revenue, and operating income." Booking Holdings' 2022 10-K similarly disclosed that 2022 revenue of $17 billion was the company's highest-ever annual revenue at that time, up 56% versus 2021. No verified public information is available quantifying what portion, if any, of this financial performance is specifically attributable to the experimentation program as opposed to broader post-pandemic travel demand recovery, pricing conditions, or other business factors, and the company's own filings attribute bookings growth primarily to travel demand recovery, room-night growth, and airline ticket volume rather than to the testing program.


Strategic Implications

The Booking.com case, as documented through Thomke's Harvard Business School research and the company's own public disclosures, illustrates several strategic implications for how firms use experimentation to pursue growth. First, the scale of experimentation at Booking.com tens of thousands of tests annually, run by a majority of the technology and product workforce without centralized gatekeeping suggests that the primary constraint on using experimentation as a growth engine is organizational and cultural rather than purely technical. Thomke's research explicitly frames the challenge as one of culture: building an environment where data is trusted over hierarchy, where failure is treated as a normal and expected outcome of most tests, and where authority to test is distributed broadly rather than concentrated in a specialist team.

Second, the comparable, independently documented practices at Netflix, Microsoft, Google, and Amazon all of which are reported by Thomke and by Netflix's own technology blog to run large volumes of controlled experiments indicate that this is an industry-wide pattern among large digital platforms rather than a practice unique to one company. The common structural feature across these documented cases is that experimentation is embedded directly into the product-development workflow, with test results treated as a required input, and in Booking.com's case, as described by an internal director quoted by Thomke, implemented "with few exceptions" once validated.

Third, the publicly available evidence supports a conclusion about process and capability-building more confidently than it supports a conclusion about precise financial causality. The documented material is strong on describing how these companies structured their experimentation systems and governance, and comparatively silent on providing audited, test-by-test return figures. This gap is itself instructive for case analysis: it highlights the distinction between an organization's demonstrated commitment to a growth methodology and the separate, harder question of isolating that methodology's standalone contribution to top-line results, which public companies generally do not disclose in a decomposed form.


Discussion Questions

Booking.com allows any employee to launch a live experiment on millions of customers without managerial approval. What organizational safeguards would be necessary to make this level of decentralization viable, and what risks does it introduce that a centralized testing function would otherwise manage?

Thomke documents that Booking.com implements statistically validated test results "with few exceptions," even when they contradict senior executives' preferences. What does this reveal about the relationship between data-driven decision-making and organizational authority, and where might this norm create tension in a traditional corporate hierarchy?

Booking Holdings' revenue and bookings growth in 2022–2023 coincided with post-pandemic travel demand recovery. Given the available public evidence, how should a strategist distinguish between growth driven by experimentation-led optimization and growth driven by macro-level market recovery?

Netflix uses geographic (market-level) experiments to measure media channels, such as television, where individual-level A/B testing is not feasible. What does this suggest about the limits of conventional A/B testing as a growth methodology, and how might other industries adapt this approach?

Across Booking.com, Netflix, Microsoft, Google, and Amazon, large-scale experimentation is consistently described as a cultural capability rather than a single campaign or initiative. What factors might explain why some large organizations successfully build this capability while others struggle to scale experimentation beyond a specialist team?


Comments


bottom of page