This strand splits into three sub-literatures of very different evidential strength.
(a) The experimentation infrastructure — settled
Continuous large-scale controlled experimentation is ordinary industry practice. Ron Kohavi, Diane Tang and Ya Xu — experimentation leaders at Microsoft, Google and LinkedIn — report in Trustworthy Online Controlled Experiments (Cambridge University Press, 2020), and in "Online randomized controlled experiments at scale" (Trials 21:150, 2020), that those firms run "over 20,000 controlled experiments/year" (they caution that counting methods vary). The feed is, literally, a permanent A/B test.
(b) Direct efficacy evidence — settled, and strong
The cleanest proof that instrumented design produces operator-intended behaviour comes from dark-pattern experiments. Jamie Luguri and Lior Strahilevitz, "Shining a Light on Dark Patterns" (Journal of Legal Analysis 13(1):43–109, 2021), ran a nationally representative US randomized trial: mild dark patterns more than doubled sign-ups for a dubious identity-protection service versus a neutral interface; aggressive patterns roughly quadrupled them; effects compounded when stacked. The FTC's staff report Bringing Dark Patterns to Light (P214800, 2022) elevates this to the regulatory record and documents Credit Karma selecting an allegedly false "pre-approved" claim because A/B testing showed it maximised clicks. (Caveat: the doubling effect is one twice-run experiment on a single service type — robust, but generalising from one well-controlled setting.)
(c) The adolescent-mental-health controversy — the open wound
Here the literature does not converge. This is the field's live methodological war over effect sizes, and it is unresolved as of this writing.
The effect-size dispute · screen use & adolescent well-being
Skeptic poler ≈ <.05
Orben & Przybylski. A specification-curve analysis across three datasets (n ≈ 355,358) finds the association "negative but small, explaining at most 0.4% of the variation" — anchored, famously, as comparable to "eating potatoes" or "wearing eyeglasses" — and "too small to warrant policy change." Their time-use-diary study finds "little clear-cut evidence that screen time decreases adolescent well-being," "far removed from the certainty voiced by many commentators."
Harm poler ≈ .20
Haidt and colleagues. Argue the near-zero result is an artefact of six "defensible" analytical choices that collectively obscured an association nearer r = .20 — larger for social media specifically (2–6× the all-digital figure), r = .15–.22 for girls and "well above r = .20" for girls in early puberty — and point to a "hockey stick" 50–150% rise in US teen mood disorders, 2009–2019. The dispute remains active into 2026 (Sigaud, Rausch, McClean & Haidt).
Both camps actually agree the correlation exists and that teen mood-disorder rates rose sharply in the early 2010s. They disagree on magnitude, causation, and whether self-reported screen time is a valid measure. This is the precise locus of the "does it work at scale" question.
(d) Primary-source leaks
The Facebook Files / Frances Haugen disclosures, entered into the US House Energy & Commerce Committee record (22 September 2021), are the field's key primary documents — distinct from peer-reviewed scholarship. A March 2020 internal slide reported that "32% of teen girls said that when they felt bad about their bodies, Instagram made them feel worse"; a 2019 slide stated "We make body image issues worse for one in three teen girls"; another reported that teens "blame Instagram for increases in… anxiety and depression… unprompted and consistent across all groups." Meta disputes the framing, not the existence of the figures.
⚠ Excluded — failed verification
The widely-circulated claim that Facebook's internal research causally tied Instagram to suicidal ideation (13% of UK / 6% of US teens with such thoughts tracing them to Instagram) was refuted 3–0 in verification and is not used here. The body-image findings above are genuine and survive; the suicide statistic, as commonly stated, does not.