The Markin ARPU report for B2C enterprisesRead now
MARKIN
Field notes
Voices9 min read

What a top data science team looks like in an agentic era

Notes from conversations with heads of growth and analytics at eight enterprises rethinking how their data teams spend the week.

Team Markin
  • #Data teams
  • #Agentic AI
  • #Operating model
What a top data science team looks like in an agentic era

Over the last two quarters we sat down with heads of growth, analytics and data science at eight enterprises rethinking how their teams spend the week. The composition of the teams has barely changed. What they do with their days has changed completely. These are notes from those conversations, with the specifics anonymized and the patterns kept.

The week that used to be

In 2022 the archetypal week at a top B2C data science team looked roughly the same everywhere. Monday and Tuesday went to pulling and reconciling data for the standing dashboards. Wednesday was spent in reviews with the business, defending or explaining last week’s numbers. Thursday was the one day of actual modeling work. Friday was the model going into a slide for the following Monday.

That week produced roughly one shippable insight per person per month, generously counted. Nobody was happy with it, including the people doing the work. It was, however, the equilibrium the tooling allowed.

The week now

At the eight teams we studied, the week has reorganized around a different unit of work. It is not the report. It is not the model. It is the shipped, causally read experiment. Every activity on the calendar traces back to raising the rate at which those get produced.

  1. Monday. Portfolio review. Which experiments shipped last week, which won, which lost, which to scale, which to retire. Thirty minutes.
  2. Tuesday and Wednesday. Hypothesis authorship. The team writes, in structured form, the fifty to eighty new hypotheses that will enter the queue for the following week. Peer review is inline.
  3. Thursday. Deep work on the underlying models, priors and guardrails that the whole system depends on. This is the day that most resembles the old week, and it is the only one.
  4. Friday. Business review focused on the metric tree, not on individual campaigns. The unit is the incremental ARPU number, not the send.

Four shifts that made it possible

1. Reports stopped being deliverables

In every one of the eight teams, the weekly BI report is no longer produced by hand. It is generated, checked and published automatically. The analyst’s job is not to produce the report. It is to interrogate it when a number looks off. That reclaimed roughly one day per person per week.

2. Data pulls stopped being ad-hoc

The pattern that used to consume Monday, a business stakeholder asks for a slice, the analyst pulls it, the stakeholder asks a follow-up, was replaced by a self-service surface over the same underlying model. The stakeholder pulls their own slices. The analyst is called in for the modeling questions, not the fetching ones.

3. Hypothesis quality became a first-class skill

A well-authored hypothesis is now the highest-leverage artifact in the team. It scopes the segment, states the expected direction, names the guardrail metrics and cites the priors. Bad hypotheses waste the entire downstream pipeline. The teams that treat authorship as a craft materially outperform the teams that treat it as an afterthought.

4. Causal reads replaced dashboards as the currency

The conversation the head of data has with the head of growth used to be about dashboards. Now it is about experiments. Nobody argues about a chart in isolation; every argument grounds out in the incrementality readout of a specific shipped variant. This is the single largest cultural shift, and the one the interviewees mentioned most often unprompted.

What the org chart looks like

The teams that made this shift did not grow. Two of the eight actually got smaller by attrition and did not backfill, and their throughput went up anyway. What changed is the internal shape.

  1. A hypothesis desk. Two to four people whose only job is to keep the hypothesis queue full, well-scoped and prioritized.
  2. A causal squad. One or two senior ICs who own the measurement stack, the holdout discipline and the incrementality readout.
  3. A model group. The classical data science bench, now smaller and more concentrated on the highest-leverage models rather than on report pipelines.
  4. An opportunity curator. Often a product manager, sometimes the head of growth, who owns the queue of revenue opportunities as a product surface.

The metric that mattered

Every team we spoke to had settled on a version of the same internal metric: shipped, causally read experiments per analyst per quarter. In 2022 the median for this group was two. In 2026 it is between twenty and forty, depending on the surface. That is the productivity delta, and it is the one number that predicts the ARPU trajectory of the business over the following two quarters.

What did not work

Three attempts came up repeatedly as false starts. Worth naming, because they are the natural first moves and they do not pay back.

  1. Rebranding the team as “AI.” No throughput change. The bottleneck was not the label on the team, it was the number of experiments the pipeline could ship per week.
  2. Adding one more BI tool. Made the reporting problem visibly worse. Reports are not the constraint; the amount of time humans spend maintaining them is.
  3. Hiring more analysts against the same operating model. Doubled headcount, roughly doubled report output, moved the incrementality curve by zero.

What the interviewees told us next

The consistent throughline from these conversations was the shift from producing artifacts to producing rate. Rate of hypotheses, rate of shipped experiments, rate of validated learnings. Every architectural choice downstream of that framing pointed in the same direction.

For the operating model these teams migrated into, see From campaign calendars to continuous decisioning.

Frequently asked

Questions readers ask about this.

How is a data science team changing in the agentic era?
The share of time spent producing dashboards and one-off analyses is shrinking. The share spent authoring, stress-testing and adjudicating hypotheses generated by agents is growing. Team structures are shifting from centralized reporting pods to reviewer roles embedded with revenue owners.
Does agentic AI reduce data science headcount?
In the enterprises we spoke with, no. It reallocates it. The number of decisions the team can support goes up materially while the headcount stays flat or grows modestly.
What skills matter most for a data scientist in 2026?
Causal reasoning, experiment design, hypothesis adjudication, and the ability to critique an agent's proposal against domain context. Coding fluency remains table stakes but is no longer the differentiator.

See it in the product

This runs in Markin today.

The same loops this note describes run 24/7 against your customer base. Watch the workspace decide, experiment and execute 1:1.

Explore the product