Reflections on How to AI UXR: What Research Leaders Have Told Us So Far
We are six episodes into How to AI UXR, the podcast series from The ResearchOps Review. Research leaders have been sharing real applications of AI in their workflow, all coalescing around some of the same conclusions on the 'how' for AI in UXR.
The short version
How to AI UXR is an eight-part video podcast from Kate Towsey and The ResearchOps Review, sponsored by Strella, running weekly from July 9 to August 27, 2026. Each episode is a 30-minute conversation with one research practitioner about something they built. Six episodes in, the guests differ on some things but agree on these principals: start with one painful task rather than a system, expect AI to move the hard thinking earlier rather than remove it, and treat the constraint as operational rather than a limitation of the models.
What the series is
Kate Towsey and The ResearchOps Review built How to AI UXR, a map for building AI-augmented research operations drawn from 562 data points and conversations with more than 50 research and ResearchOps professionals. We sponsored it, and we followed the work closely as it came together.
The map treats AI in research as a maturity path rather than a yes-or-no question.
At Crawl, teams use off-the-shelf LLMs for drafting, summarizing and clustering.
At Walk, they build custom systems: agents, RAG repositories, evals.
At Run, multi-agent pipelines reshape the practice.
The podcast asks people who are somewhere on that path to show their work. Kate sits down with one practitioner per episode. Priya Krishnan, our co-founder and COO, adds a short take on each conversation from the perspective of someone building the tools these researchers are testing.
Six episodes have aired. Here is what we keep hearing.
1. Start with one painful task, not a system
The common failure mode in these conversations isn't picking the wrong tool. It's scope.
Michelle Bejian Lotia of Ramp opened the series by arguing against the instinct to architect. Her advice was to skip the grand plan and "find one thing that's giving you pain today and try and take that pain away." Not the highest-value workflow or the one with the best ROI story. The one that hurts.
What makes that more than standard start-small advice is what came next. The system she ended up proudest of wasn't on any roadmap. It didn't exist as an idea when she opened her laptop that morning. It came out of what turned out to be technically possible against a problem in front of her. As she put it: "There was no grand strategy behind it."
That's uncomfortable for a planning-oriented team to hear, and we think it's the most practical thing in the series so far. You can't design an AI research system in the abstract, because you don't yet know what the tools are good at in your context. The system is a result of the experimentation, not a precondition for it.
Daniel Gottlieb of Microsoft CoreAI got there from a different angle. He learned by building things unrelated to work, including a website for his dad. His rule was to "find something for you to learn on and try that is fun for you, that isn't a chore," because then roadblocks stop feeling like falling behind. The skill transfers even when the project doesn't.
Read against the map, this is a case for taking Crawl seriously rather than treating it as the phase you skip. Teams that jump to Walk tend to build systems for a workflow they never examined closely enough to describe.
2. AI doesn't remove the thinking, it moves it earlier
Every guest who built something that worked described the same sequence. The hard part happened before the tool did anything.
Jordan Brinkman of ERGO NEXT was most explicit. His instruction is to stop as soon as you have a feel for the tool's constraints and map the process out first, specifically: "how would I do this workflow as a human? Because I need to then explain that to Claude in some way."
The work there isn't prompting. It's being able to articulate your own process precisely enough that someone else, human or model, could execute it. Most research workflows have never been written down at that resolution. They live in a practitioner's judgment, and they get exposed the moment you try to delegate them.
Jordan was also clear this isn't work a novice can shortcut, because evaluating what comes back requires knowing what good looks like. Rigor, in his framing, comes from the researcher going in afterward and "cleaning up those initial skills," refining the instructions based on outputs they were qualified to judge. The expertise doesn't disappear. It moves from doing the analysis to specifying and auditing it.
Nam Pham of DoorDash framed the same thing as patience. Their read on why people bounce off these tools is that the marketing promised you just use it and good things come out. The reality is closer to teaching: "you do have to be patient with it. Teach it to do the things you wanna do." Reliability is something you build over iterations, not something you receive.
They also offered a useful reframe. Much of what sounds like new territory, including harnesses, orchestration and constraints, maps onto things researchers and service designers have always done. "Sometimes it's a new language," they said, "but under the hood it can be the same thing." For anyone who feels behind, that's worth holding onto. Workflow design, constraint setting and knowing when an output is wrong are skills you already have. The vocabulary is what's new.
3. The bottleneck is operations, not models
Ask what's blocking AI adoption in a large research org and you tend to get an answer about governance rather than capability.
Daniel Gottlieb described the state of enterprise AI tooling as "the Wild West," with people building custom things faster than anyone can track where they live, who owns them, or whether they're maintained. He compared it to the early internet boom, and he isn't upset about it. He sees "a huge opportunity for ops… to find a way to better track this, keep them in one location."
That reframe is worth borrowing. Proliferation usually gets described as a mess to clean up. Daniel treats it as evidence of demand, and therefore as a mandate for ResearchOps. The teams handling this well aren't restricting tool use. They're building the connective tissue that makes scattered tools legible.
Dave Chen at 1Password answered the same problem with structure. His team sorts the research process into three zones. Automate covers work that's been meaningfully handed to AI, largely early-process tasks built on custom GPTs. Augment sits in the middle, running on a spectrum from mostly-AI to mostly-human. Keep human covers the rest. In his words, augment is "probably our biggest chunk," with "varying degrees of… this mix of AI and human in the loop."
That the middle zone is the biggest is the part we'd underline. Nobody across six episodes described a fully automated research practice, and nobody described refusing to automate anything. The question was never automate or don't. It's where each task sits on that spectrum, and who decides.
4. Researcher judgment is the job, including the right to push back
Llewyn Paine named something the field mostly doesn't say out loud. Researchers are agreeing to AI applications they privately think are bad.
Her account is that they're "afraid to push back on any AI, even patently bad AI applications," because they "don't want to look like a Luddite." Nobody feels expert enough to be certain they're right, so they second-guess the instinct and go along. Her line on the cost: "we don't want to get so focused on moving fast as researchers that we sacrifice what makes us effective as researchers."
The structural problem underneath is that a field adopting a technology this fast generates conventions faster than evidence for them. As Llewyn put it, "just because a practice is popular or common at this point doesn't mean that it's actually good research practice." Right now, common practice mostly reflects what was easy to do first.
This is the one theme in the series about organizational permission rather than technique. A researcher who can't say "this specific application is bad methodology" without being read as an obstacle isn't functioning as a researcher. Whatever else a team builds, it needs a way for that objection to be raised cheaply and taken seriously.
5. Guardrails are what make speed safe
Dave Chen's team put a research agent into a Slack channel where stakeholders ask it questions directly. The design choice worth noticing is the channel, rather than a private tool or a dashboard.
Putting it in public does two things at once. Stakeholder questions become visible, so the team can see what people want to know. And the agent's answers become visible, so errors surface in front of the people who can fix them. Dave calls this "almost the sense of public accountability," and describes a maintenance loop: they watch the responses, and "if it's like not up to snuff, we will fix it. And we have in the past."
That last clause matters. Not "it works," but "it doesn't always, and here's who corrects it." The output is visible by default, and a specific team owns fixing it. That's a more practical definition of human-in-the-loop than the phrase usually carries.
What this means for UX research right now
The argument has shifted from whether to where. Nobody in six episodes debated whether AI belongs in research. They debated placement: which tasks, which phase, how much human involvement in which part. Dave Chen is candid that his team hasn't taken up AI moderation yet, and that reads as a placement decision rather than a verdict. If your team is still relitigating legitimacy, the more useful move is to stop arguing about the category and start arguing about specific tasks, where evidence can actually settle it.
Rigor is now an operations problem as well as a methods problem. It used to live almost entirely in craft: sampling, guide design, analytical discipline. Those still matter. But at volume, rigor also depends on infrastructure, including systematic evaluation, guardrails, and traceability from a finding back to a real quote in a real interview. This is the part we care most about at Strella. A hundred AI-moderated interviews in a week without evaluation isn't more research. It's faster unreliable research, and the speed makes the problems harder to catch.
Workflow design is the skill that's appreciating. Jordan's instruction, work out how you'd do this as a human because you have to explain it to the model, is the most transferable thing in the series. It's also the thing an AI tool can't do for you. Note what that implies about seniority: the researchers getting the most out of these systems aren't the most technical ones. They're the ones who can describe their own judgment precisely.
The failure mode is ambition, not caution. The teams making visible progress picked one painful thing and solved it. The teams stuck were designing systems. That runs against the instinct of anyone under pressure to show an AI strategy, because one automated workflow is a less impressive slide than a roadmap. It's also the version that ships.
None of these guests described AI replacing research judgment. They described it raising the return on having good judgment, and the cost of not having it.
Check out every episode in the series
Episode 1. Michelle Bejian Lotia, Ramp. Aired July 9.
Start with one painful task. Don't architect a system before you've solved something.
Episode 2. Daniel Gottlieb, Microsoft CoreAI. Aired July 16.
Custom AI tools are proliferating untracked. That's ops' opening, not ops' problem.
Episode 3. Nam Pham, DoorDash. Aired July 23.
Teach the system patiently. Most "new" AI concepts are familiar craft in new vocabulary.
Episode 4. Llewyn Paine, AI expert and consultant. Aired July 30.
Researchers need permission to push back. Popular practice isn't the same as good practice.
Episode 5. Jordan Brinkman, ERGO NEXT. Aired August 6.
Map how you'd do the work as a human first. Rigor requires someone who can judge the output.
Episode 6. Dave Chen, 1Password. Aired August 13.
Sort the process into automate, augment and keep human. Put the agent somewhere visible.
Coming soon
Episode 7. Amanda Amyx, Hatch. Airs August 20.
Scaling with AI moderation, and telling customers plainly that's what you're doing.
Episode 8. Allison Robins and Brooke Sykes, Mozilla. Airs August 27.
Leading through the messy middle. Three modes of human-in-the-loop.
Listen on Spotify, Apple Podcasts, or at The ResearchOps Review.
Frequently asked questions
What is the How to AI UXR podcast?
An eight-part video podcast series from Kate Towsey and The ResearchOps Review, sponsored by Strella, running weekly from July 9 to August 27, 2026. Each 30-minute episode features one research practitioner walking through a real case study of how they use AI in one phase of their research process.
Who hosts How to AI UXR?
Kate Towsey, founder of The ResearchOps Review and a long-standing voice in research operations. Priya Krishnan, co-founder and COO of Strella, contributes a short reflection to each episode.
What is the How to AI UXR map?
A published guide to building AI-augmented research operations, based on 562 data points and conversations with 50+ research and ResearchOps professionals. It organizes AI adoption into three maturity levels, crawl, walk and run, across every stage of the research workflow. The podcast series is an extension of it. Get the map.
What are the main takeaways from How to AI UXR so far?
Three themes recur across the first six episodes. Start with one painful task rather than designing a system upfront. Expect AI to move the hard thinking earlier, into workflow design and evaluation, rather than removing it. And treat the real constraint as operational, meaning governance, guardrails and traceability, rather than a limitation of the models.
Does the series cover AI-moderated interviews?
Yes. Amanda Amyx of Hatch covers scaling research with AI moderation and customer transparency in episode 7. Several other guests discuss where moderation sits relative to work they deliberately keep fully human.
Where can I listen?
On Spotify, Apple Podcasts, YouTube, and The ResearchOps Review.



