Hi everyone! 👋
I'm a Dev/Test Engineer on a Kanban team working on a large-scale insurance platform. Here's what I'm seeing AI do for our Agile ceremonies — from the quality trenches.
We run Kanban — no sprints, no velocity charts, no two-week commitments. Instead, we plan on a quarterly cadence, setting priorities, capacity expectations, and key deliverables for the next 12 weeks.
Throughput forecasting — AI analyzes our historical cycle time and throughput to generate probabilistic forecasts ("85% chance you'll deliver 45–55 items this quarter") instead of gut-feel guesses.
Risk flagging — AI scans upcoming epics and warns us: "This area historically produces 40% more hotfixes." As a tester, that's gold for planning test effort upfront.
Bottleneck prediction — AI spots where work will pile up before it happens, helping us set smarter WIP limits.
The tester's pain point AI solves: Test complexity is finally visible in planning — not just dev effort.
Kanban retros can feel vague without sprint boundaries. AI helps by:
Correlating retro themes with flow data — "When the team flagged 'unclear requirements,' cycle time jumped 40% and escaped defects doubled." That's the evidence testers need to push for better refinement.
Surfacing hidden blockers — "Items spend 4.2 days stuck in 'Waiting for Test Environment' — up 60% from last quarter." Turns vague complaints into actionable problems.
Generating targeted prompts — Instead of "What went well?", AI asks: "7 items reopened last month — what's driving the rework?"
Are you using AI in your planning or retros? What's working?
Kanban teams — has AI changed how you manage flow?
What's one thing you'd never want AI to touch?
Let's hear it! 👇
Hi @zoltanersek _outpostlabs_dev_
Thanks for sharing your thoughts! Using AI to group related retro items sounds like a fantastic idea. It could streamline the process and help us focus on key themes and insights.
If you decide to try it out, I'd love to hear about your experience and any results you observe! 🙂
This maps closely to how I've come to think about it: a retro isn't just a session, it's a signal; and an estimate is really calibration data.
But most teams never get to actually query any of it. Your flow-to-theme correlation is the sharpest example of the payoff ("unclear requirements" -> cycle time +40%, escaped defects doubled); that's the jump from opinion-swapping, to real data-driven evidence you can take into refinement. Building on the grouping idea above, I'd push it one step further: clustering retro items within a board is handy, but the real signal shows up when you cluster across teams and quarters. That's where "one rough week" separates from "systemic issue".
Full disclosure, I work on retro and estimation tooling in the Atlassian Marketplace (Catapult Labs), and this is the direction we're building with Rovo agents: as an intelligence layer for the enterprise, so you can discover "what are the systemic issues across all our teams," correlating team-health signals with delivery patterns.
It's still early and we're auditing what the underlying data can actually support, so I'm sharing it as a direction rather than a finished thing.
On your last question: the line I wouldn't hand to an agent is facilitating the conversation itself or deciding what the team commits to. AI is great at surfacing the "what" and the "where it's stuck"; the "so what do we do about it" still needs the room. Has adding AI changed how much your team trusts the retro output, or is that still firmly the human part for you? @Anwesha Pan
Hi @Luis Ortiz - Catapult Labs
Love this and I totally agree with you.
Your cross-team clustering point is spot on too. One squad saying "flaky environments" is a gripe; three squads saying it over two quarters is an infrastructure investment case. That's where the real value lives.
On the trust question — from a QA perspective, AI has actually boosted trust in our retros because it kills recency bias. Instead of fixating on whatever broke last week, we get defect trends and cycle time shifts alongside the subjective stuff. Harder to dismiss, easier to act on.
Fully agree the "what do we do about it" stays human though. 🙂
Quick question: does your tooling pull in test-health signals (flaky rates, coverage gaps, environment stability) as part of the delivery correlation, or is it primarily planning/estimation data for now?
"Kills recency bias". That's one of our goals, moving away from the "who yelled the loudest yesterday" dynamic and toward actual historical evidence.
To answer your question directly: right now, our tooling is primarily focused on planning, estimation, and process data. We don't natively ingest raw CI/CD telemetry (like specific code coverage percentages or automated pipeline flaky rates) on our own.
However, there is a bridge: because our apps are built natively on top of Jira, our correlation relies entirely on the system of record.
If your QA team translates those test-health signals into Jira data, for example, logging escaped defects as specific issue types, tracking environment instability via labels, or using a testing app like Zephyr/Xray that maps directly to Jira tickets, using agents is an effective way (without over engineering our tooling) to cluster and surface those trends during the retrospective.
I think that if a signal matters, it needs to be visible in the Jira workflow.
Recommended Learning For You
Level up your skills with Atlassian learning
Learning Path
Apply agile practices
Transform how you manage your work with agile practices, including kanban and scrum frameworks.
Learning Path
Configure agile boards for Jira projects
Plan, prioritize, and estimate upcoming work by creating and configuring agile Jira boards for company-managed projects.
Learning Path
Registered Scrum Basics™
Manage work more effectively by learning scrum basics from a global leader in agile transformation and training—and get credentialed by Scrum Inc.®