Forums

Articles
Create
cancel
Showing results for 
Search instead for 
Did you mean: 

How are teams comparing the quality of different AI coding agents in Jira?

Chris May
August 24, 2026

Teams can now use several different AI coding agents alongside Jira, but I’m interested in what happens after the agent completes its work.

How are teams determining whether one agent performs better than another?

For example:

  • Was the output accepted without changes?

  • Did it require minor editing or substantial rework?

  • How frequently did the agent fail to complete the task?

  • Are results compared by project or type of work?

  • Or is agent selection still mainly based on individual developer preference?

I have built a small Forge prototype under Aigeeza to explore this question. It allows someone reviewing a Jira task to record the agent used and classify the result as accepted, minor edits, major rework or failed. The results are then compared across projects and time periods.

agentscore-issuepanel.jpg

 

agentscoredash.jpg

Before developing it further, I would appreciate some honest practitioner opinions:

  1. Does your team regularly use more than one AI coding agent?

  2. Would this kind of comparison help you choose between them?

  3. Is a four-level human assessment too simplistic?

  4. What evidence would you need before trusting the resulting scores?

I’m not sharing an installation link at this stage. I’m primarily trying to establish whether this is a genuine problem for other Jira teams or merely something that interests me.

For transparency, AgentScore is a prototype I have built as part of my business projects on Aigeeza.

0 answers

Suggest an answer

Log in or Sign up to answer
DEPLOYMENT TYPE
CLOUD
PERMISSIONS LEVEL
Product Admin Site Admin
TAGS
AUG Leaders

Atlassian Community Events