I've been digging through the public Marketplace API to understand what people actually complain about in paid apps, and the result surprised me enough that I want a sanity check from people who run these instances daily.
What I pulled: every 1- and 2-star review with text from 22 of the largest paid apps — Xray, Tempo Timesheets, ScriptRunner, Zephyr, draw.io, Gliffy, BigPicture and others. I ranked "largest" by installs multiplied by list price at the 50-user tier, which is a proxy and not actual revenue. 431 reviews in total.
What's in them: roughly half mention a technical failure — broken, won't load, slow, data lost, capability missing. I'll be honest about the fragility here: depending on which words I count as "technical", I get 48.7% with a narrow keyword list and 52.4% with a wider one, and about a third (30-33%) are purely technical with no support or pricing complaint attached. 32.9% mention support or pricing. If I restrict to 2025-2026 reviews only, the technical share goes up rather than down, to 61.2% — though that's n=67, so treat it as directional.
The part I keep coming back to: I read those 67 recent ones individually rather than trusting my keyword matching, and the single biggest cluster isn't missing features. It's the app simply not working, for days or weeks, while the admin hears about it from their own users. Some of them, quoted as written:
"Plugin is not loading from last two weeks and there is no resolution yet."
"We've been facing a critical issue for nearly two weeks now where our users are completely unable to use the app."
"Outage for over 22hrs, no communication, even the status page is sending mixed signals."
"Error rendering web panel" — quoted verbatim by two reviewers three weeks apart, one writing in Portuguese, and a third describing the same testing board failing to render, in Spanish.
So I tried to check that third one. I took the outage complaints that carried dates and looked at whether the vendor's own status page had declared an incident around that date. Within the window each status feed actually covers, 6 of 7 had no declared incident within plus or minus 7 days.
That n is 7. It proves nothing, which is exactly why I'm asking here instead of announcing a finding. There's an innocent explanation that fits just as well: those may have been single-tenant problems, which a global status page correctly reports as "all systems operational" — and the admin still has no way to know.
What I'd genuinely like to hear from you:
1. Has a Marketplace app gone down on you before the vendor acknowledged it? How did you find out — users, monitoring, or the status page?
2. Would a public, independent record of "this app was actually unreachable on these dates" be useful to you? For choosing apps, for renewal conversations, or for your own reporting?
3. If yes, which apps would you want watched?
I'm recording it either way, comparing what a vendor's status page says against a direct unauthenticated probe of their API — no credentials, nothing installed on anyone's site. Happy to share the raw review data and the collection method with anyone who wants to check my numbers.
And if the honest answer is "this isn't a real problem, you're reading noise" — that's a useful result too, and I'd rather hear it now than in six months.
Nick — thank you for the RFC pointers specifically. I read RFC-124 and RFC-133 afterwards, and you are right that A4A requires public SLOs, a public status page and SOC 2 or ISO 27001. That closes the gap I thought I had found, at least for badged apps, and I would not have known to look without your comment.
The SOC 2 route interests me more than the pinger ever did, because it is an audited answer rather than an inferred one. One practical question, if you are willing: does requesting the report actually work outside of large enterprises? You framed it as "if you are an admin in a big company". For a fifty-seat customer, does a vendor typically send the Type 2 report under NDA, or is that a conversation only procurement teams get to have? That difference decides whether your alternative is universal or only available to people with leverage.
Your suggestion about scoping to unbadged apps with a vendor dispute path is the only shape I could see working, and I think it needs someone with standing in this ecosystem to run it fairly. That is not me.
From my conversations with app developers (I'm not a developer myself), the biggest challenge for them is the constant change that's happening in the Cloud environment. It is way harder to keep an app stable on Cloud than on Data Center. There are regular API Endpoint deprecations, currently there's the move to Forge, core features are being renamed, others deprecated etc.
Even with a really good developer team it is likely that things break or you simply cannot keep up with all changes required.
Luckily, a longer outage of a Marketplace app we are using has not occurred yet.
As a result of what I wrote, I think a public record like #2 you suggested would have a character of "name and shame" which is not fair under those circumstances.
Sebastian — one line in your comment turned out to matter more than anything I brought: "a longer outage of a Marketplace app we are using has not occurred yet". That is a data point against the problem I went looking for, from someone actually running the apps, and it should count as much as any complaint I pulled from the API.
Your framing also changed my mind about fairness rather than just about method. If instability is largely driven by Cloud moving underneath vendors — endpoint deprecations, the Forge migration, features renamed — then a public record measures the platform's rate of change as much as any vendor's diligence, and attributing it to the vendor is simply wrong.
It has become worse than that since I wrote the post. I tested my own claim that technical complaints are rising, and it fails: 2019-2024 sits at 50.8%, 2025-2026 at 61.2%, permutation p of 0.14, and the yearly series bounces with 2020 as high as 2025. So I do not even have evidence of deterioration to name anyone for. Thank you for pushing on the fairness question before I had thought it through.
You don't have to got to marketplace apps for that phenomena, because Statuspages are a better tool for reporting the past than hinting to live problems. AFAIK even Atlassians Statuspage is not reporting live about incidents. I would go so far to say, that whenever we were affected by a large scale incident, that was mirrored on their page, we always felt the impact some hours before the statuspage displayed it.
The reason for that is pretty simple: The display on the Statuspage is not automated and there are reasons for that. As Statuspages normally are subscribed by a lot of customers, the one thing you never want to happen is that it shows an outage, that never happened. So there will always be a human deciding to change the product status on the page. This only happens after
a) a monitoring system makes the assumption that something is down or behaving badly
b) some customers open a ticket to make you aware of a problem
So a Statuspage is a tool that should reduce your ticket intake in bigger outages and makes reaching all you customers easier in these cases, not a real-time monitoring.
Regarding the marketplace apps I can confirm, that some of the apps we use had more difficulties in the near past than in the last years. But in our case it were rather small problems that outages of any kind. We also have no app currently that has bad support or is giving us pricing headaches. I think part of this development is also to blame on a general social change that's influencing how people react to difficulties.
I think some of these reviews would have been phone calls or e-mails in the past. In these times people value the option more that gives them the biggest echo chamber for their emotions rather then finding the correct entry to a process that solves their problem.
Richard — two separate things in your comment, and the second one is the part I keep coming back to.
The first is mechanical: a Statuspage is human-gated on purpose, because the failure you must never have is showing an outage that did not happen, and its job is to reduce ticket intake during a real one rather than to report in real time. I had been treating it as a monitoring feed that was underperforming. It is not that thing at all, so a divergence between it and a probe is not a finding.
The second is the one I did not expect: "some of these reviews would have been phone calls or e-mails in the past. In these times people value the option more that gives them the biggest echo chamber". That is a better explanation of my data than my own hypothesis was. A rising share of technical complaints in reviews is equally consistent with a shift in where people take their frustration, and I have no way to separate the two.
I then tested the trend itself and it does not survive: 2019-2024 at 50.8%, 2025-2026 at 61.2%, permutation p of 0.14, with 2020 as high as 2025. So the pattern I thought I was explaining may not be there at all — and if it is, your channel-shift reading fits it as well as any reliability story. I have retracted the claim in the thread.
I provide technical services supporting Atlassian customers of all sizes (from small businesses to large enterprises). It seems to me that you're reading noise:
"Top paid apps" are usually (always?) published by metal-tiered vendors (Silver, Gold, Platinum). Many of them are SOC 2 certified. All of them have support requirements imposed by Atlassian as part of their tiering. I just don't believe outlandish reviews claiming extended outages.
Regarding this:
I'm recording it either way, comparing what a vendor's status page says against a direct unauthenticated probe of their API.
I'm really curious what that means. Which API(s) are you probing? How is the API an indication of overall app availability to users? I'm especially curious about apps that have migrated (partially or fully) to the Atlassian Forge platform (which is a lot).
Aaron — you raised a testable hypothesis, so I tested it rather than argue. You suggested the two reviews might be competitive sabotage. I pulled 500 reviews for Zephyr Essential (com.thed.zephyr.je) with author names and dates.
Result: the two are from different people, Kshitij Semwal on 2025-06-30 and Pushpkant Garg on 2025-07-09, each with exactly one review in that sample. Single-review accounts are the norm here, so that alone distinguishes nothing. But the timing does: negative reviews ran 3 in May, 3 in June, 2 in July 2025, against a median of 2 per month across the app's whole history. There is no burst. A coordinated campaign would normally show one.
So the sabotage reading is not supported by the data. Your other explanation is untouched, and I think it is the stronger one: a localized error in those two sites, which the vendor could not reproduce and a global status page would correctly report as operational.
On your API question, and the part of it I had not properly reckoned with: you asked specifically about apps that have migrated to Forge, which is a lot of them. Forge apps do not expose a public endpoint of their own — I checked, and Atlassian confirms there is no supported way even for an installed app to enumerate other apps on a site. So for exactly the population you are asking about, my probe has nothing to probe. That is a harder limit than the noise problem, and it holds regardless of how carefully the checks are designed.
To your framing overall: I have since tested my own claim that technical complaints are rising and it fails, permutation p of 0.14 with 2020 as high as 2025. "Reading noise" was a fair description.
Thank you, all four of you. This is exactly the check I was hoping for, and it went against me on the main point.
Nick, on the pinger — you are right, and I think it is decisive. An unauthenticated probe from outside sees rate limits, edge behaviour and one tenant's trouble, and cannot separate any of that from a real outage. My "6 of 7 with no declared incident" is at least as well explained by your reading as by mine. I had written "global status is not tenant health" in my own notes and then failed to draw the obvious conclusion from it.
Richard, on status pages — this dismantles the framing rather than the method. I was treating a divergence between the status page and a probe as a signal. If the page is deliberately human-gated, precisely so it never shows an outage that did not happen, then divergence is the expected behaviour of a process working correctly, not evidence of anything. My load-bearing assumption was that the page was trying to be real-time. It is not, and it should not be.
Aaron, you asked what I was probing. Unauthenticated GETs against api.tempo.io/4/worklogs, api.zephyrscale.smartbear.com/v2/testcases and prod-play.zephyr4jiracloud.com/connect/public/rest/api/1.0/serverinfo, reading a 401 in about 200ms as alive and a 5xx or silence as not. Your implied objection lands: a 401 proves an edge gateway answered, not that the app works.
On the two reviews you doubt, the precise shape, because it is fairer to put it on the record than to defend it: both are Zephyr Essential, dated 2025-06-30 and 2025-07-09, nine days apart, from two different reviewers. So not a single outlier. But two customers of one app is entirely consistent with your localized-error explanation, and I have nothing that distinguishes the two readings. I will drop the "offline for weeks" framing.
What survived: Nick and Richard independently said the same thing — more technical trouble in 2025-2026, tied to the Connect to Forge move and API deprecations. That is what the data shows too; the technical share of complaints rises in recent reviews rather than falls. I will take that as the real finding and let the rest go.
Sebastian, your "name and shame" point is why I would not publish a league table. A tracker that cannot tell a rate limit from an outage would produce exactly that.
So I am not building the uptime tracker. Four people with more context than me pointed at the same structural flaw within a day, which is a far cheaper answer than six months of building. Nick, scoping it to unbadged small apps with a vendor dispute path is the only version I can see working, and I do not think I am the right person to run it.
The 431 reviews with dates, ratings and app names are yours if useful, and I am happy to post the extraction script. If anyone wants a cut by app or by year, say which and I will run it.
I think there is a fundamental issue with using negative reviews as a proxy for the overall quality or reliability of an app.
People are much more likely to leave a review when something goes wrong than when everything works as expected. A satisfied user normally just continues using the product. A frustrated user, on the other hand, has a much stronger motivation to post a one- or two-star review. That creates an obvious selection bias.
This becomes even more noticeable with popular apps. If an app has tens or hundreds of thousands of users, even a very small percentage of unhappy customers can generate a significant number of negative reviews. That does not necessarily mean the product is getting worse. It may simply mean that the user base is larger, the number of use cases is broader, and there are more opportunities for users to encounter edge cases, integration problems, configuration issues, or platform limitations.
For that reason, I would be very cautious about concluding that app reliability has deteriorated simply because the percentage of technical complaints in negative reviews has increased. There may be a real signal there, but the review data alone does not seem strong enough to prove it.
Reviews are useful, of course. If many unrelated users repeatedly report the same issue, that is certainly worth investigating. But user satisfaction should be treated as one indicator among many, not as an objective measurement of service quality. Ideally, it should be compared with other data such as incident frequency, response times, support performance, bug resolution times, SLA/SLO compliance, release stability, and actual availability measurements.
The uptime question is more interesting to me.
If there is a concern that vendor-operated status pages do not always reflect what customers are experiencing, then I think the best answer is not to infer outages from reviews. It is to measure availability independently.
An external monitoring service could provide a genuinely useful outside-in perspective. But it would need to be designed carefully. Simply checking whether an endpoint returns 200, 401, or 500 would probably create more confusion than useful information. A meaningful monitor should ideally reproduce actual user workflows, check multiple components, distinguish between complete outages and partial degradation, run from different locations, and account for things like authentication, rate limiting, tenant-specific problems, and dependencies on Atlassian itself.
That kind of independent monitoring would also make the discussion much more objective.
Instead of debating whether a vendor's status page is accurate or whether a negative review represents a real outage, you could compare independently collected availability data with the provider's own incident history. If both sources agree, great. If they repeatedly diverge, then there is something concrete to investigate.
I also think this would be fairer to vendors. A public "uptime score" based on weak signals could easily produce false positives and damage the reputation of an app that was actually operating normally. But a transparent monitoring methodology, with clearly defined checks and reproducible data, could be genuinely valuable for both customers and vendors.
So for me, the most interesting outcome of this discussion is not whether 50% or 60% of negative reviews mention technical issues. That is interesting context, but it is still subjective data from a self-selecting group of unhappy users.
If we really want to answer the question "How reliable are the major Jira apps?", I think we need independent, measurable, reproducible operational data. Reviews can help explain the story, but they probably should not be the measurement itself.
Have a nice day,
--Hugo
Hugo — thank you, and I have to concede the main point rather than argue it.
You said the review data alone is not strong enough to prove reliability has deteriorated. I went and tested my own claim instead of defending it, and you are right. The year-by-year technical share of negative reviews in my sample: 2019 55.4%, 2020 62.9%, 2021 33.3%, 2022 59.0%, 2023 51.7%, 2024 48.3%, 2025 65.4%, 2026 46.7% (partial year, n=15).
It bounces, and 2020 is as high as 2025. Comparing 2019-2024 (n=364, 50.8%) with 2025-2026 (n=67, 61.2%) gives +10.4 points at a permutation p of 0.14. Not significant. So I retract "the technical share rises in recent reviews". I put that in the original post and again in my reply above, and it does not hold.
Worse: I wrote that Nick and Richard had independently confirmed the increase. They confirmed their own impression, which is worth having, but I treated two anecdotes as validation of a difference that fails its own significance test. That is exactly the error I had been criticising.
On selection bias generally, you are right that I cannot separate "more complaints" from "a larger user base meeting more edge cases". Marketplace installs are a point-in-time snapshot, so I have no way to normalise by installed base over time. I should have said so up front.
On your monitoring spec — real workflows, full outage versus partial degradation, multiple locations, auth, rate limiting, tenant-specific failures, Atlassian's own dependencies — I agree that is what it would take, and I think that is also the honest reason not to attempt it as a side project. Your spec and Nick's objection are the same observation from two directions: the cheap version produces noise, and the version that does not is a real synthetic monitoring platform.
So the accurate summary for anyone who finds this thread later: the review data shows what people choose to complain about, it does not measure reliability, and nothing here established a trend over time. The 431 reviews with dates and app names remain available to anyone who wants to check that themselves, including checking me.
Recommended Learning For You
Level up your skills with Atlassian learning
Learning Path
Improve user experience across Jira with global settings
Learn how to set up and configure a Jira site, manage Jira permissions, and configure Jira apps and integrations.
Learning Path
Streamline projects across Jira with shared configurations
Build Jira work items with reusable configurations called schemes, and reduce administrative work with automation.
Learning Path
Become an effective Jira software project admin
Set up software projects and configure tools and agile boards to meet your team's needs.