We cannot access our instance neither sandbox since late Monday EST.
Any updates besides the ones showing on the Atlassian status page?
Thanks,
Olimpia Estela
Hi Everyone,
In an Easter miracle of Biblical proportions, our site has been resurrected!
Here is what Atlassian support have sent us:
We have successfully completed site restoration! Please follow the steps below to get started.Step 1: Validate your sitePlease validate your site and respond to this ticket with any outstanding issues or concerns and we’ll make it our priority to work through them with you. Our engineering and support teams are standing by to address any additional issues.Step 2: Enable mailNote that your incoming and/or outgoing mail will be disabled at this point. Once you have thoroughly validated the site, you can enable it by going to the link below:Outgoing mail - https://<YourSite>.atlassian.net/secure/admin/OutgoingMailServers.jspaIncoming mail - https://<YourSite>.atlassian.net/secure/admin/GlobalMailSettings.jspaIf you would prefer us to turn this functionality back on for you, please reach out to us via this support ticket.Step 3: Sync groupsIf you have user provisioning configured at your Organization you might need to resolve potential conflicts with syncing groups to ensure the changes from your Identity Provider during the outage get reflected in your site. To achieve that, please follow the instructions below:Access https://admin.atlassian.com/Select your organizationNavigate to Directory > User provisioning. If you don't see User provisioning under the Directory tab, go to Settings and it should be there.You might see a banner asking you to resolve the syncing conflicts. Banner will show: <group count> groups pending sync(If you don't see a banner, no action is needed.)Click on ‘Review groups before sync’.Expand the groups to see what will change (who will be added/removed from that group). You can always add/remove users later at the Identity Provider administration dashboard.This process will link these groups to the ones at your site(s) and allow you to add/remove people based on what we receive from your Identity Provider.
We have successfully completed site restoration! Please follow the steps below to get started.
Step 1: Validate your site
Please validate your site and respond to this ticket with any outstanding issues or concerns and we’ll make it our priority to work through them with you. Our engineering and support teams are standing by to address any additional issues.
Step 2: Enable mail
Note that your incoming and/or outgoing mail will be disabled at this point. Once you have thoroughly validated the site, you can enable it by going to the link below:
Outgoing mail - https://<YourSite>.atlassian.net/secure/admin/OutgoingMailServers.jspaIncoming mail - https://<YourSite>.atlassian.net/secure/admin/GlobalMailSettings.jspa
If you would prefer us to turn this functionality back on for you, please reach out to us via this support ticket.
Step 3: Sync groups
If you have user provisioning configured at your Organization you might need to resolve potential conflicts with syncing groups to ensure the changes from your Identity Provider during the outage get reflected in your site. To achieve that, please follow the instructions below:
Access https://admin.atlassian.com/
Select your organization
Navigate to Directory > User provisioning. If you don't see User provisioning under the Directory tab, go to Settings and it should be there.
You might see a banner asking you to resolve the syncing conflicts. Banner will show: <group count> groups pending sync(If you don't see a banner, no action is needed.)
Click on ‘Review groups before sync’.
Expand the groups to see what will change (who will be added/removed from that group). You can always add/remove users later at the Identity Provider administration dashboard.This process will link these groups to the ones at your site(s) and allow you to add/remove people based on what we receive from your Identity Provider.
In case it is useful for other customers, here is the timeline of events regarding our restoration:
We are beginning to test the site and verify that everything has returned to its proper place. At first glance, it seems ok, but we are going through it in detail.
Thank you to everyone in this thread for sharing information about this event. It has been very helpful for our organization.
Best of luck to everyone, and thank you all once again.
@Olimpia Estela Cáceres-Brown -
The best way I would recommend is for you to submit a formal support request to Atlassian Support (https://support.atlassian.com/) for direct assistance.
Best, Joseph Chung Yin
Jira/JSM Functional Lead, Global Infrastructure Applications Team
Viasat Inc.
Thank you @Joseph Chung Yin for your response. I already submitted a request to Atlassian Support.
Further: the information presented in the status page shows no progress whatsoever. Check the screenshot below.
I am fairly certain progress is being done behind the scenes, but please inform us so we can in turn inform our customers about what's going on. Compared to other providers of professional services, the information level here is abysmal. As a former president would put it: Sad!
When trying to create a support ticket for the issue, the ticket service is out of commission. Likely because my instance is down. Oh, the irony!
The irony is that we use Jira to provide our customers with a higher quality of service, while we see that Atlassian, in particular, leaves us completely in the dark.
As a customer, I expect on a regular basis meaningful updates, including something like an predefined update interval, when we could expect the next update.Posting an update without useful information here is pointless and only arouses more irritation. So please tell us a bit more about what you are working on and when you expect to resolve this issue.
Have the same problem. Can't find our site, but then there is absolutely NO OTHER WAY TO CONTACT THEM!!! And as previously mentioned, their updates are completely useless. Just the same repetition of the last update with someone using a thesaurus on a couple of words, but not actually providing useful information on timeframes. And the link to "Reach out to us if you have any questions or comments" dumps you right back to the same useless web page that we can't get pass because they can't find our Atlassian URL. Hoping someone from Atlassian sees this and reaches out to us, since we have no way to "Reach out to them".
Engineers tend to (only) focus on solving the problem, which sadly means missing informing the customer about what's going on. This is a big mistake.
Example: imagine you're on a train from city A to B. Suddenly, the train stops in the middle of nowhere. No public information announcements whatsoever. Twenty minutes later, the train starts rolling again and you arrive your destination.
Now imagine the same stop scenario but as the train stops in the middle of nowhere, the train driver/engineer gets on the speaker and informs everyone she's got a red light ahead and needs to call dispatch. Throughout the unscheduled stop, the engineer keeps you informed every ten minutes about what's going on. Thirty minutes after the unscheduled stop, the train starts rolling again and you arrive your destination.
Which of the above two scenarios would have you fuming as you arrive? I would much rather prefer a 30-minute delay being fully informed, rather than a 20-minute delay and no information given.
ATLASSIAN, PLEASE LEARN FROM THE ABOVE.
I totally agree with you!
Hello, We're sorry you've been impacted by this incident. We have sent email communications to all affected customers which provide further details of the incident as well as a link for live status updates. Reach out to us at https://support.atlassian.com/contact if you have any questions or concerns.
Additionally, we’re checking and validating anyone who has posted in Community to report their site has been impacted and we are creating a support request on their behalf. If you’re unsure if an issue has been created for your site we suggest reaching out to your site admins or technical contacts.
Regards,Stephen Sifers | Product Lead, Community
I am unable to report an issue, because the url for our Jira Service Management is not recognized (due to the error you are experiencing).
So please help the ones impacted with solutions different than you are proposing!
Our URL is: https://hwbgsupport.atlassian.net
Agreed. Seeing this often repeated suggestion to contact support when the support page blocks further contact because of the very problem we are trying to get help for.... is intensely frustrating. We have an entire engineering team idled and our Service Desk team is back to handling calls via direct emails.
To rub salt into the wounds, my colleagues at other companies are happily using their Atlassian instances since this was evidently only impacting a "small segment" of customers.
Dear Stephen,
Thank you for your response.
I assume you did not read this thread. I will try to summarize what has been written for you.
In summary, technical mistakes and glitches do happen. As to how to handle them, check out the train story above.
We've also been affected and can't create a support ticket. Nearly 2 full workdays lost for 2 departments, paperwork is going to be fun! Good luck with the fixes.
Hello Mr. Sifers,
Thank you for replying in the forum; unfortunately your reply doesn't provide us with much information.
At my company, we're in the same position as @Karim Abrik and @Ulf Sahlin and many others. We can't even interact with support because the support form cannot find our site URL (presumably because of this incident).
Yes, you've sent out an email to the affected customers, but that email said that recovery would take several days and it strongly suggested that there was the potential for data loss. It very pointedly did not confirm that the data was safe. We're already at the point where the downtime is measured in days instead of hours.
You can imagine that this has caused a tremendous amount of stress for the affected customers. The updates on the status page are sporadic and devoid of anything that would give a hint as to an ETA.
The way that these updates are phrased as to be intentionally oblique also ratchets up the stress. Whenever people talk around an issue instead of directly about it, one starts to think that they aren't giving the unvarnished facts.
All of your customers work in IT. We can see that this is a gigantic screwup somewhere along the line in Atlassian. Ok, stuff happens to the best of us. Nobody is perfect. Even big screwups stay a minor irritation if they are resolved quickly and communicated clearly.
But, given the length of the outage and the way we're being communicated with, this is looking like something that is even more serious than it already appears.
Your customers are pretty unhappy right now.
We're unable to contact support, we're unable to access our instances, the communication from Atlassian has been terrible (although sending the email was a good idea) and we all have our own teams to answer to about this.
We would appreciate it if we could get at least a solid confirmation that no data has been lost and an estimate of when we will have access to our instances again.
We need to communicate that information to our stakeholders, and it is very reasonable for us to expect you to communicate that to us.
If you don't have that information, then please just say "We don't know if any data has been lost and we don't know when or if your site will return to normal."
Thank you again for replying. We do appreciate it.
Mr. Sifers / Atlassian Support Team:
We're still stuck here. The fact that you don't even have someone regularly monitoring your community forums in an incident this significant does not inspire confidence.
You wrote "If you’re unsure if an issue has been created for your site we suggest reaching out to your site admins or technical contacts."
I am the site admin for our company. I would imagine most of the people in this thread are also site admins.
We are already asking you for the status of this issue. I (and I suspect most or all of us) haven't received any communication from Atlassian saying that a support ticket has been created.
We all understand that this isn't something a first-tier support agent will help with. What we're desperately trying to get from Atlassian are answers of some kind, even if the answer is an honest "we don't know how bad the damage is yet".
What we're getting right now is disinterested silence, broken support request forms and boilerplate "we're working on it" updates to the status page.
At my company at least, this outage has already caught the attention of the finance team that pays the bills.
It has also agitated everyone from middle-management down who haven't been able to submit or update any requests to our IT or marketing departments for a very long time now.
In short, Atlassian's failure to communicate about this in any meaningful way has managed to turn pretty much everyone in our company against them in the span of a little over a day.
Your site admins and technical contacts are the people who advocate for Atlassian within our respective organizations, and you're hanging us out to dry.
Please reply with meaningful information that we can communicate to the many thousands of people across our collective organizations who are directly affected by this situation.
Thank you for your time and assistance.
I would recommend handing out a copy of the book Radical Candor to the Atlassian support management team.
I'm really looking forward to the broken SLA reimbursements.
Likely they will offer a 50% service credit for one (1) month. Not automatic however: one actually needs to APPLY for it.
Imagine how that resonates with the actual cost of all the non-operational organizations.
Service Credits | Atlassian
I have tried submitting a request on Atlassian support, but it will not allow me to do so because our site is down. This is ridiculous - our entire organization depends on this and it has been down for ~36 hours with no information about when it is expected to be restored.
And don't tell us you will post an update in an hour or in three hours and then not post an update!
So, now that we are going on Day 3 of not being able to access our JIRA Ticketing System, Confluence Pages, or the Insight instance that we've been working on, and being provided with very little to no information about how much longer this outage might be, I'm becoming more and more curious about the extent of the small population mentioned in "While conducting routine maintenance, an action caused a small population of our customers to be unable to access their products and data. Please reference our statuspage for updates." are effected.
Who (or how many) is on the list of this subset of the small population that has been without their ticket system for ~60+ hours based on "Update - We are continuing work in the verification stage on a subset of instances. Once reenabled, support will update accounts via opened incident tickets. Restoration of customer sites remains our first priority and we are coordinating with teams globally to ensure that work continues 24/7 until all instances are restored." Since this has been the only provided update at Apr 7, 12:27 UTC, Apr 7, 09:35 UTC, Apr 7, 04:58 UTC. Is my company part of that subset, or are we farther down the list??
What is the severity of the outage. Is there data loss? Was something compromised? I've seen messages indicating the routine maintenance issue as well as issues with Atlassian AWS cloud. What is going on so we can prep our customers.
When will we be able to contact support through a means that actually functions? How do we get updates and timeframes that we can share with our customers to assure them that JIRA IS a legitimate ticketing system solution that is the right solution for them (my company was SUPPOSED to have a live demo of that system yesterday to a customer, but instead had to show them screenshots of the system that we use to support customers 24/7, because it was down and had no ETA on a return to service).
And what (if anything) will Atlassian learn from this outage to ensure that they have methods in place to provide their customers with usable, factual, and adequate updates to inform and manage the expectations of their customers' customers.
---
We have been down since Tuesday morning. It's beginning to seriously affect our teams. My IT team is coping but we can't access our asset tracking now and are having to handle support requests via email (which is super not ideal). I'm also curious to know where in the queue we stand for restoration of services. I have seen some big name companies complaining and fear that we might be considered "too small" to be addressed in a timely fashion.
Our engineering group is already planning contingencies as this is starting to impact release schedules and seriously hampering their work.
We have lost a LOT of good will towards Atlassian and I would not be surprised if leadership begins to ask for competitors and/or lawyers get involved.
Just give us more information so we (those who represent you at our companies) can give useful information to our leadership and users. Saying "we don't know when it will come back...if it will come back" is horrible.
It is incredibly discouraging that, after this many days and this many posts requesting some kind of real information, that nobody from Atlassian has responded after Mr. Sifers non-answer above.
Today is April 7th. Our sites have been inaccessible since April 5th.
This is the third day of downtime on the products that we use in order to *minimize* our downtime.
The level of disinterest Atlassian is displaying towards its paying customers is hard to comprehend.
You accidentally deleted our sites, might not be able to restore our data, and can't even be bothered to pop into the forums twice a day to say "We're still here, we understand your frustration and we haven't forgotten about you."
We're left to shout into the darkness and commiserate with our fellow customers who have had the rug pulled out from under us.
When problems like this occur, reputational damage control is just as important as technical damage control. If, as the response below indicates, you have hundreds of engineers working on this problem, how do you not have a support team proactively trying to communicate with the affected customers, in the forums or directly?
How do we continue to advocate for using Atlassian products? What should our response be to the executive teams and clients who depend on our recommendations for software solutions?
At least BitBucket isn't down, so we can continue with some development, but how do we answer when our organizations ask "How long until this happens to BitBucket, too?"
Somehow, Atlassian has a market cap north of 70 billion dollars but can't spare anyone to proactively communicate with its customers, even in a worst-case scenario outage such as this.
I have no doubt that I speak for everyone in this thread when I say that we are more than incredibly frustrated here; we are angry.
Three days, and we don't have real information about what happened.
We know that the original statement was not honest or complete because, if our sites were just disabled, they would have simply been re-enabled and this would have been a blip, not a three-day outage.
We know that Atlassian's own statements have gone from "We don't believe any data has been lost at this point." to "We are working on minimising any potential data loss."
We know that we've had to spend the last three days answering questions from our stakeholders asking us "What happened to the service desk requests?" and "Where are our Confluence documents that we've spent three years accumulating?"
Nobody has ever heard of a publicly-traded, industry-standard company in Atlassian's market position taking its core products offline for three days due to human error. It is quite literally unbelievable for IT service managers to hear this.
How do we rely on Atlassian for incident management software when this is how you handle incidents?
Why are we paying for a managed solution when you don't have automated disaster recovery procedures?
Someone has created a support ticket on our behalf now, which is great. For anyone who has not received such a ticket, I will include it below. It is not exceptionally helpful.
Hi Team,
We’re sorry for the continued frustration this incident is causing. We are continuing to move through the various stages for restoration. The team is currently in the verification stage on a subset of instances. Successful verification will then allow us to move to reenabling those sites or identify any other steps needed for restoration.
Once reenabled, support will update accounts via opened incident tickets. Our efforts will continue 24x7 through this process until all instances are restored.
This is our top priority and we have mobilised hundreds of engineers across the organisation to work around the clock to rectify the incident. These restoration efforts may be visible on your sites today, however please wait until support notifies you when the site is fully available and works with your teams to confirm the recovery.
The restoration is expected to continue over a number of days. We are working to reinstate access to all products with the priority on our key products. We are working on minimising any potential data loss.
We can confirm this incident was not the result of a cyberattack and there has been no unauthorized access to your data.
As we work to restore access you can look to us to continue to provide updates:Every 3 hours, or sooner if we have a material update, at: http://status.atlassian.com
Direct contact via support tickets once we’re able to reinstate your access and your site becomes usable
You can continue to reach out to us at https://support.atlassian.com/contact for any questions, concerns or updates. If you have any issues opening a technical support ticket please open a billing question ticket and we will transfer it into the support teams.
Regards,Daniel Soo
I received the EXACT same response (verbatim) to my ticket, which was auto created as a result of my comment above. All I can do is SMH and roll my eyes. You hit the nail on the head, Shane!
New day, new hope.
I would assume Atlassian would be able to get that "handful" of sites up very quickly, as "a handful" to me means <100. Now it seems to me the number of affected customers is way higher than that.
I received this nothingburger in the mail this morning. It is now FRIDAY and still no reasonable word on ETA or potential data loss. Out of those "hundreds of engineers" working on it, maybe Atlassian could set aside a few to deal with giving information to customers?
Scott Farquhar here, I want to personally apologise for the Atlassian outage that you are experiencing. We understand how mission‑critical our products are to your business, and want to make sure you know we are doing everything we can to resolve this. We hold ourselves to the highest standards in dependability, transparency and customer service, and over the past few days, we have failed to live up to that standard.
On Tuesday morning (April 5th PDT), we conducted a maintenance procedure designed to clean up old data from legacy capabilities. As a result, some sites were unintentionally deactivated, which removed access to our products for you and a small subset of our customers. We can confirm this incident was not the result of a cyberattack and there has been no unauthorised access to your data.
We are working 24/7 to restore your service and will alert you when your products are available. We have already restored partial access for some customers and will continue to restore access into next week. Please know that once we have recovered all of our customers' access, we will review our processes to conduct a complete post incident review. We will make an overview of this post incident review available to you.
In our efforts to restore your site as quickly as possible, there may be some limitations when we make it available to you such as 3rd party app functionality. We will be sure to inform you of these in our direct communications with you.
When your site is available, we will directly notify you via your support ticket along with any details on the limitations mentioned above, as well as guidance for follow‑up support.
We'll continue to provide updates on status.atlassian.com as new information becomes available. If you have further questions, please reach out to us at https://support.atlassian.com/contact. If you have any issues opening a technical support ticket, please open a billing question ticket and we will transfer it into our support teams. It is my and my team's priority to do what we can to make things right.
Were any of you guys using the old Insight and were forced to move to the new/built-in one? The end-of-life of the old Insight was March 31, 2022.
We moved from the old to the new Insight and we are affected by this incident.
I really am frustrated and I believe within ITSM, this is what you don't want for your customers!
We are (were?) insight users but I think Atlassian have had the last of any money from me. Not for the outage, more for the completely amateur way it's been handled. I can no longer trust them with any of our processes.
How can a company write ~20 emails that say nothing?
How can you need "hundreds of engineers" to fix "a handful of sites"?
This whole thing stinks.
Yesterday evening I had a direct conversation with Mr. Sifers, who had responded above, and with Mr. Spells in escalations at Atlassian.
They were quite open about the causes of the current problem, and what Atlassian is doing internally to rectify the situation.
I was asked not to paraphrase what they said (it honestly wasn't much more than the email sent by Mr. Farquhar; it was just a bit less corporate-speak), but I strongly recommended that they participate in the community discussions around this. I can definitely confirm that Mr. Sifers is in this forum and has read everything that has been said. I have no explanation as to why he doesn't actively participate, and, believe me, I asked.
I can also say that both Mr. Sifers and Mr. Spells were very informative, professional and courteous with me. They both seemed to genuinely care about the situation. Mr. Spells in particular went above and beyond to answer my questions and I respect the way he handled the interaction. He was exceptionally accommodating and pleasant.
I don't have any more information about return-to-service times than anyone else, and all the questions I articulated above still stand, but at least there was an attempt to communicate with an affected customer.
Further, as much as the support team is steadfastly refusing to communicate in any kind of normal way publicly about this, my impression is that the they legitimately do not have answers much more helpful than they've already given. I just don't think that anyone there has concrete information yet about when this will all be over.
Again, I strongly disagree with the corporate decision to let people languish in the forums without official responses -- in a way, it's almost worse to know that the powers-that-be can see your frustration and are choosing to do nothing.
For now, I will say that I appreciate both Mr. Sifers' and Mr. Spells' time and explanations very much, and I will reiterate how professional they were with me. I could see they were both feeling the weight of a very heavy week.
I'm still frustrated, I'm still angry that all of our projects are at a standstill because our project management platform ate our site, I'm unhappy that they've chosen to stay silent in the forums...there's no satisfaction to be found yet.
I'm writing this post so that maybe knowing Atlassian is actually reading this thread will give a tiny spark of solace to my fellow customers who have been burned by this the same as me.
There will be hard questions for Atlassian to answer at the end of this, and I am not sure that my company will even continue using their services -- once bitten and all that.
We pay for a managed solution from one of the world's largest, most trusted providers of IT service management software, and they have let us down catastrophically. It's hard for me to go to stakeholders and pair that information with the sentence "but I'm sure it won't happen again."
This situation has made everyone in my IT department question the disaster recovery procedures in place at Atlassian, across all their products. It's a very large and complex platform, and the fact that they don't have automated, tested bare-metal recovery processes in place is just inexcusable.
How can we be sure that they have DRP in place at BitBucket? If our pipelines there go down, the situation would be dinosaur-asteroid bad for us.
I don't know if we've lost all our data yet or not, but I do know we've lost a week of work across the company. There is a cost to that which exceeds our subscription fee by a large margin.
If we aren't operational on Monday, then we quite literally have nothing left to lose by switching away from Atlassian.
Even if we are operational on Monday, we can no longer trust Atlassian the way we used to.
That means we need to prepare our own DRP and create some kind of on-premise fallback in case Atlassian goes down again.
And if we need to spend time and money doing that, what do we gain from using Atlassian?
We store data in their formats, in their systems, in exchange for the resiliency and professional capabilities of a 72 billion dollar company specializing in IT service management.
The problem isn't just that there was a technical failure at Atlassian; it's that they weren't prepared for it.
Thank you again to Mr. Sifers and Mr. Spells for reaching out to me. I hope that they are continuing to do that with as many affected customers as possible.
Well said, on all counts
Thank you, @Karla Keefe. I hope that your instance is restored soon.
(Saying that feels like some strange Viking blessing..."May your instance always be online and your data always be backed-up.")
And 2 points gryffindor for @Ulf Sahlin putting together that the insight retirement was to blame for all of our outages 4 DAYS AGO. As far as I’ve seen, that was something that was only actually confirmed by atlassian today. The amount of deductive reasoning that I’ve seen by this forums community members has been nothing short of impressive. Thank you.
It looks like you're new here. Sign in or register to get started.