Hi community,
I'm working in two different migration projects, and in both cases, CCMA is extremely slow and sometimes unresponsive.
Both instances of Confluence are very fast and work properly, the problem seems to be specific with CCMA.
In the first case, the client has a Microsoft server environment (Confluence hosted in MS Server, and database MSSQL). In that case, the bottleneck apparently is the communication with the database server that causes a spike in the server CPU.
Words from support:
we could find on the catalina logs (Tomcat) a lot of stuck threads where the database was not able to handle the SQL queries responses putting them in a queue. That's why the CPU did a spike of 95% because it is receiving many more requests from the CCMA than it can handle.
In the second case, the client has a Linux based server environment hosted in AWS. Our test server is an c5.2xlarge (CPU 8, Memory 16GB) with the database installed in the same server.
In this case, when the problem happens, the server CPU goes to 500% or 800%.
We can see this in the logs
17-Aug-2022 00:25:36.452 WARNING [Catalina-utility-4] org.apache.catalina.valves.StuckThreadDetectionValve.notifyStuckThreadDetected Thread [http-nio-8090-exec-23 url: /rest/migration/latest/stats/usersGroups; user: gmuller] (id=[799]) has been active for [67,496] milliseconds (since [8/17/22 12:24 AM]) to serve the same request for [http://34.211.xx.xxx:8090/rest/migration/latest/stats/usersGroups] and may be stuck (configured threshold for this StuckThreadDetectionValve is [60] seconds). There is/are [18] thread(s) in total that are monitored by this Valve and may be stuck.
java.lang.Throwable
at org.hibernate.event.internal.AbstractVisitor.processValue(AbstractVisitor.java:106)
REPRODUCING THE PROBLEM
After the check for errors phase, we proceed to the "Review your migration" phase, and this is where the problem happens. The estimated time keeps spinning forever, if I leave and let CCMA working, it will show a connection error eventually.
The rest of the instance gets unresponsive while CCMA is in that phase. For example, if I open another tab and try to navigate to any Confluence page, it will not work and will forever spin. If I close that CCMA tab, Confluence will get back to life at the same moment.
I did test with a very minimal number of spaces, like 5 spaces only, and it took like 10m to finish the Estimated time. For a customer that has 500, 800 spaces, migrate in very small chunks is not a viable solution.
Any ideas? Is Confluence Cloud Migration Assistant really poorly optimized?
Thanks!