If you are using JIRA and/or Confluence in a corporate environment you should expect to hit scalability issues quite fast.
Here our current list of configuration changes that are needed in order to be able to scale the system for more users:
- Switch from VM to baremetal, get a machine with SSD and plenty of RAM.
- Install Oracle JVM 1.7, and keep it updated (you can to this with apt-get).
- Never week JIRA or Confluence files on network drives / NAS. You can safely mount the attachments directory from NAS, but nothing else.
- Also, keep Atlassian products updated, not older than 6 months.
- Use Linux and PostgreSQL, don't waste precious time with other configs, these are the ones that work the best and that are used by Atlassian on all of their instances.
- 50% of the memory should be reserved to the JVMs, and at least 30% should be free / used for caching by the OS.
- enable validationQuery on JDBC connection, sooner or later you will lose DB connections and you will hate your life if you do not do this.
- Increase the number of max database connections in both JDBC and the database engine, always configure a max in potgresql +20-30 greater than your worst case.
- Monitor performance, for example we use DataDog and monitor:
- Machine memory
- IO Usage
- CPU Load
- JVM memory
- JVM threads
- Postgresql connections per database
- We use nginx in front of these services, which also adds the SSL layer. Nginx is used to provide temporary out of service messages and allowing us to throttle or even ban some HTTP clients, when needed.
Example of working setups that are sharing a server with 48GB RAM, 32 cores, and SSD:
- JIRA, 300k issues: 7500 MB RAM, 512 max perm size, maxThreads=250, JDBC maxActive=120
- Confluence, 50-100 real users: 6000 MB RAM, 512 max perm size, maxProcessors=140
Number of cores is not so important, it seems that Atlassian products are not able to really use them effectively, the average load is less than 3, and a 100% usage would translate into 32.
Memory and IO speed are essential, SSD added a speed up of almost 10x when it comes to indexing or service start time.
While I do have full HTTP logs that do include backend response time for each request, I am still looking for a tools that is able to parse these and to extract some meaningful information.
We are not fully pleased about the performance and I want to have some realistic data regarding what is causing the slowdowns.
Feel free to add your own hints, so we can build a better tuning tutorial.