Dear Community,
Backstory
I am at loss here and seeking help. For 2 month now our Jira Data Center has an indexing problem. We tried twice to run a full re-index, which usually took 30-40 minutes to complete. Now, it doesn't even complete after hours and hours. We had multiple downtimes within our business hours with 1.000+ users.
Root Cause
Today, we think we found the culprit: Scripted fields with Scriptrunner.
Since Scriptrunner is well established, it must be something we are doing wrong / using it wrong.
Pseudo Code
Here is some pseudo code of the field causing issues:
def ISSUETYPE1 = ...
def ISSUETYPE2 = ...
def FIELDID1 = ...
def projectKey = ...
// search for issues
try {
issues = Issues.search(
"issuetype IN (${ISSUETYPE1}, ${ISSUETYPE2}) " +
"AND cf[${FIELDID1}] = ${projectKey}"
)
.stream()
.toList();
}
catch (Exception ex)
{
return null;
}
// GUARD: none found
if (issues.size() == 0) { /*do something...*/ return null }
// GUARD: more than one found
if (issues.size() > 1) { /*do something...*/ return null }
// DEFAULT
return issues.first();
One example of many. This is one of the least complex fields causing the issue.
The problem
While performing a full re-index, the index is first deleted and then rebuild.
While rebuilding the index, the scripted field is queried to add to the index.
This field accesses a search, which in turn accesses the index, which in this case DOES NOT EXIST!
The indexer is waiting for the field to complete.
The field seems to raise an exception, log it, create insane disk I/Os, ...
The indexer throws an exception that is waited for 5000 ms and is still waiting, creating even more disk I/Os for logging
Workaround
We ended up deleting those fields.
before: index recovery was at 30% after 5 minutes and hanging to index the delta changes.
after: index recovery completed in 25 seconds.
We will try a full re-index in 2 or 3 weeks, since we are super afraid of the next yet another downtime. Each downtime costs us around 125k USD. And we had 2 of those already.
Question
Any ideas on how to resolve this? Or should we avoid HAPI search in scripted fields altogether?
are we using the stream/toList wrong? I wonder...
We spent huge time to debug our Jira since the logging is non-existing, not only from Scriptrunner but from almost every plugin. We often see "Scriptrunner throw an Exception" with no hint on where, which line, ... It's the search for a grain of sand in the known universe.