Hi Community,
I am designing a scalable data pipeline to ingest Jira data into Google BigQuery for advanced analytics. My goal is to move beyond simple reporting and build a robust Data Warehouse structure.
The challenge I am facing is extracting complete data (Issues + All Comments + All Worklogs + Full Changelog/History) without hitting rate limits or facing severe performance degradation on large instances (50k+ issues).
The Architecture I am considering:
Incremental Loading: Using JQL updated >= last_run_date to fetch only modified issues.
The "Search" API for bulk data: Using POST /rest/api/3/search with fields like ['*all', 'worklog', 'comment'] and expand=['changelog'].
My Bottlenecks/Questions:
Truncation of Nested Arrays: The Search API seems to truncate nested fields (e.g., it only returns the first 20 comments or worklogs, and limited history).
Changelog Performance: fetching expand=changelog significantly slows down the Search response.
Deletions:
I am trying to avoid the "N+1 request" pattern where I loop through every issue ID, as that won't scale.
Any advice on the most efficient "Bulk Export" strategy for these datasets would be greatly appreciated.
Thanks!