Hi everyone,
We’re working on building Retrieval Augmented Generation (RAG) for our Gen-AI applications using our Confluence, but we’ve hit a few roadblocks and would love some guidance.
Specifically, we’re using frameworks like LangChain and LlamaIndex, which successfully pull context from all pages within a Space. However, we’re facing challenges when it comes to extracting large documents. Additionally, some queries through the Confluence REST API are returning 500 errors, likely due to unknown limits or restrictions.
Has anyone else encountered these issues? What’s the best approach to efficiently retrieve both pages and large documents from Confluence in aspect of both cloud and on-prem?
Thanks in advance for any insights!