Hi Team,We have use case to crawl all spaces and pages from confluence site.But we have encountered an issue where crawler account is not able to discover pages defined with restrictions. We are using Content API as mentioned below:-Sample Endpoint: https://<SiteURL>/wiki/rest/api/content/search?expand=body.storage,restrictions.read.restrictions.user,restrictions.read.restrictions.group,space,ancestors,history,history.lastUpdated,history.contributors.publishers.users,children.attachment,metadata.labels&limit=25&start=0&cql=type in (page,blogpost) and space = <spaceName>. It gives empty response.This crawler account (app authorized using OAuth2 with required scopes) has access to all spaces but not part of restrictions defined at page level.
Basically, we want crawler account to access all spaces and pages (even restricted ones).Can you please help in answering below:1. How crawler account can access to restricted pages?
2. Is there any other endpoint which can be used to get restricted pages by crawler account?
As far as I know, Confluence Cloud admins can't view restricted spaces to which they have not been granted view permission. This is different than Confluence server.However, as the admin you can still grant yourself view permission, and then scrape the space.
Welcome to the Atlassian Community!
No.
The whole point of restrictions is to stop people who should not see pages from seeing them.
You will need to grant your crawler access to the restricted pages if you want it to be able to see them.
Please don't cross-post using multiple user accounts.
You asked the question here in the Developer Community using the account RolIT and got the same answer.
It looks like you're new here. Sign in or register to get started.