I am trying to parse the RemotePage.getContent into a JDOM2 document and then output that document again.
I have so far run into the following problems:
- The results of RemotePage.getContent() is the content of XHTML <body> but without a top element. I got around this by wrapping the getContent() results in a <body> tag
- The next error message I got was 'The entity "oslash" was referenced, but not declared.'. I got around this by slurping xhtml-lat1.ent into the string lat1Entities and prefixing "<body>" with "<!DOCTYPE body [" + lat1Entities + "]> "
- The third error message was a parsing error caused by the "as" namespace for a macro (this one I haven't figured out a workaround for yet)
I also perceive a possible problem with getting XMLOutputter to output the content without the body element (possible workaround: iterate over all element children of Document and output each element separately (but more cumbersome than it has to be)).
Also, I suspect that Confluence will prefer the lat1 (and other special characters) as character entities, and I don't know what XMLOutputter will do here.
Is there a simpler approach? Is there a way I can get the full XHTML for the page, including namespace, and DOCTYPE declarations?
Are there more correct ways of handling the problems I have encountered than the ones I've used so far?