How Atlas extracts source content
Atlas processes a source before its text can support summaries, search, chat answers, and citations. Processing quality depends on the source format and whether its content is accessible.
What this page covers
Explain extraction, processing, indexing, unsupported content, and the difference between visible source content and retrievable evidence.
From source to evidence
- Atlas validates the file or URL.
- Atlas extracts readable text or obtains an available transcript.
- Atlas divides the text into retrievable passages.
- Atlas indexes those passages inside the current project.
- Atlas marks the source ready when retrieval can use it.
The source viewer and the retrieval index serve different purposes. A page can be visible while some text is not extractable. A citation can only target evidence that Atlas processed.
Extraction limits
Scanned PDFs can require machine-readable text. Websites can block automated access or place content behind authentication. YouTube sources require an available transcript. Layout, tables, equations, images, and embedded media may not convert perfectly to text.
If a source completes but answers miss important content, open the source and confirm that Atlas can select or search the relevant text. Use a clearer copy of the source when possible.
