How Atlas extracts source content
Atlas prepares a source in several parts. It extracts or receives text, renders pages when needed, creates a summary, and builds indexes that help chat and search find relevant passages.
A source becomes readable and searchable in stages
Files, websites, and YouTube videos enter Atlas differently:
- A PDF, DOCX file, or PPTX file must be opened, read, and prepared for display.
- A website must return readable page text.
- A YouTube video must provide a transcript.
For PDF-backed files, the processing panel separates content reading, AI context, page rendering, citation checking, layout analysis, and search indexing. A page may become readable before every later stage finishes.
Visible content and retrievable evidence can differ
A document can display readable pages even when Atlas has little searchable text. This happens when a PDF contains page images without a usable text layer or when part of a page cannot be extracted cleanly.
In that state, you can read the rendered page, but chat and semantic search may not find its wording. The processing panel can report No searchable text, Text unavailable, or Semantic search unavailable when the corresponding stage cannot be prepared.
Format limits affect extraction
Websites that require interaction before showing their article can return little or no text. YouTube videos without an available transcript cannot become transcript sources. Complex PDF layouts, damaged files, and image-only pages can reduce the text available to search and chat.
When Atlas cannot prepare the content you need, add a version that exposes that content as readable text. See Troubleshoot source processing for the next action for each visible symptom.
