Knowledge Base Ingestion
Everest ingests content through native connectors for Confluence, Jira, and Heretto CCMS, alongside website crawl and direct file upload. Sync frequency is configurable per source: hourly, daily, weekly, or manual. When the same topic exists in multiple sources, the newest revision wins by default, and administrators can set a per-source priority override to pin an authoritative source.HTML and PDF Manual Ingestion
Everest parses and indexes HTML pages, PDF manuals, and product documentation, including multi-page PDFs, embedded tables, and deeply nested HTML structures. Extracted content is chunked with structure preserved so that table rows and section hierarchy remain intact at retrieval time. Images inside documents are processed with OCR so text in screenshots and diagrams is searchable.Website Crawl and Sitemap Ingestion
Public sites such as product and support portals are indexed via sitemap-driven or URL-scoped crawls.- Scope controls: restrict crawls by sitemap, URL prefix, or explicit include and exclude rules.
- Crawl frequency: configurable per source (hourly, daily, weekly, or manual).
- Freshness: each recrawl re-indexes changed pages so the newest published revision is what retrieval returns.

