Ask an AI tool to read a document and it will almost always give you an answer. What it rarely tells you is how much of the document it actually understood.
Most tools read the easy part — the flowing paragraphs — and quietly drop the hard parts: the pricing table, the dimensioned drawing, the performance curve, the clause written in another language. You don't get an error. You get a confident answer built on a fraction of the file.
Sebtember is built on the opposite principle. Lose nothing.
What gets lost — and what Sebtember keeps
Most tools read the flowing text and quietly drop the parts that are hardest to read. Here is how the tools people compare — NotebookLM, Claude, OpenAI, and the specialist parsers — handle each dimension of a document, and what Sebtember does instead. Scroll the table sideways to see every provider.
| Dimension | Sebtemberthis pipeline | NotebookLMconsumer, managed | Claudeprojects + files | OpenAIfile_search | SpecialistsLlamaParse · Reducto · Azure DI |
|---|---|---|---|---|---|
| What it can read | |||||
| Scanned-PDF OCR | Cloud Vision per page; per-page confidence surfaced as a warning | Managed, opaque | Answer-time page vision; no stored OCR layer | Weak on scans | Best-in-class |
| Digital tables → retrievable structure | Per-row vectors + structured JSON + row-level citation | Structure-lossy, no row index | Read from the page at answer time | Flattened into prose chunks | Schema + bbox tables |
| Multi-column reading order | Gutter detector → VLM; measured on digital and scanned fixtures | Opaque | Visual reading handles it | Text-order dependent | Dedicated reading-order models |
| Technical drawings — value ↔ meaning | Paired at ingestion and stored (“overall height: 737.1 mm”) | Summary only | At answer time, per query | Numbers without referents | Transcription, not interpretation |
| Charts & curves — data-point reading | Axes + units + per-curve points, at ingestion | Descriptive | At answer time | — | Varies by vendor |
| How it gets indexed | |||||
| When comprehension happens | Ingestion, once — and the interpretation is itself retrievable | Managed; long-context at answer time | Answer time, every query | Ingestion, but shallow | Ingestion — output is a file, not an index |
| Chunking sophistication | Contextual enrichment + boundary bridges + heading-aware splits | Long-context stuffing instead | Managed | Fixed 800 / 400 default | Chunking is yours to build |
| Visual-vector retrieval (page embeddings) | Multimodal page vectors in the same index as text, in production | — | — | — | Emerging |
| Hybrid lexical + dense retrieval | Vector + keyword + rank fusion + cross-encoder rerank | No exposed retrieval layer | Managed | Dense-led | Not a retrieval product |
| Multilingual | Original-language storage + bilingual keywords → verbatim exact-match | Strong | Strong | Dense-only — no verbatim match | Strong |
| Corpus-level synthesis | Rollup vectors, so “what’s in here?” has a retrieval target | Long context sees the whole source | Context-window bound | Chunk-level only | |
| Whether you can trust the result | |||||
| Failure honesty | Warnings taxonomy on the document — which pages read poorly, which caps were hit | A source just fails generically | Request-level errors | File status API | Job status APIs |
| Self-healing after a hard kill | Bounded auto-resume from checkpoint; the recovery path is itself tested | Managed, invisible | Managed | Managed | Managed |
| Ingestion observability | Live page-by-page progress + ETA, per-file token breakdown, routing per page | A spinner | — | Status polling | Job status |
| Ingestion-side eval harness | Golden fixtures run against the deployed function; claims are measured, not asserted | Impossible by construction | Impossible by construction | Impossible by construction | Not shipped |
| Prompt-injection hardening at ingestion | Document text is treated as data; structural tags stripped before indexing | Not disclosed | Model-level | Model-level | Parsers don’t prompt |
| What it costs and where it runs | |||||
| Per-file cost attribution | Itemized, per-payer, debt-aware — what this document cost, and who owes it | Seat pricing | Seat pricing | Storage + query metering | Per-page pricing |
| Scale ceiling per document | 25 MB upload · 540 s per run · checkpoint cliff on the largest PDFs | Very large per source | Per-request page limit | Very large per file | Unbounded per-page jobs |
| Source sync & connectors | Manual upload only | Drive, web, YouTube, audio | Connectors, tier-dependent | Build your own | |
| Data residency & tenancy | Runs in a project you control, in a region you choose | Vendor-managed | Enterprise options | Enterprise options | Some offer on-prem |
The Sebtember column is verified against the deployed pipeline. The other columns are our reading of each product in mid-2026; these tools evolve quickly, so treat them as directional.
The row that matters most is "when it reads the file." Reading the page on every question sounds fine until you realise you cannot search by something that was never written down. Sebtember reads once, at upload, and stores the interpretation — so the depth is already there the moment you ask.
We have yet to find another tool that keeps all of this. Most keep the prose and quietly let the rest go.
How it works: many technologies, one job
No single model reads every kind of page best — so Sebtember doesn't rely on one.
Each page is routed to the tool that fits it. A clean digital page uses its own native text layer. A scanned page goes through dedicated OCR. A diagram or a complex table goes to a vision model that interprets it rather than merely transcribing. The results are embedded with best-in-class models, indexed for both meaning-based and exact-keyword search, and re-ranked so the right passage rises to the top.
It is a pipeline of many specialised technologies and providers, each doing the one thing it does best — and it is constantly evolving. Extraction quality isn't a fixed feature; it's a moving target we keep chasing. When a better engine, model, or technique appears, it goes into the pipeline. So the promise isn't "the best extraction of last year." It's the best result available on the day you upload.
How do you know it actually worked?
A promise to lose nothing is worth only as much as your ability to check it. So Sebtember does two things most tools don't.
First, it tells you how well it read your file. Instead of a spinner that ends in "done," you get live page-by-page progress and, when the file finishes, a plain record of anything that went less than perfectly — which pages scanned poorly, where a limit was reached. "Your file is ready, and pages 12–14 were low quality" is a far more honest answer than "your file is ready."
Second — and as a company built on trustworthy answers, this is the part we care about most — the claims in this article are measured, not asserted. The extraction pipeline is checked by golden-set tests that run against the actually-deployed system: known documents in, known values expected out. "It reads tables correctly" isn't a slogan here; it is an assertion that passes or fails on every change we make.
And every answer traces back to the exact page, the exact row, and the exact passage the model saw — so you are never asked to take the result on faith.
The bottom line
Most AI tools lose the hardest, most valuable parts of your documents and never mention it. Sebtember is built to lose nothing — to read the table as a table, the drawing as real measurements, the chart as real data, your language as your language — to store all of it so it stays findable, and to show you how well it did.
Because an answer is only as trustworthy as the reading behind it. You should never have to wonder how much of your own document the AI actually saw.
Upload the file everything else struggles with — the dense price list, the dimensioned drawing, the contract in two languages — and ask it the question you actually care about.