A lot of people who own a book in Green Room also own it on paper, and the two copies have always been strangers: coming in from a commute-listen meant skimming pages to find the spot, twice a day, forever. Page Match is the bridge. There is a new camera button on the reader's rail, and it opens a sheet with three doors. Scan a page opens the phone's camera; photograph the page you're holding and the app lands on that passage. Read it aloud listens while you read a few lines from the page and locks on as soon as it is sure — fifteen to twenty-five words of ordinary prose is enough, and a chime tells you when, because your eyes are on the paper, not the screen. Find my page goes the other way: when the ebook's publisher shipped a print page list it answers "you're around page 214" directly; when it didn't, it plays hot-or-cold with a scan ("flip forward about 30 pages"). Spotify shipped the camera half of this for audiobooks in February and called it Page Match too; we kept the name because it's the right one.
The part worth being precise about is what leaves the device. The photo goes to our API, which asks a vision model to transcribe the printed text and hands back the words; the photo is held for that one call and never written anywhere. Speech goes to a transcription service over a socket the app opens with a token our server mints — fifteen minutes, single use, worth exactly one listen; the vendor key never touches the client. What comes back in both cases is text, and the text is matched against the book on your device. The server never sees the book, and it never learns where you landed. That is a reversal of the design we wrote in August, which insisted on on-device text recognition too. We changed our minds for a reason we should say out loud: the recogniser we could ship across five platforms was much worse at a curled, shadowed page than a vision model is, and the vendor we already pay for voices has no camera product at all, so the two halves were always going to come from two places. What survives of the original posture is the half that matters for the boundary: a position is not prose. Nothing here calls the context assembler, and the check that counts its callers still says one.
Matching is a small thing in the reader core, and its tests are against real books. Every section's words are cut into overlapping three- and four-word shingles, hashed to numbers (string keys over a 500,000-word novel would cost more memory than a phone webview will give us), and a transcript votes for the place where its shingles line up. Three-word shingles survive the one word the recogniser garbled that kills every four-word one around it. Two gates stand between a vote and a jump: an absolute floor a six-word refrain can never clear however often it recurs, and a two-and-a-half-times margin over the runner-up. When two places tie — Project Gutenberg prints the same licence sentence at both ends of Alice, which turned out to be the only genuine long duplicate in our whole test corpus — the sheet shows both and asks. It never picks between distant twins for you, and it never moves your position on its own. A match that lies ahead of where you've read asks for a tap first, because position is the spoiler boundary and sync only ever moves it forward.
It costs us money per use, so it is metered in the units you already see. A voice session draws sixty seconds from your narration hours the moment the token is minted; a photo counts as one companion question. Both doors sit behind the paid tiers, the same wall as asking the Green Room anything. The page-list doors are free on every tier, because they run entirely on the device — the plan's reviewer caught that the first draft had paywalled those too, and the rule here is that reader features stay free and only marginal-cost features are paid. The mint-time charge has a sharp edge we knew about and accepted: the server cannot see how long you actually spoke, so a session that fails before it listens still costs the minute. During the first day's testing that edge cost the founder's own account seven minutes. More on why below.
Two bugs made it through a full round of tests, mutation testing, a spoiler audit and two
code reviews, and were found within an hour of the build reaching a real phone. The camera
door opened the photo library. Not a permission problem — the manifest declared the camera
and the phone even asked for it — but package visibility: on modern Android an app cannot
see that a camera app exists unless its manifest declares the intent it means to send, and
the chooser Tauri generates checks for one before launching it, falls back to the picker
when it finds none, and logs the reason where no reader will ever look. One
<queries> block fixed it. The voice door failed instantly with the
wrong sentence on screen: it said it couldn't place what you'd said in the book, when in
fact it had never heard you. The socket opened normally and the vendor then rejected the
session because we forwarded the book's own language tag, en-US, and the
service accepts only three-letter codes like eng. We found that by attaching
Chrome's debugging protocol to the app's webview over a USB cable and reading the first
frame the server sent. The language is now normalised, and a session that dies before it
listens says whether it was the microphone or the connection. Neither bug could be seen in
a browser, in the emulator's test page, or in any unit test; both are now pinned by tests
that assert the declaration our side has to supply.
Two more things came straight from using it. The "Choose a photo" alternative under the camera row is gone — one door is enough, and on a desktop without a camera that same row simply opens a file picker. And the lock chime exists because the person testing it pointed out the obvious: you cannot watch a screen and read a page at the same time. Two rising notes mean it found you; one low note means the minute ran out without a match.
What isn't done. Windows has not had its microphone prompt verified, iOS has no build yet,
and Linux needed a native fix because the webview framework never turns media streams on —
getUserMedia simply hung there until we did. None of the dozen books we
fetched from our test corpus ship a real print page list, so the direct "page 214" answer
is tested only against a synthetic fixture. Pride and Prejudice does carry 906 page
markers, but in its older navigation file, which the EPUB engine at our pinned version
reads only when the newer one has no table of contents — so it reports none. And the
mint-time minute is honest but blunt; if failed handshakes keep costing people, the
accounting will have to move to the end of the session, which means the client telling the
server how long it listened, and trusting it. We'd rather write that down now than
discover it in a support thread.
Join the waitlist — beta invites go out in order, and the devlog keeps you honest company until then.