It is 12:40. You have a meeting at 1:15. You were going to read the paper your advisor sent at lunch, the one that is fourteen pages of dense two-column IEEE typesetting with three equations on the first page, and now you are walking to grab a sandwich instead. You could sit down with the PDF on the bus on the way back. Or you could pop in your headphones, hit play, and have the paper read to you on the walk.
For most people, that second option is theoretical. It exists on paper — every PDF-to-speech app technically supports academic documents — but in practice the audio is so badly garbled that you give up by page two. The abstract runs straight into the introduction. Citation brackets are read aloud as “open square bracket twelve close square bracket.” A subscript drops mid-sentence and you cannot tell whether you missed an exponent or a footnote.
This post is about why papers are particularly hard for text-to-speech, what structure-aware reading actually does for them, and a few practical tactics that work whichever app you end up using.
Why papers are hard for TTS
A research paper is not a long-form essay typeset as a Word document. It is a specific format that academia has evolved over fifty years, and almost every feature of that format makes life harder for a generic TTS engine.
Dense two-column LaTeX layouts. Most papers in computer science, physics, and engineering are typeset in two columns. Naive text extractors read down the left column to the bottom of the page, then jump back up to the top of the right column — but more naive ones interleave the columns line by line, producing word salad. Even when extraction is correct, the reading order is interrupted by figures and tables that span both columns and arrive mid-thought.
Equations rendered as Unicode plus images. LaTeX equations get exported to PDF as a mix of Unicode math characters, font swaps to Computer Modern, and embedded raster images for the parts that won’t fit in a font. A TTS engine sees “ε₀ ∂E/∂t” and produces noise. The visual reader sees an integral and skims past it; the audio listener gets ambushed.
Citations in bracket notation. “We extend the approach of [12, 14] to handle the case described in [3].” Visually, your eye glides over the brackets and lands on the content. In audio, an unmodified engine says “open square bracket twelve comma fourteen close square bracket.” Three of these per paragraph and you cannot follow the argument.
Footnotes. PDFs encode footnote markers as superscript digits that the extractor often dumps inline. You get “the result generalizes³ to higher dimensions” emerging as “the result generalizes three to higher dimensions.” Five footnotes in and you have lost the thread.
The references list at the end. Eight pages of densely packed bibliographic entries, each one a perfect storm of last names, initials, titles in title case, journal abbreviations, and DOI URLs. Reading the references aloud is rarely what you want; most apps will happily plow through them anyway.
Section headings at three levels. Papers use numbered sections — 1, 1.1, 1.1.1 — to nest arguments. The numbering carries real information: “3.2” is a subsection within section 3, and your brain uses that hierarchy to keep track of where you are. A flat-text TTS engine reads “three point two related work” as if it were a single sentence and the structure is lost.
Each of these on its own is a small annoyance. Together, they are the reason most people who try to listen to a paper give up by page two.
What structure-aware reading actually does for papers
Structure-aware reading is what Audris does — and the longer post on the pipeline explains the mechanics in detail. Here is what it means specifically for an academic paper.
A 1.5-second pause before each new top-level section. When the audio gets to “1. Introduction” or “3. Related Work,” there is a brief, deliberate beat of silence before the heading. Your brain registers the boundary the same way it would visually — you know a new section just started, and you adjust your attention.
A shorter pause and a slight rate slowdown on the heading itself. The title of the section is read at 0.97 of the normal rate with a touch of emphasis on the colon or period that follows. It lands differently from body text, the way it lands differently visually.
Citation-aware emphasis. When the analyzer sees a bracket-citation pattern, it can read it as “in reference twelve” or skip it entirely on a setting, instead of barking out the punctuation. Footnote markers can be suppressed and recovered later via a tap-to-jump table of contents.
Section skip via the headphone or steering-wheel button. Because the audio knows where one section ends and the next begins, the “next chapter” button on your Bluetooth headphones can take you to the next section heading — not three minutes forward at random.
The aggregate effect is that you can follow a paper end-to-end during a thirty-minute walk in a way you genuinely could not with flat-text TTS. The structure is what your brain uses to keep track; restoring the structure in audio is what makes the audio listenable.
The commute test
Concretely: it is 8:15 in the morning. You have a thirty-minute walk to the office. You queue up a fourteen-page survey paper in the app, lock your phone, and put your headphones in. Here is what should happen.
The phone reads the title at a slightly slower rate, pauses, then reads the author list. Another pause. Then “Abstract,” with a beat of silence before the body of the abstract begins. The abstract is read straight through at normal pace, ending with the last sentence dropping slightly in pitch and tempo to signal the section is done. Beat of silence.
“1. Introduction.” Beat. Body of the introduction. You are walking past a coffee shop now and not really thinking about the audio — it is just narrating in the background, the way someone would.
You hit a citation block. The reader either skims past it or replaces it with a brief acoustic marker. You did not lose your place.
You are coming up on an intersection. You triple-tap the right headphone button — section skip. The audio jumps to “2. Related Work” with the same 1.5-second lead-in pause. You wanted to skip the introduction because you already know the area; the app honored that. This is the difference between an audio reader and a robot reading text.
Twenty minutes later you are at section 5, the conclusion. You arrive at the office. You unlock the phone, tap pause, the playback position is saved. You sit down with the paper visually for the bits you want to re-read carefully — the equations, the figure captions, the references for one particular result. The walk did the work of the first pass.
This entire workflow depends on the audio knowing where it is in the document. Without structure, you cannot skip cleanly, you cannot tell where a section ended, and you cannot resume meaningfully.
Tips for academic readers
A few tactics that work in any app, framed neutrally — these are about how to listen to a paper, not which app to use.
Read the abstract first; then decide if you want the full paper. This is the universal advice and it applies double in audio. The abstract is usually 150 to 250 words. It will take ninety seconds to listen to. If after the abstract you do not want to hear the rest, the paper was not the right one — skip to the next one in the queue.
Use section skip aggressively. Papers are designed for non-linear consumption. The way an academic reads a paper visually is: title, abstract, conclusion, then back to introduction if interested, then specific sections by need. The same strategy works in audio. Skip to the conclusion early to decide whether to invest the time in the methods section.
Bookmark equations and figures for later. Audio is bad at equations. Always. Even with structure-aware reading, an equation is going to be a list of letters and operators that you cannot really visualize from sound alone. Mark them, then come back to them visually when you have a screen.
Queue up multiple papers for a review session. If you are doing a literature survey and have eight or twelve papers to triage, load them all into your library and play through abstracts only. Twenty minutes of audio gets you through the abstracts of a small literature review; you can decide which papers warrant a full visual read after.
Pick a voice you can listen to for an hour. Voice fit is personal. A voice that sounds great in a thirty-second demo can grate after fifteen minutes. Before you start a long paper, pick a voice and listen to a full minute of a different paper in it. If you are not happy at the end of that minute, switch.
Use slower playback for unfamiliar territory. A paper in your field at 1.0× is comprehensible. A paper in an adjacent field at 1.0× will move past you faster than you can absorb. Slow it to 0.9× or 0.85× and you’ll catch more on the first pass. Speed has a real comprehension cost.
These are workflow tactics, not app-specific tricks. They will improve your listening experience whichever PDF-to-speech tool you use.
Where Audris fits
Audris was built with the academic-paper case in the test set. The structure analyzer specifically handles the patterns that make papers hard: numbered headings at three levels, two-column reading order, citation-bracket detection, footnote marker handling, references-list suppression as an option. It is not magic — equations will still defeat any audio reader — but the pieces that can be solved with formatting analysis are solved.
The free tier supports all of this. Structure-aware reading is the core of the product, not a paid upgrade. If you want to try it on a paper you are already meaning to read, the download links are on the home page. Load a PDF, hit play, see whether the cadence helps.
If a different app fits your workflow better, that is the right outcome. The tactics above will work either way. What matters is that you actually read the papers you have been meaning to read.
Built by one person, for people who actually read.