How to convert a web article to a clean PDF
Paste the article URL or the raw article text into the Primary input at the top of the page, choose whether the source is a URL or pasted text, then pick a template, a paper size, and a colour theme. Press Convert article and Web2PDF sanitises the source, rebuilds the document model (headline, byline, standfirst, headings, paragraphs, pull quotes, code blocks, lists, and figures with captions), and typesets it against the page geometry you selected. The output pane shows the contract result, the document metrics, and a live preview sheet rendered with the real reading measure. From there you can copy the textual export, download a paginated PDF, download clean Markdown, or print the styled sheet.
What clean extraction removes from a web page
Most articles arrive wrapped in machinery: a sticky masthead, a navigation menu, a cookie consent bar, a newsletter interstitial, an autoplay video, a floating share rail, a related-posts grid, and a footer of legal links. Web2PDF classifies every block it can see and keeps only the editorial spine. Promotional paragraphs, subscription prompts, comment sections, and repeated bylines are dropped, tracking parameters are stripped from every retained link, and the reading order is rebuilt from heading structure rather than from DOM order, so a sidebar that appears before the article in the markup does not jump to the front of your PDF.
Typesetting controls that make a PDF readable
The typography panel drives the page geometry that actually decides whether a PDF is pleasant to read. Margin sets the reading measure in millimetres, so the line length stays near the 60–70 character sweet spot instead of running the full text width of a wide monitor. Line height controls the leading between baselines. Base font size sets the body type scale, and every other size — headline, subhead, code, caption — is derived from it, so the document stays in proportion. Paper size switches between A4, US Letter, and Kindle, each with its own page box, running head, and page-number treatment, and you can toggle whether figures are kept or stripped for a text-only archival copy.
Reader-view HTML, Markdown, and PDF from one conversion
The same document model produces three deliverables, so you never convert twice. The reader-view HTML card is a self-contained, style-inlined fragment you can drop into a CMS, an email, or a notes app. The Markdown export preserves the structural semantics — heading levels, fenced code with a language tag, block quotes, ordered and unordered lists, figure captions as italic alt text — so it round-trips cleanly into any static site generator or knowledge base. The PDF export renders the styled sheet through the browser print pipeline, which means the fonts, page boxes, running heads, and page-break rules you configured are exactly what lands in the file, with no server round-trip and no upload.
Page breaks, widows, and orphans handled for you
A paginated document fails in recognisable ways: a heading stranded at the foot of a page, a three-line paragraph split across a page boundary, a figure separated from its caption, a code block sliced in half. Web2PDF applies a break policy to the document model before export — headings keep with their following block, paragraphs avoid single-line widows and orphans, figures and tables keep together with their captions, and long code blocks break at a safe boundary rather than mid-identifier. The pagination summary reports the resulting page count, the estimated reading time, and any block that had to be relaxed to fit the page geometry, so you know why a break landed where it did.
Is converting a web article to PDF private?
Yes. Web2PDF is a client-side tool: the article text, the URL, the template you pick, and the generated document stay in your browser tab, and the saved library is persisted only to your own browser storage so a reload does not lose your work. There is no sign-up, no email gate, and no upload step. Because a browser cannot fetch a cross-origin page on your behalf without a proxy, URL mode honestly reports the cross-origin restriction and falls back to the article text you paste, rather than pretending a download succeeded. The tool ships an Axle API adapter boundary so a self-hosted deployment can point the same article and note operations at a Go/SQLite backend, and when no backend is configured the adapter serves the identical list, get, create, update, and delete surface from a deterministic local store.
Frequently Asked Questions
How do I convert a web article to a clean PDF for free?
Paste the article URL or its text into the Primary input, choose the source mode, paper size, and template, and press Convert article. Web2PDF strips the navigation, adverts, cookie banners, and related-post rails, rebuilds the article into a document model, and typesets it against your page geometry. Download PDF renders the styled sheet through your browser print pipeline, so the conversion is free, needs no account, and never uploads your article.
Can I paste an article URL, or do I need the text?
Both are supported. URL mode extracts the domain, slug, and publication trail from the link and builds the document from the article text you paste alongside it, because a browser cannot read a cross-origin page without a proxy. If the fetch boundary is hit, Web2PDF renders a visible extraction warning explaining exactly what happened instead of silently failing, and the rest of the workflow continues unchanged.
What is the difference between the PDF, HTML, and Markdown exports?
They come from the same document model, so the content is identical. The PDF export paginates the styled sheet with running heads, page boxes, and break rules. The reader-view HTML card is a self-contained style-inlined fragment for a CMS, email, or notes app. The Markdown export keeps the structural semantics — heading levels, fenced code, block quotes, lists, and captions — so it round-trips into any static site generator or knowledge base.
How do the margin, line height, and font size controls affect the PDF?
Margin sets the reading measure in millimetres so the line stays near the 60–70 character optimum instead of spanning a wide monitor. Line height sets the leading between baselines, and base font size sets the type scale from which the headline, subhead, code, and caption sizes are derived. Because the preview sheet renders with the same measure and leading values, what you see in the preview is what the print pipeline paginates.
Are images and captions kept in the converted document?
Yes, when you leave the figures toggle on. Each retained figure keeps its caption and its alternative text, and the break policy keeps a figure together with its caption rather than letting a page boundary separate them. Turn the toggle off for a text-only archival copy: the figure is replaced by its caption as an italic standfirst so the surrounding argument still reads correctly.
Is Web2PDF private and does it work on mobile?
It is private and runs entirely in the browser: your article, your settings, and the saved library stay in local storage on your own device, and nothing is uploaded. It is also responsive — the source module and the output module stack into a single column on a phone, the typography controls stay reachable without horizontal scrolling, and the preview sheet scales down while keeping its true page proportions.