Python HTML to PDF: the libraries compared
Python has more PDF libraries than most languages, and several of them are excellent. They split into three families: renderers that read HTML and CSS themselves, toolkits where you draw the page in code, and wrappers that drive an external program such as a browser. Choose by the kind of documents you make and by what you are willing to install on every server.
| Library | How it works | Worth knowing |
|---|---|---|
| WeasyPrint | Its own layout engine turns HTML and CSS into PDF, with strong support for paged media CSS | Runs no JavaScript; needs Pango and related system libraries in every image |
| xhtml2pdf | Pure Python, converts HTML through ReportLab | Understands a subset of CSS, so modern layouts usually need simplifying |
| ReportLab | A toolkit for drawing pages in code (canvas or the Platypus flowables) | Precise and fast for fixed layouts; designers cannot edit the result as HTML |
| pdfkit | A thin wrapper that calls the wkhtmltopdf binary | The wkhtmltopdf repository was archived in January 2023 and its WebKit engine is old |
| Playwright | Drives headless Chromium and calls page.pdf() | Modern CSS, but you host, update and watch a browser process |
| CastPDF | One HTTPS request; Chromium with print features runs on our side | A network call per document and a monthly plan above the free allowance |
If your PDFs are short, styled with simple CSS and WeasyPrint is already in your Dockerfile, keep it. The API earns its place when documents get long and data driven: a 40 row table that must repeat its header on page two, a footer with "page 3 of 5", fonts that look the same on a laptop and in production, and no browser or C libraries to rebuild every time the base image changes.
A first PDF with requests
Sign up, copy a test key from the dashboard (it begins with cpdf_test_) and export it as CASTPDF_API_KEY. Test documents cost nothing, never touch your monthly allowance and carry a small watermark, so you can experiment freely. Install the client with pip install requests, save the script below as timesheet.py and run it.
import os
import requests
TEMPLATE = """
<h1>Timesheet: {{ person }}</h1>
<p>Week {{ week }}</p>
<table>
<thead><tr><th>Day</th><th>Project</th><th>Hours</th></tr></thead>
<tbody>
{% for row in rows %}
<tr><td>{{ row.day }}</td><td>{{ row.project }}</td><td>{{ row.hours }}</td></tr>
{% endfor %}
</tbody>
</table>
"""
payload = {
"html": TEMPLATE,
"data": {
"person": "Lena Fischer",
"week": "2026-W40",
"rows": [
{"day": "Monday", "project": "Data migration", "hours": 7.5},
{"day": "Tuesday", "project": "Data migration", "hours": 8},
{"day": "Wednesday", "project": "Client workshop", "hours": 6},
],
},
"mode": "print",
"filename": "timesheet-2026-W40",
}
res = requests.post(
"https://api.castpdf.com/v1/pdf",
headers={"Authorization": f"Bearer {os.environ['CASTPDF_API_KEY']}"},
json=payload,
timeout=60,
)
if not res.ok:
err = res.json()["error"]
raise SystemExit(f"{err['code']}: {err['message']}")
with open("timesheet.pdf", "wb") as fh:
fh.write(res.content)
print("saved timesheet.pdf with", res.headers["x-pages"], "page(s)")Passing json= makes requests serialise the dict and set the Content-Type: application/json header for you. Because the body includes data, the HTML is treated as a Liquid template: the {% for %} loop builds one table row per entry before Chromium lays out the page. Note that Liquid uses the same double braces as Jinja2, so do not run this string through Jinja first. "mode": "print" switches on the print engine, which repeats the thead if the week ever spills onto a second page. The body of a successful response is the PDF itself, and res.content holds its raw bytes.
Fill a saved template from a dict with httpx
Pasting HTML into Python strings gets old once a designer wants to change the layout. Move the markup into a saved template in the dashboard instead: you edit HTML and CSS with a live preview, and colleagues can change the sample values in Simple mode without touching code. Your Python then sends only the template ID and the data. The example below fills a weekly operations report and asks for a signed link rather than the file, which is handy when a Celery task stores the link on a database row for someone to download later.
import os
import httpx
report = {
"team": "Platform reliability",
"week": "2026-W40",
"deploys": 38,
"incidents": [
{"opened": "2026-09-29", "service": "checkout", "minutes": 14},
{"opened": "2026-10-01", "service": "search", "minutes": 6},
],
}
timeout = httpx.Timeout(75.0, connect=10.0)
with httpx.Client(timeout=timeout) as client:
res = client.post(
"https://api.castpdf.com/v1/pdf",
headers={
"Authorization": f"Bearer {os.environ['CASTPDF_API_KEY']}",
"Idempotency-Key": f"ops-report-{report['week']}",
},
json={
"template_id": os.environ["OPS_REPORT_TEMPLATE_ID"],
"data": report,
"filename": f"ops-report-{report['week']}",
"response": "url",
},
)
body = res.json()
if res.is_error:
raise RuntimeError(f"{body['error']['code']}: {body['error']['message']}")
print(body["url"], body["pages"], "pages, link expires", body["expires_at"])The Idempotency-Key names the document in your own terms, here the report week. If the worker crashes after sending and the task is retried, CastPDF recognises the key and returns the first document instead of rendering and billing a second one. The data dict replaces the template’s sample data completely, so include every field the template reads. Prefer async code? httpx.AsyncClient takes the same arguments; just await client.post(...) inside an async with block.
The JSON reply carries id, pages, bytes, the signed url and expires_at. The link works without an API key until it expires, so you can hand it to a browser. Keep the id too: GET /v1/pdf/:id returns a fresh link later, and the API reference documents every field of both calls.
Handle errors and retries in Python
Every failure comes back as JSON with an error object holding a stable code, a human message and a docs_url. Only a few statuses deserve a retry: 429 (wait the number of seconds in Retry-After), 503 (the service is briefly busy, so back off) and network failures. Everything else in the 4xx range describes the request itself, so sending it again gives the same answer. This helper uses one requests.Session for the whole process and raises a typed exception you can catch in views or tasks.
import os
import random
import time
import requests
API_URL = "https://api.castpdf.com/v1/pdf"
session = requests.Session()
session.headers["Authorization"] = f"Bearer {os.environ['CASTPDF_API_KEY']}"
class CastPdfError(Exception):
def __init__(self, status, code, message):
super().__init__(f"{status} {code}: {message}")
self.status = status
self.code = code
def render_pdf(payload, idempotency_key, attempts=5):
for attempt in range(1, attempts + 1):
try:
res = session.post(
API_URL,
json=payload,
headers={"Idempotency-Key": idempotency_key},
timeout=(10, 75),
)
except (requests.ConnectionError, requests.Timeout):
if attempt == attempts:
raise
time.sleep(2 ** attempt)
continue
if res.ok:
return res.content
if res.status_code in (409, 429, 503) and attempt < attempts:
wait = res.headers.get("Retry-After", "")
delay = int(wait) if wait.isdigit() else 2 ** attempt
time.sleep(delay + random.random())
continue
error = res.json().get("error", {})
raise CastPdfError(res.status_code, error.get("code"), error.get("message"))A 409 idempotency_conflict means the first request with that key is still rendering, so waiting and repeating it is safe too. A network timeout is the case idempotency exists for: you cannot know whether the first attempt finished, and resending the same key makes the question harmless. The small random jitter stops a fleet of workers from retrying in lockstep.
Production checklist for Python services
- Always pass a timeout.
requestswaits forever by default, and a hung worker is worse than a failed one. Use something liketimeout=(10, 75): a short connect limit and a read limit above 60 seconds. - Allow for the render budget. A document may take up to 30 seconds to render, so a 30 second gunicorn worker timeout is too tight for long reports. Move big documents to a Celery or RQ task instead of rendering inside the request.
- Reuse connections. Create one
requests.Sessionorhttpx.Clientper process, not per call, so TLS handshakes are not repeated for every PDF. - Read the key from the environment, a
.envfile loaded by your settings, or your secrets manager. Never commit it, and never send it to browser JavaScript. Swapcpdf_test_for acpdf_live_key at launch; the code stays the same. - Build the
Idempotency-Keyfrom your own identifiers (order number, report week, user ID plus date) so retries from task queues never create duplicates. - Watch the size limits: up to 50 pages and 40 MB per PDF. Resize photos before embedding them; they are the usual cause of heavy files.
- Log the
x-request-idheader of failed responses. It identifies the request if you ever need to ask support about it.
Common errors from Python code
| Symptom or code | Likely cause | Fix |
|---|---|---|
KeyError: CASTPDF_API_KEY | The variable is set in your shell but not in the service, container or IDE run config | Add it to the environment where the process actually starts |
invalid_api_key (401) | A missing Bearer prefix, a stray newline from a file, or a revoked key | Strip the value and send Authorization: Bearer <key> |
invalid_request (400) | An unknown field (for example margin instead of margins) or both html and template_id sent | Compare the payload keys with the API reference |
template_render_error (422) | A Liquid tag failed, often Jinja syntax such as {{ x|default("") }} copied into the template | Check details.line and details.column and use Liquid filters |
content_lost (422) | Print mode found content that would be cut off, such as a wide table | Let the table wrap, shrink it, or switch to landscape |
requests.exceptions.ReadTimeout | The client gave up before the render finished | Raise the read timeout above 60 seconds and retry with the same key |
rate_limited (429) | Too many requests per minute (test keys allow 20) | Sleep for Retry-After seconds, then retry |
The quickstart covers the dashboard side: keys, starters and your first template. Building a Django app? The Django page shows a view that returns the PDF as a download. Running JavaScript next to your Python? See the Node.js page for the same calls with fetch.