Why Chromium on Lambda is hard work
Lambda is a natural home for document jobs: they arrive in bursts, run for seconds and then stop. The awkward part is the renderer. A full Chromium build does not fit comfortably within Lambda’s deployment limits, so every headless browser approach on Lambda is some kind of workaround.
| Approach | How it is packaged | The catch |
|---|---|---|
Puppeteer with a compressed Chromium (for example @sparticuz/chromium) | A layer or the deployment zip | Must fit the 250 MB unzipped limit for code plus layers; the browser decompresses into /tmp on cold start; versions must match puppeteer-core |
| Chromium in a container image | An image of up to 10 GB from ECR | No size squeeze, but larger images and a browser launch make cold starts slower, and you patch the image yourself |
| PDFKit or pdf-lib | Plain JavaScript in the zip | Small and fast; layouts are drawn in code and multi page tables are your problem |
| CastPDF | A few lines in the zip | One HTTPS request per document; a plan above the free allowance |
Memory is the other hidden cost. Chromium needs far more of it than a typical API function, and Lambda ties CPU share and price to the memory setting, so every invocation pays for the browser even when the document is tiny. Fonts are another recurring surprise: the Lambda base image ships very few, so text renders in a fallback face until you package your own. With an API call, the function can run on the minimum memory setting and the renderer already has Noto, Liberation and DejaVu fonts installed.
A Node 20 handler that writes to S3
The example renders a monthly usage report for each customer of a SaaS product. An EventBridge schedule or a Step Functions map state invokes the function once per customer with the figures for that month. The handler reads the CastPDF key from Secrets Manager (cached across warm invocations), renders the report from a saved template, writes it to S3 and returns a presigned link valid for seven days.
import { S3Client, PutObjectCommand, GetObjectCommand } from '@aws-sdk/client-s3';
import { getSignedUrl } from '@aws-sdk/s3-request-presigner';
import { SecretsManagerClient, GetSecretValueCommand } from '@aws-sdk/client-secrets-manager';
const s3 = new S3Client({});
const secrets = new SecretsManagerClient({});
let cachedKey;
async function castpdfKey() {
if (!cachedKey) {
const out = await secrets.send(new GetSecretValueCommand({ SecretId: process.env.CASTPDF_SECRET_ID }));
cachedKey = out.SecretString;
}
return cachedKey;
}
export const handler = async (event) => {
const { customer, month, usage } = event;
const res = await fetch('https://api.castpdf.com/v1/pdf', {
method: 'POST',
headers: {
Authorization: `Bearer ${await castpdfKey()}`,
'Content-Type': 'application/json',
'Idempotency-Key': `usage-report-${customer.id}-${month}`,
},
body: JSON.stringify({
template_id: process.env.REPORT_TEMPLATE_ID,
data: { customer, month, usage },
filename: `usage-${customer.id}-${month}`,
}),
signal: AbortSignal.timeout(75_000),
});
if (!res.ok) {
const body = await res.json().catch(() => ({}));
throw new Error(`CastPDF ${res.status} ${body.error?.code}: ${body.error?.message}`);
}
const Key = `reports/${customer.id}/${month}.pdf`;
await s3.send(
new PutObjectCommand({
Bucket: process.env.REPORT_BUCKET,
Key,
Body: Buffer.from(await res.arrayBuffer()),
ContentType: 'application/pdf',
ContentDisposition: `attachment; filename="usage-${month}.pdf"`,
}),
);
const url = await getSignedUrl(s3, new GetObjectCommand({ Bucket: process.env.REPORT_BUCKET, Key }), {
expiresIn: 7 * 24 * 3600,
});
return { key: Key, url, pages: Number(res.headers.get('x-pages')) };
};Creating the clients outside the handler lets warm invocations reuse their connections, and caching the secret avoids a Secrets Manager call on every run. Throwing on failure is deliberate: Lambda retries failed asynchronous invocations, and because the Idempotency-Key names the customer and month, a retry after a timeout returns the report that may already exist instead of creating a second one. A presigned URL signed with the function’s role credentials stops working when those temporary credentials expire, so for links that must last the full week, sign them with a longer lived identity or generate them on demand.
Timeout, memory and secrets
A new Lambda function times out after 3 seconds, which is far too short for document rendering. Raise it above 60 seconds. Memory can stay low, since the function mostly waits on the network. Store the key in Secrets Manager (or, for a quick prototype, an encrypted environment variable) and give the execution role secretsmanager:GetSecretValue on that secret plus s3:PutObject and s3:GetObject on the bucket.
aws secretsmanager create-secret --name castpdf/api-key --secret-string "$CASTPDF_API_KEY"
aws lambda update-function-configuration \
--function-name usage-report-pdf \
--timeout 90 \
--memory-size 256 \
--environment "Variables={CASTPDF_SECRET_ID=castpdf/api-key,REPORT_BUCKET=acme-usage-reports,REPORT_TEMPLATE_ID=$REPORT_TEMPLATE_ID}"Use a test key (it begins with cpdf_test_) while you wire everything up: test documents are free and watermarked. Switching to a cpdf_live_ key later is a single put-secret-value call, and cached warm instances pick it up once they are recycled.
Behind API Gateway: return CastPDF’s link
When a user clicks a button and waits, the Lambda usually sits behind API Gateway or a function URL. Returning the PDF bytes through that path means base64 encoding, binary media type settings and a 6 MB response payload limit for synchronous invocations. Skip all of it: ask CastPDF for "response": "url" and return the signed link it gives you, which works without a key until the stored file expires.
export const handler = async (event) => {
const report = JSON.parse(event.body ?? '{}');
const res = await fetch('https://api.castpdf.com/v1/pdf', {
method: 'POST',
headers: { Authorization: `Bearer ${process.env.CASTPDF_API_KEY}`, 'Content-Type': 'application/json' },
body: JSON.stringify({
template_id: process.env.REPORT_TEMPLATE_ID,
data: report,
filename: 'usage-report',
response: 'url',
}),
signal: AbortSignal.timeout(25_000),
});
const body = await res.json();
if (!res.ok) {
return { statusCode: 502, body: JSON.stringify({ code: body.error.code }) };
}
return {
statusCode: 200,
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ url: body.url, expires_at: body.expires_at, pages: body.pages }),
};
};The shorter abort signal is there because API Gateway’s integration timeout is about 30 seconds by default. Most single documents come back well within that, but a long report might not. For those, invoke the S3 version of the function in the background, then let the browser poll for the stored link or notify the user from your own app when it exists.
The report template and its JSON
The report starter is a good base: a cover summary, highlight bullets and a table that can run across pages with its header repeated. Copy it in the dashboard and adapt the HTML and CSS with a live preview. Each invocation then sends only the numbers, like this:
{
"template_id": "8c0e2a4b-6d8f-4a1c-9e3b-5a7c9e1b3d4f",
"data": {
"customer": { "id": "cus_2041", "name": "Lumen Analytics GmbH", "plan": "Growth" },
"month": "2026-09",
"usage": {
"api_calls": 1284330,
"active_users": 412,
"storage_gb": 86.4,
"by_week": [
{ "week": "2026-W36", "api_calls": 301200 },
{ "week": "2026-W37", "api_calls": 322870 },
{ "week": "2026-W38", "api_calls": 318040 },
{ "week": "2026-W39", "api_calls": 342220 }
]
}
},
"filename": "usage-cus_2041-2026-09"
}Use {{ usage.api_calls | number: 0, "de-DE" }} in the template to print large counts with the right separators for each customer’s locale. The data object replaces the template’s sample data completely, so include every field the layout reads.
Production checklist for Lambda
- Set the function timeout to 90 seconds or more and keep the fetch abort a little below it. A render can take up to 30 seconds, plus queue time on a busy minute.
- Bundle the AWS SDK clients you use with esbuild or your bundler instead of relying on the copy in the runtime, so a runtime update cannot change their behaviour under you.
- Configure a dead letter queue or an on failure destination for asynchronous invocations, so a report that fails every retry is recorded rather than lost.
- Limit concurrency on scheduled fan outs. A thousand customers rendered at once will hit rate limits; reserved concurrency or a Step Functions map with a max concurrency smooths the burst.
- Block public access on the report bucket and share files only through presigned URLs. Add a lifecycle rule to expire reports you no longer need to keep.
- Stay within 50 pages and 40 MB per document. Split year long reports by month rather than producing one huge file.
Troubleshooting PDF Lambdas
| Symptom | Likely cause | Fix |
|---|---|---|
Task timed out after 3.00 seconds | The function still has the default timeout | Raise the timeout above 60 seconds |
AccessDeniedException from Secrets Manager | The execution role cannot read the secret | Allow secretsmanager:GetSecretValue on its ARN |
The presigned URL returns ExpiredToken | It was signed with temporary role credentials that have expired | Sign links on demand, or with a longer lived identity |
Cannot use import statement outside a module | The handler file ends in .js without "type": "module" | Rename it to index.mjs or set the type in package.json |
rate_limited (429) during a fan out | Too many renders started in the same minute | Cap concurrency and retry after the Retry-After seconds |
render_timeout (504) | The document took more than 30 seconds to render | Check for slow external images, or split the report |
The quickstart covers keys and starter templates, and the API reference describes every field and response header. The Firebase page solves the same problem with a Cloud Function, the Supabase page with an Edge Function, and the Node.js page shows the plain calls these handlers are built on.