A Python and Playwright pipeline that turned 2,200+ scattered Shopify help pages into a structured, AI-ready knowledge library for Gemini Notebook (formerly NotebookLM). ChatGPT was used for research and compiling, Claude and Gemini Notebook were used to document and explain it.
At a Glance
- Source: 2,200+ URLs extracted from Shopify’s public sitemap.xml
- Scope: Filtered to 738 URLs covering Shopify Admin functionality only
- Output: 12 clean, category-organized reference documents
- Stack: Python, Playwright, custom JS/CSS content stripping, session-based authentication
- Use case: AI-ready study library for retrieval and Q&A in Gemini Notebook
The Problem
I wanted to get proficient in Shopify Admin quickly, without opening a new store just to have something to practice on. Shopify’s help center covers the material, but it’s built for browsing, not studying: content is spread across hundreds of pages, wrapped in navigation, sidebars, and promotional elements. That structure works fine for a human skimming one article at a time. It breaks down as a study resource, and it breaks down completely as input for an AI tool.
Gemini Notebook and similar tools read documents as flat text. They can’t distinguish an article’s body from the navigation menu wrapped around it. Feed them a page-accurate PDF of a live website, and the AI treats sidebar links and footer boilerplate as part of the content, which degrades summarization and search. Building something actually usable meant solving a content architecture problem, not just a scraping problem: extract the right material, strip it down to signal, and organize it so both I and an AI tool could navigate it.
Requirements
- Cover Shopify Admin functionality only. Exclude storefront setup, theme design, and anything outside the scope of managing an existing store.
- Organize by functional topic, not as one undifferentiated pile of URLs.
- Strip every layout element that isn’t article text: navigation, sidebars, images, icons, promotional boxes.
- Consolidate each topic into a single reference document instead of leaving hundreds of individual files.
- Run unattended across hundreds of pages. No manual copy-paste.
Building the Pipeline
1. Source extraction. Pulled Shopify’s public sitemap.xml and extracted every URL with CLI tools, yielding over 2,200 candidate pages.
2. Scoping and categorization. Filtered the list down to 738 URLs strictly within Shopify Admin scope, removing store setup guides, theme/design content, and anything outside day-to-day store management. Grouped the remaining URLs into 12 functional categories, each defined by its own manifest file listing the URLs in scope.
3. Access layer. Built a Python script using Playwright to load each page and render it to PDF. Shopify’s bot detection blocked unauthenticated automated access, so the script reused a saved browser session-state file to authenticate without manual intervention on every run.
4. Content transformation (“reader mode” pipeline). The first pass produced pixel-accurate PDFs of the live site, complete with menus, icons, and sidebars, unusable for an AI parser. I added a transformation step that ran before each page was printed:
- Removed images, SVGs, videos, canvases, iframes, and figures
- Cleared background images
- Stripped promotional boxes, related-article links, and comment sections
- Flattened hyperlinks to plain text
- Blocked image, media, and font network requests to speed up rendering
- Applied custom CSS to hide headers, footers, and navigation bars
The result was a clean, text-only document per article, with only the substantive content intact.
5. Assembly. Rewrote the script to read each category’s manifest file, generate a clean PDF for every URL in it, then merge the set into a single category-level PDF. That turned hundreds of individual article files into 12 consolidated reference documents.
Validation Before Scaling
Before running all 12 categories, I tested the full pipeline against a single one, Analytics, to confirm the reader-mode transformation was working as intended. Once that output held up, formatting clean, structure intact, I removed the restriction and ran the remaining 11 categories.
Results
The pipeline reduced 2,200+ scattered, navigation-heavy web pages to 12 clean, category-organized reference documents covering the full functional scope of Shopify Admin. Each document is text-only, consistently formatted, and ready for direct ingestion into Gemini Notebook, where I use it for AI-assisted quizzing, summarization, and search instead of manually re-navigating Shopify’s help center.
Why This Generalizes
The core pattern isn’t Shopify-specific: pull a full site’s public documentation from its sitemap, scope it to what’s actually needed, strip everything but article content, and assemble it into topic-level references. That’s a repeatable approach for converting any SaaS platform’s public documentation, or an organization’s internal knowledge base, into a clean corpus for AI-assisted search, training material, or retrieval-augmented workflows.
Lessons Learned
- Start broad, then filter. Pull everything from the sitemap first, then scope down to the actual goal.
- Validate on one segment before committing to the full run. Catching a formatting problem in one category is cheap; catching it after processing all 12 is not.
- Print CSS alone isn’t enough. True reader-mode output requires stripping DOM elements and blocking resource loads, not just hiding them visually.
- Organize before you capture. Structuring URLs into category manifests up front made merging straightforward and the final library far more usable.
Next Steps
The pipeline isn’t limited to Shopify. It can be pointed at any platform with public documentation and a sitemap. Adding change detection, comparing source pages against the last capture, would let the library update automatically as articles are added or revised, closing the last manual step in the process.