Document decomposition and page extraction require dissecting a master PDF container while isolating the exact content streams and font sub-dictionaries referenced by target pages. Rather than simply deleting unwanted pages, an advanced extraction engine reconstructs a lightweight, independent root catalog containing only the specific page resource nodes, color spaces, and media boxes requested by the user. During selective extraction, orphaned indirect objects—such as unused embedded fonts, invisible metadata layers, and detached annotation arrays from non-selected pages—are pruned through unreferenced object garbage collection. This architectural optimization drastically reduces the resulting file size, producing lean, standalone documents ideal for email distribution and database indexing. Common enterprise applications include extracting signed signature addenda from 100-page commercial lease agreements, separating confidential tax schedules from multi-entity financial returns, and isolating specific chapters from technical manuals for targeted distribution.