XML Splitter
XML Splitter
Cutting an XML file into byte ranges gives you one document and a pile of fragments. Every part has to be a document on its own — same prolog, same root, same namespace declarations, and the channel-level metadata repeated rather than left behind in part one. That is the whole job, and it is why split is not the tool for this.
How it decides where to cut
It looks for the element that repeats: <item> in an RSS feed, <product> in a merchant feed, <row> in a database dump. Every repeating path in the document is listed in the Record picker with its count, deepest first, and the deepest one is selected for you — that is nearly always right, and when it is not, the count in the picker tells you at a glance which path is.
The cut then happens between whole records. A record is never split across two files, so no part can be half a product.
What every part keeps
- The XML declaration — written as UTF-8, which is what the parts are encoded in.
- The root element with all of its attributes, including every
xmlnsdeclaration. A part of a namespaced feed is still in that namespace. - The ancestor chain down to the records —
<rss><channel>, not just<rss>. - Every non-record sibling. An RSS
<channel>holds<title>,<link>and<description>next to its items; each is copied into every part, because a feed missing them is not a feed. This is the thing generic splitters drop.
Which is why the bytes out are more than the bytes in — the status line says by how much. The wrapper is paid for once per part.
The four ways to split
- Records per file — the predictable one. 100 records per file gives files you can reason about, and the last one is short.
- Number of files — when the constraint is the count rather than the size: eight parts for eight workers.
- File size — records are packed until the next one would cross the limit, so no part goes over. Useful against an importer with a hard upload cap. The limit applies to the records; a very large wrapper is added on top of it.
- One file per record — for a per-item pipeline. Capped at 500 files in one go; past that, split in two passes.
Getting the parts out
Every part is listed under the panes with its record count and size. Click one to preview it in the right pane; shift-click to download just that one. Download all packs them into a plain ZIP built here in the page — stored entries, no compression step, which every unzip tool reads.
Privacy
100% client-side. The file is parsed, split and zipped in your browser; nothing is uploaded. See the privacy policy.
Related
To go the other way, format or minify what comes back. To check a part before shipping it, use the validator, or the schema validator if you have the XSD. If the records are what you actually want rather than the XML, XML to CSV flattens the same repeating element into rows.