XML Merge
XML Merge
cat a.xml b.xml produces a file with two roots in it. That is not a document, and every parser you hand it to says so on line one. Merging XML means picking one wrapper, moving the records from every other file into it, and carrying the namespace declarations along so the moved elements still resolve.
How it reads several documents
Paste them one after another, or drop a whole selection of files onto the box. The input is scanned as a stream and cut into documents on element depth — a document ends when its root closes — so the second and later files do not need their own <?xml> declaration, and comments, CDATA sections, processing instructions and a DOCTYPE with an internal subset are stepped over rather than counted as markup.
Each document is then parsed on its own, so a broken one is named by number instead of poisoning the whole run.
What it keeps
- The first document's wrapper — its root element, attributes, and the channel-level metadata beside the records: an RSS
<title>and<description>stay exactly as they were. - Every record, from every file, appended into that wrapper in file order.
- Namespace declarations from the other roots. If file two declares
xmlns:mediaand file one does not, the declaration is copied onto the merged root — otherwise every<media:content>you just merged in becomes unresolvable. A prefix bound to two different URIs cannot be merged; the first file wins and the status line says so.
Which element counts as a record
The Record picker lists every repeating path in the first document with its count, and picks the outermost — <item>, <product>, <row>. That is the same choice the splitter makes, and for the same reason: in an RSS feed the deepest repeating element is <title> inside <item>, and merging on that would be nonsense.
Later files are searched at the same path first. If a file nests its records differently — <item> directly under the root rather than under <channel> — they are found by element name instead, so an exported feed and a hand-written one still merge.
Duplicates and order
- By a field — the usual way to merge feeds.
guid,linkoridfor a feed;@idreads an attribute on the record itself. First occurrence wins. A record with no value for that field is kept, not folded in with every other blank one, and the count is reported separately. - Identical records — compares the record's whole markup with whitespace collapsed. Use it when the files genuinely overlap and there is no id to trust.
- Order — sort on a child element or an attribute. Values that are numbers compare as numbers, and dates as dates: RSS's RFC 822 stamps and Atom's ISO 8601 both parse, so a merged feed can be put in real reverse-chronological order rather than alphabetical. Everything else compares as text, and equal values keep their file order.
- Keep first — trims the merged list after sorting. 0 keeps everything. Set it to 50 with a descending date sort and you have a rolling combined feed.
Privacy
100% client-side. Every file is read, parsed and merged in your browser; nothing is uploaded. See the privacy policy.
Related
To go the other way, the splitter cuts one document back into well-formed parts. Check the result with the validator, tidy it with the formatter, or squeeze it with the minifier. If the merged records are really a table, XML to CSV flattens them into rows, and XML sort reorders a single document in place.