XML Canonicalizer (C14N)
XML Canonicalizer
Two XML documents can mean exactly the same thing and share not one byte: attribute order, quote style, empty-element syntax and redundant namespace declarations are all free choices a parser throws away. C14N is the agreed way to make that choice for everyone, so the two documents come out identical — which is what makes signing and comparing XML possible at all.
What the canonical form fixes
Every one of these rewrites happens, and none of them changes the document's meaning:
- Attributes are sorted — namespace declarations first (by prefix), then attributes by namespace URI and then local name.
- Attribute values are always double-quoted, with
&,<,"and the tab, newline and carriage-return characters escaped. <a/>becomes<a></a>. There is no such thing as a self-closing element in the canonical form.- CDATA sections become ordinary text, character references are resolved, and
&,<,>and carriage returns are re-escaped. - The XML declaration and the DOCTYPE are dropped; the output is UTF-8 by definition.
- Namespace declarations that repeat one already in force are not written again.
Inclusive or exclusive
This is the choice that actually matters, and getting it wrong is the usual reason a signature that verified yesterday fails today.
Inclusive C14N 1.0 gives every element the full set of namespace declarations in scope for it — including ones inherited from ancestors it never uses. That is fine while the document stays whole. Lift a signed fragment into a different envelope, though, and the ancestors are different, so the canonical form changes and the signature breaks.
Exclusive C14N 1.0 writes only the declarations an element visibly uses: its own
prefix and the prefixes of its attributes. A fragment then canonicalises the same way wherever it
is embedded, which is why SOAP and SAML use it. The InclusiveNamespaces PrefixList box is
the escape hatch for prefixes that matter but are not visibly used — typically because they appear
inside attribute values, like a QName in an xsi:type or an XPath expression.
List them space-separated; #default means the default namespace.
Signing: the digest
In XML-DSig, DigestValue is the base64 of the SHA-256 of the canonical bytes — exactly
the string in the output pane, encoded as UTF-8. It is printed under the panes, together with the
hex form for command-line comparison. If your digest does not match the one in an existing
signature, the difference is almost always one of three things: inclusive where exclusive was
meant, a missing prefix in the InclusiveNamespaces list, or comments included when the algorithm
URI says they are not.
The algorithm URIs, for the reference you are probably filling in:
http://www.w3.org/TR/2001/REC-xml-c14n-20010315
http://www.w3.org/TR/2001/REC-xml-c14n-20010315#WithComments
http://www.w3.org/2001/10/xml-exc-c14n#
http://www.w3.org/2001/10/xml-exc-c14n#WithComments
Canonicalising one part of the document
A signature almost never covers the whole file. Put an XPath in the Subtree box —
//Signature, /Envelope/Body, //*[@Id='Body'] — and only the
first element it matches is canonicalised. Namespace declarations inherited from the ancestors you
excluded are handled the way the algorithm you picked handles them: inclusive mode writes them onto
the fragment's root, exclusive mode writes only what the fragment uses. That difference, on that
one element, is the whole reason exclusive canonicalisation exists.
Common pitfalls
- Whitespace is data. C14N does not strip or normalise whitespace between elements. A pretty-printed document and its minified twin have different canonical forms and different digests. Canonicalise the bytes you received, not the bytes you reformatted — if you have already lost them, the formatter cannot put them back.
- Line endings are already normalised. The parser turns CRLF into LF before the
canonicaliser ever sees it, as the XML spec requires. A literal carriage return that survives —
because it was written as

— is re-escaped rather than emitted raw. - Entities must be resolvable. A document referencing an external DTD entity cannot be canonicalised here; the browser's parser will not fetch it. Expand it first.
- Comments are excluded by default, matching the plain algorithm URI. If the
reference's URI ends in
#WithComments, tick the box.
FAQ
Is this C14N 1.1?
No — 1.0 inclusive and exclusive. Version 1.1 differs in how it handles xml:base
when a subtree is lifted out of its document, and is rare outside of a few specifications that
call for it explicitly. Everything in XML-DSig, SOAP and SAML you are likely to meet uses one of
the two here.
Why did my document get bigger?
Because self-closing elements were expanded and inherited namespace declarations were written out where they first take effect. Canonical form is not compact form — the minifier is the tool for size.
Two files look identical but the digests differ.
Diff the canonical forms rather than the originals: run both through here, then paste the results into the XML diff. It is usually trailing whitespace inside an element, a tab where the other file has spaces, or one file having a BOM.
Does anything leave my machine?
No. Parsing, canonicalisation and the SHA-256 all run in the page. The digest needs the browser's WebCrypto, which is only available on HTTPS pages — which this one is.