English

Developer tools · Syntax converters

XML to JSON mapping: attributes, text nodes and the one-vs-many problem

· How it works

xml json data-formats

One XML item mapping to an object while repeated items map to an array
Original ToolAcre vector illustration

There is no single correct way to turn XML into JSON, because XML has attributes, mixed content and ordered children that JSON lacks. This post explains the common mapping conventions and the traps in each.

Why one <item> became an object and two became an array — the XML feed whose JSON shape changes depending on how many entries it has

A feed with `<item>one</item>` produces `"item": "one"`; adding a second sibling changes that property to `"item": ["one", "two"]`. The parser cannot infer that one item was conceptually a list of one because both meanings have identical XML syntax. Code built against only the first sample can therefore fail when production sends the second.

ToolAcre does not hide this instability behind an always-array option. It records a name once as one value and repeated siblings as an array. That direct mapping is easy to inspect, but consumers needing a stable collection shape must use schema knowledge or normalize the result themselves after conversion.

What XML has that JSON does not — attributes, text mixed with elements, ordered siblings, namespaces, comments and processing instructions

XML separates attributes from child elements, preserves sibling order, permits text between elements, carries namespace prefixes, and can contain comments and processing instructions. JSON offers objects and arrays but has no built-in equivalents for those node categories. Any XML-to-JSON result is consequently a chosen projection, not a universal translation.

This reader drops the declaration, comments and processing instructions. Namespace prefixes stay verbatim rather than being resolved: `<ns:item>` becomes the key `ns:item`, while `xmlns:ns` becomes `@xmlns:ns`. Mixed text is joined under one key, so its original position around child elements is lost and a warning says the conversion cannot round-trip.

Attribute conventions — prefixes such as @ or $, why they exist, and how an attribute and a child element with the same name collide

Attributes use an `@` prefix. `<user id="7"><id>other</id></user>` becomes an object with `@id` equal to `"7"` and child `id` equal to `"other"`. The prefix prevents two different XML constructs from colliding in a single object property. It also becomes part of the converter’s contract when JSON is written back to XML.

Other libraries may use `$`, an attributes object or another convention. ToolAcre supports only the visible `@` mapping. Changing that prefix in application code without changing the writer would turn attributes into elements, so preserve it when using the converted value as an intermediate inspection form.

Text node conventions — #text or _ for element content, and what happens when an element has both text and children

An element containing only text collapses to that string. When attributes or children are also present, text lives under `#text`; CDATA is kept separately under `#cdata`. A self-closing tag becomes an empty string. These reserved keys let the writer distinguish ordinary child names from content categories that JSON itself does not define.

Mixed content remains lossy. In `<p>before<b>bold</b>after</p>`, the position of “before” and “after” relative to the child cannot be reconstructed from a joined `#text` property. The tool detects that structural pattern and warns. Use a node-preserving XML API when document order is part of the meaning.

The one-versus-many problem — repeated elements becoming arrays only when repeated, and why consumers must code defensively

Repeated siblings turn into arrays only after repetition is observed. A single `<book>` is one object; two books are an array of objects. This is sometimes called the one-versus-many problem, but it is not a parser defect. The source document simply does not contain a list declaration independent of its occurrences.

Defensive consumers can normalize known paths with schema knowledge: wrap `catalogue.book` when it is not already an array. Do not apply that rule to every property, because an ordinary scalar should not become a list merely for symmetry. The converter deliberately avoids inventing such domain information.

Worked example: converting a small RSS-style document — attributes, a repeated element and a namespace, with the resulting JSON annotated

Convert `<feed xmlns:m="https://example.invalid/meta"><item id="1"><m:title>One</m:title></item><item id="2"><m:title><![CDATA[Two & More]]></m:title></item></feed>`. The root is `feed`; `@xmlns:m` retains the namespace declaration; `item` is an array; each `@id` is text; and the second title contains `#cdata`.

Type inference is off by default, so even `id="2"` stays the string `"2"`. Enabling inference lets the parser read numeric and Boolean text as those JavaScript types, but XML did not declare that intention. The option is a guess controlled by the user, not evidence supplied by the document.

What this does not cover — schema-driven conversion that knows an element is always a list, which requires an XSD or a manual mapping

No XSD is loaded, and no schema-driven list information is available. The converter cannot know that an element is repeatable when only one appears, validate required children, resolve namespace URIs into application types or generate a typed client. A successful parse only establishes well-formed XML accepted by the configured reader.

DOCTYPE declarations are refused before parsing, including harmless ones. That boundary prevents external entity requests, local-file reads and entity expansion. Removing a DOCTYPE may also remove declarations the document depended on, so do that only when you own the data and understand the consequences.

Takeaway: XML to JSON is a mapping, not a translation — and how the Syntax converters panel lets you inspect that mapping in your browser

Treat the result as ToolAcre’s documented mapping: `@` for attributes, `#text` for mixed element text, `#cdata` for CDATA, arrays after repeated siblings and literal namespace prefixes. Those rules make the output predictable without pretending XML and JSON share one data model.

For inspection, this projection is fast and readable. For durable integration, test one and many occurrences, attributes sharing names with children, empty elements, mixed content and namespaces. If element order or schema constraints matter, parse XML against that contract rather than depending on a generic converted shape.

Keep the raw XML fixture beside the normalized expectation. That pairing preserves evidence if a dependency upgrade changes array handling, whitespace trimming or entity decoding. It also gives reviewers a place to see distinctions the JSON view cannot carry. A converted object alone cannot prove whether an empty string came from a self-closing element, paired tags or another convention in the source.