Developer tools · Syntax converters
XML's lineage from SGML: why it has attributes, namespaces and DTDs
· Background
xml data-formats security
XML's oddities (attributes versus elements, namespaces, DTDs, mixed content) make sense once you know it was designed as a simplified SGML for documents, not for data. This post traces that lineage and what it means when you convert XML to JSON.
Why does this data format have attributes? — a developer converting XML to JSON and meeting a distinction JSON never needed
JSON has one kind of object property; XML distinguishes attributes, child elements and text. ToolAcre maps that distinction with `@` attributes and ordinary child keys. The difference is visible when an element has both `id="7"` and an `<id>` child: both survive under separate properties.
This convention explains the resulting shape without claiming that attributes have one universal semantic role. XML authors choose them under their own schemas. The converter preserves the node category it can observe, not the business reason it was chosen.
SGML and the document tradition — ISO 8879, markup for publishing, and the idea of tagging text rather than encoding records
Mixed content, ordered children, attributes, comments and processing instructions are document-oriented features present in XML source. The implementation demonstrates their handling but does not contain evidence for SGML chronology, ISO publication history or publishing-industry motivations.
A source-bounded article therefore starts from executable behaviour. Text around child elements is lossy; comments and processing instructions are dropped; CDATA remains marked. Those facts explain why a document tree does not fit cleanly into a JSON value tree.
Document-oriented XML features visible in the implementation, without an SGML history claim
The parser validates well-formed XML and reports line and column. It ignores the declaration in the resulting value and preserves the five built-in entities and numeric character references as decoded text. It does not establish working-group goals or a ten-principle design account.
Historical claims need external editorial sources, which this task does not add. Omitting them is more accurate than inventing compliance or provenance from a dependency name.
The parser handles XML syntax; repository evidence does not establish the 1998 design history
Attributes become keys prefixed with `@`; elements retain their names. Text inside a structured element moves to `#text`, while a text-only element collapses to a string. This keeps structural categories apart as far as an object model permits.
Writing XML reverses the convention: `@id` becomes an attribute. Invalid attribute or element names are refused rather than sanitized. That decision keeps a malformed mapping from becoming plausible XML with silently altered names.
Attributes and elements remain distinct under ToolAcre’s @ convention
Namespace prefixes are retained verbatim in element and attribute names, including namespace declarations. They are not resolved, removed or rewritten. This preserves source spelling but is not namespace-aware type resolution.
Every DOCTYPE is rejected before fast-xml-parser receives the document. The mask avoids false matches inside comments and CDATA. This prevents external entity requests, local-file entity reads and entity-expansion attacks, with no option to bypass the refusal.
Namespace prefixes are retained literally and every DOCTYPE is refused
In `<p>before<b>bold</b>after</p>`, text placement matters. ToolAcre detects mixed content, joins fragments under `#text` and warns that their positions relative to children are lost. JSON object properties cannot reproduce an ordered sequence of alternating text and element nodes.
Repeated same-name children become arrays, but differently named siblings remain properties. Code requiring exact document order should use an XML node representation rather than treating the converted object as complete.
What this does not cover — XSLT, XPath and XQuery, the processing languages that grew around XML
XSLT, XPath and XQuery are not imported or exposed. Nor does the panel validate XSD, process DTD declarations or construct typed domain objects. It is a data projection with explicit security and fidelity boundaries.
A query or transformation language can preserve and navigate node order in ways this plain-value conversion cannot. Choose that tooling when the document features are part of the work rather than incidental packaging.
Takeaway: XML is a document format that learned to carry data — and how the Syntax converters panel shows what survives its translation into JSON
XML’s observable data model includes distinctions absent from JSON. ToolAcre marks attributes, CDATA and structured text, retains namespace prefixes, drops non-data nodes and refuses DOCTYPEs. Each choice is visible and tested.
Use the panel to learn what survives a projection, not as historical authority or a full XML processor. When ordering, schemas or namespace semantics matter, keep the original tree and use purpose-built XML tooling.
Security and fidelity also intersect at the DOCTYPE boundary. Refusing the declaration prevents entity processing, but deleting it from an arbitrary legacy document may change entity references or validation assumptions. If you own the source, replace required entities with explicit safe text and validate the resulting document. If you do not own it, use an approved XML workflow rather than weakening the refusal or presenting a partial conversion as the original content.