A JSON-first Python SDK for creating, reading, validating, and converting AutomationML CAEX 3.0 documents.
The SDK is based on the CAEX 3.0 XSD shape, but its primary developer experience
is Pydantic and AML JSON: readable Python attributes in code, canonical
AutomationML names in serialized JSON, and XML when an existing AML toolchain
requires .aml output.
Project status: Alpha. The core model, JSON/XML round trips, schema generation, and semantic validation are usable, but the public API may still change before the first stable release.
- Python 3.11 or newer
- Pydantic 2
Install the current development version from a local checkout:
python -m pip install -e .Install development and validation dependencies:
python -m pip install -e ".[dev]"Once the first package release is available, the intended installation command will be:
python -m pip install automationmlInstall the optional OPC UA conversion runtime when UANodeSet export is needed:
python -m pip install "automationml[opcua]"from automationml import CAEXFile
document = CAEXFile.from_aml_xml(aml_xml)
nodeset = document.to_opcua_nodeset()
nodeset_xml = document.to_opcua_nodeset_xml()
archival_nodeset = document.to_opcua_nodeset_xml(include_roundtrip=True)
# Source-preserving only for exports that explicitly embed an AML payload.
archival_recovery = CAEXFile.from_opcua_nodeset_xml(archival_nodeset)
# The normal NodeSet uses the Python semantic reverse profile.
semantic = CAEXFile.from_opcua_nodeset_xml(nodeset_xml)
# The semantic path can also be selected explicitly for an app export.
semantic = CAEXFile.from_opcua_nodeset_xml(
nodeset_xml,
prefer_embedded_source=False,
)
# Exercise and inspect both directions without source embedding.
roundtrip = document.round_trip_opcua(publication_date="2026-08-17")
recovered = roundtrip.assert_equivalent()
# The lifecycle may start from an OPC-UA-authored graph as well.
opcua_roundtrip = nodeset.round_trip_automationml(
publication_date="2026-08-17"
)
canonical_nodeset = opcua_roundtrip.assert_equivalent()The default forward mapper builds a typed Pydantic OPC UA graph and validates
the serialized UANodeSet without output repair. The strict profile covers file
metadata, instance trees, attributes and datatypes, all four class-library
kinds, inheritance/type relations, contextual roles and MappingObjects,
interfaces, directed InternalLinks, constraints, mirrors, facets, and real OPC
UA Part 19 dictionary-entry Objects. Unsupported or ambiguous features fail
explicitly. The pinned, patched AutomationML/OPC Foundation XSLT is retained as
a development comparison engine in evaluation/, not as SDK API: it scores
1/121 against the Python mapper's 121/121, so it is evidence rather than a
conversion path anyone should ship. The SDK reads NodeSets produced by that
stylesheet like any other supported graph. The
Python semantic reverse maps supported OPC UA graphs to canonical AML without
requiring an embedded source payload. This is a versioned AutomationML mapping
profile, not a claim to convert arbitrary third-party NodeSets.
The OPC UA authoring layer uses immutable Pydantic value objects and an explicit
mutable UANodeSetBuilder. UANodeSet.from_xml() and to_xml() allow the
supported graph to be inspected independently of either mapper. Forward and
reverse conversion share validated datatype and contextual role-reference
rules, including their documented canonical inverses. This prevents the two
directions from silently growing separate lookup tables.
The versioned mapper comparison test plan defines the independent oracles, 120-case minimum corpus, OPC-UA-origin cases, and publication gates used to compare the Python profile with the unmodified working-group XSLT. The companion strict-implicit mapping profile documents every induced and canonical reverse choice.
The frozen release corpus declares 145 cases, of which 121 are scored. The
current evidence is PY-STRICT 121/121 versus unmodified XSLT-RAW 1/121,
with 30/30 versus 1/30 critical cases. Inspect the
release manifest,
complete result matrix,
scale report, and
expert-review checklist.
Regenerate the release JSON, JUnit, and Markdown evidence with:
python tools/compare_opcua_mappers.py \
--manifest tests/opcua-comparison/release-manifest.json \
--output-dir build/opcua-release
python tools/benchmark_opcua_scale.py \
--repeats 1 \
--output-dir build/opcua-scaleBoth roundtrip results expose differences as frozen Pydantic records. Each
record has an RFC 6901-style JSON Pointer such as
/InstanceHierarchy/0/InternalElement/0/Attribute/0/Value, a change kind, and
the source/recovered values. semantically_equivalent is computed from this
evidence rather than stored as an independently writable Boolean.
from automationml.opcua_nodeset import UANodeSet
graph = UANodeSet.from_xml(nodeset_xml)
# Equivalent Python-first path, without an XML serialization step:
graph = document.to_opcua_nodeset()
file_node = graph.node("ns=1;s=CAEXFile")
properties = graph.outgoing(file_node.node_id, "HasProperty")- Python API:
snake_casefields and small builder helpers. - JSON serialization: canonical CAEX names such as
InstanceHierarchy,InternalElement,RefBaseSystemUnitPath, andSchemaVersion. - XML serialization: explicit adapter, namespace-aware, ordered according to the CAEX object model.
- Semantic validation: structured issues for unresolved class paths, duplicate IDs, attribute type references, role links, interface references, and internal link partners.
- XSD conformance: practical by default so existing JSON examples load, with
stricter checks available through
assert_caex_valid(strict_xsd=True).
from automationml import attribute, caex_file, instance_hierarchy, internal_element
doc = caex_file("plant.aml")
motor = internal_element(
"Motor",
id="ie-1",
ref_base_system_unit_path="ExampleSystemUnitClassLib/Motor",
attributes=[attribute("speed", 1500, unit="rpm", data_type="xs:double")],
role_paths=["ExampleRoleClassLib/Drive"],
)
doc.instance_hierarchies.append(
instance_hierarchy("Plant", internal_elements=[motor])
)
json_text = doc.to_aml_json()
xml_text = doc.to_aml_xml()
for issue in doc.caex_validation_issues():
print(issue.code, issue.path, issue.message)The JSON output stays compact and canonical:
{
"SchemaVersion": "3.0",
"FileName": "plant.aml",
"SourceDocumentInformation": [
{
"OriginName": "AutomationML Python SDK",
"OriginID": "automationml-python",
"OriginVersion": "0.1.0",
"LastWritingDateTime": "2026-06-28T00:00:00Z"
}
],
"InstanceHierarchy": [
{
"Name": "Plant",
"InternalElement": [
{
"ID": "ie-1",
"Name": "Motor",
"RefBaseSystemUnitPath": "ExampleSystemUnitClassLib/Motor",
"Attribute": [
{
"Name": "speed",
"Value": "1500",
"Unit": "rpm",
"AttributeDataType": "xs:double"
}
],
"RoleRequirements": [
{
"RefBaseRoleClassPath": "ExampleRoleClassLib/Drive"
}
]
}
]
}
]
}Beyond the basic model and serializers, the SDK now provides extension-safe
AdditionalInformation XML, immutable ID/path and reverse-reference queries,
SystemUnitClass instantiation, automatic CAEX 2.15 import, reversible change
sets and copy-on-write edit transactions, atomic policy-driven merging, and a
filesystem-confined external-reference resolver.
from automationml import CAEXFile, FileSystemResolver, diff_documents
imported = CAEXFile.import_aml_xml(source_xml)
document = imported.document
motor = document.query().find_by_path("Equipment/Motor")
instance = document.instantiate_system_unit_class("Equipment/Motor", name="M1")
edit = document.edit()
with edit as working:
working.instance_hierarchies[0].internal_elements.append(instance)
updated = edit.result.document
assert edit.result.inverse.apply(updated).to_aml_dict() == document.to_aml_dict()See the core-platform API guide for lookup semantics, migration diagnostics, conflict policies, transaction guarantees, and external reference security rules.
python -m automationml.schema --output schemas/automationml.caex3.schema.jsonThe generated schema is intentionally keyed by AML names. That makes it useful for non-Python producers and validators while preserving a clean Python SDK.
Use caex_validation_issues() when an application needs structured validation
results. Each issue includes severity, code, path, message, and optional
target / suggestion fields. Use reference_index() when authoring tools
need to resolve AML paths such as SystemUnitClassLib/Motor or
RoleClassLib/Drive.
Validation is intentionally non-destructive. Invalid or incomplete documents can still be loaded and inspected, while applications decide which issue severities should block their workflows.
Ready-to-upload AML JSON and AML XML examples live in
examples/validation-suite. The suite includes small, medium, and big positive
and negative cases, plus a manifest.json with expected SDK validation issue
counts. The SDK test suite reads these same files so the public examples stay
locked to the validator.
Three self-guided notebooks in learning/ introduce AutomationML through the
same public SDK APIs used by applications:
- Build a drive station progressively from concrete equipment to stable class semantics.
- Round-trip one validated model through AML JSON and AML XML.
- Explore SystemUnitClass inheritance, interface materialization, semantic warnings, repairs, and InternalLink partners.
Install the local learning environment and open the notebooks with:
python -m pip install -e ".[learning]"
jupyter lab learningpython -m pip install -e ".[dev]"
pytestBuild the source distribution and wheel:
python -m pip install build
python -m buildsrc/automationml/: public SDK models, builders, serialization, schema, and validation APIs.tests/: unit, round-trip, schema, and semantic validation tests.examples/validation-suite/: catalog of valid, invalid, and warning-focused AML JSON/XML documents.learning/: executable, self-guided AutomationML lessons.schemas/: generated AML JSON Schema artifacts.docs/: design and SDK documentation.tools/: repository maintenance and generation utilities.
Issues and pull requests are welcome. Before submitting a change, install the development dependencies and run the complete test suite:
pytestChanges to validation behavior should include a focused test and, when useful to SDK consumers, a matching example in the validation suite.
The optional xml extra is reserved for pydantic-xml integration work. The
current XML adapter is kept separate from the public model so XML concerns do
not dominate the JSON-first model design.
This project is licensed under the MIT License. See LICENSE.