Nink AI Engineering © Nink AI Engineering GmbH

Filed Under the Wrong Headline: A Licence, a Deferral, and a Benchmark

Three things landed between early July and early August 2026, and all three were reported under headlines that do not describe them. A licence change dated to July that actually shipped in May. An AI Act deadline that "did not happen" while a different one quietly closed. A benchmark whose headline finding the authors themselves tell you not to generalise. In each case the decision-relevant fact sits in the second paragraph. That is where reading stops.

SysML v2's Reference Implementation Is Becoming a Library You Embed

First, a correction that matters for dating artefacts. The relicensing of the SysML v2 pilot implementation from LGPL-3.0 to EPL-2.0 did not ship with the July bundle. It shipped in the 2026-04 releases: PR #752, "ST6RI-922 Change licensing to the EPL", 808 files changed, merged 13 May 2026, with the PR body recording that "All copyright holders agreed" (PR #752, LGPL-3.0 at tag 2026-03, current LICENSE). API & Services followed on 14 May (releases). The 2026-05 bundle is simply the first combined release to carry the change at the top of its notes. And note the trap underneath: release naming lags publication by roughly two months. Pilot 2026-05 published 2 July 2026 as Eclipse plugin 0.60.1; the combined SysML-v2-Release bundle followed on 22 July 2026 (2026-05).

The more consequential engineering story ran alongside it and got almost no coverage. 2026-04 added org.omg.sysml.model and its bundle, described in the notes as the first step in refactoring the pilot implementation so its code can be used "outside the framework of Eclipse plugins" (2026-04); 2026-05 continued the decoupling with org.omg.sysml.logic (2026-05). EPL-2.0 is also, not coincidentally, the licence that Eclipse SysON (Obeo + CEA) and Sensmetry's SysIDE already run on (SysON, SysIDE). Licence and packaging moved in the same two-month window, which is what makes "SysML v2 parsing as a library" realistic rather than aspirational: CI-side model validation, textual-notation linting in a build pipeline, a SysML v2 → ArchiMate bridge. I have argued for years that no single tool or meta-model does it all and that what matters is whether the tools already in the organisation integrate—an embeddable parser is exactly that kind of move.

What is in Beta 1 of KerML 1.1 / SysML 2.1 / Systems Modeling API 1.1 is less exciting, and that is the good news. The release notes publish no total; a count of the enumerated list gives 118 resolutions—31 KerML, 84 SysML Part 1, 1 SysML Part 2, 2 API—and the notes state the metamodel changes were "entirely to documentation comments and OCL", that the new metamodel is "structurally the same as the KerML 1.0/SysML 2.0 metamodel", and that the metamodel URI is deliberately held at version date 20250201 (2026-05). Your SysML 2.0 XMI and your existing tool bindings keep working. I have argued before that your tool's pilot release number is the real version number, not the spec version printed on the box; nothing in Beta 1 disturbs that. Two declared backward incompatibilities do deserve a targeted check: new validation that assign action usage referents are time-varying, and a revised derivation of Definition::ownedMetadata / Usage::nestedMetadata (SYSML21-419). Anyone running the pilot implementation as a validator in a pipeline should expect new findings after upgrading to 0.60.1.

The governance picture is honest and mixed. Three RTFs are running with roughly 680 open issues between them—about 498 on SysML v2.1, 112 on KerML 1.1, 68 on the API (task force index)—and issues are still arriving: SYSML21-658 was filed on 20 July, two days before the bundle shipped. Meanwhile the public OMG tracker shows every issue as open with a blank disposition, including ones the release notes say Ballots 1 and 2 resolved (spot-checked on SYSML21-540 and KERML11-144). GitHub, not issues.omg.org, is currently the only public window into what the RTF has decided.

One smaller thing worth knowing before you write a licence attestation for a client: the pilot README still points at a LICENSE-GPL file that no longer exists in the repository (README.adoc), and whether EPL-2.0's Secondary Licenses option is in play could not be established. Ask the maintainers rather than guess.

The AI Act Deadline That Moved, and the One That Shut

Any AI-governance roadmap baselined on 2 August 2026 is now partly wrong, and saying so is the story rather than an embarrassment. Three weeks ago Member States were obliged to have at least one national AI regulatory sandbox operational by that date, and the high-risk classification rules stood deferred to August 2027. Both changed. Regulation (EU) 2026/1744, the Digital Omnibus on AI, was published in the Official Journal on 24 July and entered into force on 27 July 2026—three days later instead of the standard twenty, deliberately, so the amended rules were settled before the AI Act's general application date (Commission, Hunton). Sandboxes now move to 2 August 2027. Annex III stand-alone high-risk obligations move to 2 December 2027, Annex I embedded-in-product to 2 August 2028 (Cooley, Orrick).

"The AI Act was delayed" is the wrong sentence to carry into a steering committee. The general application date did not move: it is still 2 August 2026, and Article 50 transparency applies now, backed by penalties of up to €15m or 3% of worldwide annual turnover (ComplianceHub). Nor is a further slip plannable. The Commission's original November 2025 proposal tied the deferral to a conditional trigger on harmonised-standards availability; the co-legislators replaced that with hard dates in the 7 May 2026 political agreement (Gibson Dunn). The escape hatch was removed. The Commission's Article 50 guidelines of 20 July add a detail worth noting in the current agentic enthusiasm: AI agents must identify their human operator (Mayer Brown).

The Omnibus also changed answers, not only dates, and this is the part that gets missed. The Machinery Regulation (EU) 2023/1230 moved from AI Act Annex I Section A to Section B, routing AI in machinery products through the sectoral regime instead, with a Commission delegated act on Machinery Annex III due by 2 August 2028. "Safety component" was narrowed in Art. 3(14) and Art. 6(1)—AI whose intended purpose is user assistance, performance optimisation, efficiency, automation, convenience or quality control is not a safety component unless failure could endanger health or safety. A new small mid-cap category (<750 employees, ≤€150m turnover) brings simplified documentation and fine caps (Orrick, Cooley). For a DACH Maschinenbau client, that first item is larger than the deferral. And Annex XIV now carries code AIH 0401 for "emerging AI technologies not covered by other codes, including agentic AI"—reportedly the first binding EU text to name it, with no definition or obligation attached (NicFab). An empty placeholder is still a hook for later.

Then the counter-intuitive one. Two sources report that the Omnibus left the Article 111(2) grandfathering cut-off untouched: high-risk systems lawfully placed on the market before 2 August 2026 stay out of scope absent significant design change (NicFab, Cooley, Nov 2025). If that reading holds, the deferral did not extend the grandfathering window—a system placed on 1 August 2026 is grandfathered, the same system placed on 3 August 2026 is fully in scope on 2 December 2027. The caveat is real: this could not be confirmed against the Official Journal text, and it needs counsel to verify against the consolidated wording before it drives anything.

For an architect, the practical shape is five workstreams rather than one slipped milestone, and the nearest real deadline is 2 December 2026: machine-readable marking of synthetic content for generative systems already on the market before 2 August, under the new Art. 111(4), plus two new Art. 5 prohibitions. Article 50(2) is not a UI requirement—it is a provenance requirement at the point of generation, which makes it a data-flow and component-boundary decision rather than a presentation-layer one. Classification should be re-run, not paused, because the amended text changes outcomes. And the temptation to skip levels is strong here: sixteen months of relief invites deferring the slow work.

One modelling approach is worth trying: in ArchiMate, AI systems as application components carrying a classification property and a provider/deployer role property, with AI Act articles as motivation-layer requirements traced to the components they constrain; the same trace works in SysML v2 as requirement-satisfy relationships. That produces a dated, re-runnable classification trace instead of a compliance binder.

No Prompt Strategy Wins Everywhere—and the Authors Say Don't Quote That

Stein et al., eight authors from the Institute for Engineering Design at TU Braunschweig, published the most systematic public comparison so far of how you ask versus what you get for SysML v2 generation: four models, five prompt strategies, three modelling tasks, 62 configurations on a traction-battery case study, open access on 23 July 2026 (DOI). The models were ChatGPT-5.4 Pro and Gemini 3.1 Pro through their web interfaces, Llama 4 Maverick and DeepSeek-v3 through APIs; the strategies were Zero-Shot, Few-Shot, Chain-of-Thought, Chain-of-Verification and RAG, plus a RAG-Upload condition using platform-native document upload that existed only on the two proprietary web products. Syntax validity was gated by actually importing the generated model into Cameo (Magic Systems of Systems Architect / Cameo Enterprise Architecture 2026x), and quality was scored by F1 against a reference model plus an LLM-as-a-Judge run on Claude Opus 4.6.

The numbers that made it into the prose are worth having. For requirements, ChatGPT and Gemini both reached an overall F1 of 0.80, ChatGPT climbing from 0.47 under Zero-Shot and Gemini from 0.45—the largest single effect in the paper, and it comes from feeding in the source documents rather than from clever phrasing. For state machines, Chain-of-Thought reached the highest average normalised score of 0.88 and was the only strategy where all four models cleared 0.75; ChatGPT averaged 0.95 on judge score, DeepSeek 0.85, Gemini 0.80, Llama 0.55. For block definition, Chain-of-Thought and Chain-of-Verification led and RAG scored lower (DOI). So: context beats cleverness for requirements, scaffolding beats context for structure and behaviour.

Does this tell you something real about prompt strategy for MBSE? The answer is a resounding yes, but it is not the whole story—and the honest part is that the authors got there first. Their own §5.5 says it: "Because each configuration was generated once, these findings describe the investigated generation sequences and should not be interpreted as general rankings of prompt strategies." Each cell was run once, with no variance reported and up to five syntax-repair rounds allowed before scoring, so a 0.05 gap between strategies is indistinguishable from resampling noise. The judge is a single un-calibrated model with no expert-agreement study behind it. The requirements headline is confounded, because "RAG Upload wins" is inseparable from "the two closed models on consumer web UIs win"—only they had the feature. And requirements were scored by F1 while the other two tasks were scored by a judge; three tasks on two incommensurable scales will tend to produce different winners. It is a good stick to hold next to a vendor demo, and a weak basis for a procurement decision.

Which is where it lands squarely on ground I have argued before. The authors' bottom line is that "LLMs should be used as human-in-the-loop modelling assistants rather than autonomous model generators. They can support early modelling and reduce initial effort, but expert validation remains necessary"—which is the copilot position, arrived at empirically. Autonomy and trust remain two different axes, and this paper is a measurement on the trust axis, not the autonomy one. The sharpest number in the whole area is not in this study at all: Al-Shami et al. report that semantic fault repair in SysML v2 runs at under 3% unaided, and rises above 91% with a fine-tuned small model plus a domain knowledge graph across 1,184 samples (arXiv:2606.23395). Compilers catch syntax; they do not catch domain-rule violations. Reliability should be earned before independence is granted, and on the semantic axis it has not been earned yet. Four questions to take into the next demo: which task, which metric and who wrote the reference model, how many runs and may I see the failures, and are those numbers before or after automated repair.

Worth a look