After I wrote up the cardiac digital twin in July, one of the questions that came back had nothing to do with pharmacology or MCP. It was about Git. The demo raises a metoprolol dose from 50 mg to 60 mg, and someone asked how that change would get reviewed if it lived in a team repository instead of on my laptop. Who looks at it, and what are they looking at?

I started typing an answer and realized my own repo dodges the question. CardiacDigitalTwin.slx is generated by create_cardiac_model.m and deliberately left out of Git, so there is never a binary in a pull request to argue about. Most teams that build in Simulink commit the .slx. When one changes, GitHub gives the reviewer a file name, a size, and “Binary file not shown.” Somebody still clicks approve. What did they just approve?

A diff nobody can read

If you spend your days in Simulink, you already know the model is the source. A gain value, a port on a subsystem, a data dictionary entry: those are the lines of code. In a text codebase the reviewer sees the line that changed. With a model, the reviewer either checks out both commits and opens them side by side in Simulink, which needs a licence and a free afternoon, or trusts the pull request description.

MathWorks has a good answer for part of this. Their Simulink Model Comparison for GitHub Pull Requests example runs the licensed Simulink Comparison Tool in GitHub Actions and publishes the official visdiff HTML report. If you have MATLAB on your runners and want the real visual comparison, use it. What I wanted was the layer around that report. Which models changed in this pull request, how much should I trust what the tool extracted, what else in the repository depends on them, and which one do I look at first?

Unzipping the slx felt like progress

An .slx file is a zip package with XML inside. So the first thing I did was the obvious thing: unzip the base and head versions and diff the XML.

It produced a diff. That turned out to be the problem. The XML inside an .slx is not a documented API, and nothing promises its shape stays the same between releases. A moved block and a changed gain both show up as edited attributes, and the diff has no opinion about which one changes behavior. Worse, a text diff always looks complete. Nothing in it says “this is the part I could not interpret.” A reviewer reading it would reasonably assume that if it is not in the diff, it did not change.

That last part is what shaped everything after. Whatever extracted the model had to declare how much of the model it actually saw, and the rest of the pipeline had to respect that declaration.

I got it wrong a second time in the canvas, too. The first version rendered every report it loaded as if it were ready for a decision, including reports where extraction was only partial. Version 0.3.0 reworked it to stop at the trust state first. Version 0.1.0 and version 0.3.0 shipped on the same day, September 29, which tells you how quickly I was finding things to fix.

What runs on the pull request

The project is samueltauil/simulink-model-diff, and the main piece is a GitHub Action. If you live in Simulink more than in GitHub: an Action is a step GitHub runs on one of its own machines every time a pull request opens or updates. You configure it with a YAML file in .github/workflows/, and the result shows up on the pull request as a check with a summary page.

The trigger only fires when a model file changes:

on:
  pull_request:
    paths:
      - "**/*.slx"
      - "**/*.mdl"
      - "**/*.model.json"

permissions:
  contents: read

What to notice: it is pull_request, never pull_request_target, and the only permission is reading the repository. A pull request from a fork runs with no secrets and no write access. Then the steps, trimmed of their names:

      - uses: actions/checkout@v7
        with:
          fetch-depth: 0
          persist-credentials: false

      - id: drift
        uses: samueltauil/simulink-model-diff@v0.5.0
        with:
          fail-on: error

      - if: always()
        uses: actions/upload-artifact@v7
        with:
          name: simulink-model-drift-${{ github.run_id }}
          path: build/model-drift

fetch-depth: 0 is there because the analyzer compares the exact base and head commits of the pull request, so it needs the full history instead of a shallow checkout. The Action itself requests no permissions and uploads nothing, so the workflow that calls it decides what gets stored. And the upload runs with if: always(), because the reports are kept even when analysis or policy fails. A blocked check that throws away its evidence is just a red X, and the reviewer has to go dig.

What comes out is a review plan. Every changed model gets a priority: blocked when analysis failed, evidence is incomplete, or policy failed; high for interface or functional drift; normal for other semantic drift; low when nothing semantic changed. The whole pull request rolls up into one review-status output of blocked, review-required, or clear. The artifact holds an aggregate model-drift-index.json, a Markdown summary, and per-model JSON, Markdown, and SVG reports, plus an optional SARIF file. SARIF is the format GitHub code scanning reads, so findings can show up as annotations if you choose to upload it in a separate job.

Repository impact was the last piece I added, in 0.5.0. A change to one model is rarely only about that model. The Action bundles a scanner built on MathWorks’ BSD-licensed data-explorer-core, which reads the saved package bytes and finds model references, linked dictionaries, and external data sources without starting MATLAB. From that I compute direct and transitive dependents, unresolved references, and cycles. The report labels all of it structural context. It tells you what might be affected. It does not prove anything about behavior.

Not starting MATLAB is a deliberate choice, and it is the one I would defend hardest to a Simulink team. Loading a model can run callbacks, scripts, custom code, and whatever the model references. On a pull request from someone outside the team, that means running their code on your runner. So the default path reads bytes and never executes anything. Licensed extraction, where MATLAB actually opens the model, belongs in a separate workflow that someone approves, on a runner you throw away afterwards.

Tests and other tools plug in the same way. If your workflow runs MATLAB tests and produces JUnit, or another analyzer produces SARIF, the Action imports both into the same review plan. Failed tests, errors, and required evidence that never showed up all block. Warnings keep the result at review-required.

Four words that decide the review

Every extraction ends in one of four states: complete, partial, unsupported, or failed. Only complete can support the conclusion that nothing drifted. The other three stay visible in every report, push the model into blocked, and cannot be upgraded by anything else in the pipeline. A passing test suite does not make a partial extraction complete.

A lot of my posts this year have been about healthcare, so the analogy I keep reaching for is a lab panel. A blank next to potassium means the test was not run. It does not mean potassium was normal, and no clinician would read it that way. Model drift needs the same rule. “No drift recorded” is a finding only when the extraction was complete. Otherwise it is a blank, and the review tool should print it as a blank.

The canvas is where the engineer decides

The Action decides whether the pull request passes policy. Whether the change is correct is still an engineering judgement, and that is the job of the second piece: a canvas extension for the GitHub Copilot app. A canvas is a panel the app opens next to the agent conversation, and the repository defines what goes in it. This one lives in .github/extensions/simulink-model-diff-canvas/. You copy that folder into your repository, and there is no build step and no node_modules.

To review a real pull request, you pull the artifact into your checkout:

gh run download <run-id> \
  --name simulink-model-drift-<run-id> \
  --dir build/model-drift

Then you open the repository in the Copilot app and ask: Open the Simulink Model Diff canvas for build/model-drift/model-drift-index.json.

A GitHub Copilot app session with the Simulink Model Diff canvas open beside the agent conversation, showing extraction trust, the changed model queue, evidence, and the merge assessment

The panel runs in the order I would want a colleague to walk me through a model change. Trust comes first, and if the extraction is partial the canvas says so before showing anything else. Then scope: the queue of changed models in priority order, plus dependents from the repository scan. Then evidence, with imported tests and quality findings ahead of a ledger of recorded before and after values that you can filter by functional, interface, or structural changes. Fields the report never populated are hidden instead of rendered as “unknown,” which sounds like a small thing until you are scanning a long ledger full of them.

It ends with a merge assessment: analysis unavailable, qualified extraction needed, policy gate failed, reviewer decision required, or no drift detected. The agent can drive the panel while you read, through three capabilities, load_report, select_model, and refresh. You can ask it to jump to the model a reviewer flagged, or to reload after a new analysis run, without losing your place.

The canvas also has hard limits. It does not open the .slx. It does not call the GitHub API, post a review, or change the check result, and it cannot approve anything. Branch protection and the Action’s policy result stay in charge.

My own model got blocked

The demo in the repo is the cardiac model again. The 32-second recording follows a real pull request that changes the dose parameter from 60 to 65. The Action job comes back blocked, the report upload still completes, and the same change opens in the Copilot app canvas with a partial trust state and a merge assessment of “qualified extraction needed.”

There is also a checked-in report at samples/cardiac-digital-twin/drift.json for the original 50 mg to 60 mg change, and it is marked partial too. The snapshots were built from the MATLAB builder script, not extracted by a licensed Simulink runtime, so that is all they can honestly claim.

I have mixed feelings about that demo. A green “no drift detected” would look much better in a 32-second video, and I thought about building a sample that produced one. But the first time my tool looked at my own model, it refused to call it clean, and it was right. That is the behavior I would want from it on a Friday afternoon when someone’s pull request is the last thing between the team and a release.

Why I would keep the boundary even if it costs the demo

This is not a replacement for the MathWorks comparison. visdiff is still the qualified visual comparison, and the project docs say plainly that a package-level result is not equivalent to it. The two fit together. A team with licensed infrastructure can run a MathWorks-backed extractor, emit complete evidence, and let this project carry it through policy, impact, SARIF, and the canvas.

The obvious counter-take, once Copilot is in the picture, is “just let the model read the diff and approve it.” I went the other way on purpose. The agent in this setup can navigate the evidence, filter it, and explain it. It cannot turn partial evidence into a green check, and it cannot click approve. When the model ends up in a car or an infusion pump, I would rather have an assistant that helps a human read carefully than one that reads on their behalf.

A review tool that cannot say “I don’t know” will eventually approve a change it never saw. The repo is at samueltauil/simulink-model-diff, the Action is on the Marketplace, and the end-to-end review guide walks the whole flow if you want to try it on your own models.