A publicly documented experiment
12 months, documented in the open. Started 2026-09-10.
Agents today live in the digital world. For one to be allowed to act in the real one, it has to master four further realities – the physical, the normative, the economic and the institutional. The Battery Reality Pack is the first case that shows how far that carries: a schema through which a foreign agent understands a real battery. It is the first case, not the goal.
The hypothesis
Computing what to do with a battery is a solved problem, sold by several vendors. Computing whether it may be done is not.
The constraints that decide it sit elsewhere: with the manufacturer, the grid operator, the regulator, the market and the owner. They are written for people, not machines. They bind a context – this metering point, this registration, this moment – rather than the object itself. And the sources disagree with each other.
The consequence is not that agents fail loudly. It is that they act inside an envelope their own builder hand‑coded, and nobody can tell from the outside which limits were checked and which were assumed.
The missing input can be supplied as a property of the object, not as knowledge inside the agent.
If a real asset is described across five layers – digital, physical, normative, economic, institutional – then an agent that knows nothing about batteries can be refused correctly when it proposes an action it may not execute, and the refusal names the layer and a source a human can check.
That is the whole claim. It is stated this way so that it can be false.
Such a description is a Reality Pack: one real object, five layers, machine‑readable. The one for a grid‑scale battery is Pack #001 – the Battery Reality Pack, and it is set out in full further down this page. Other packs are conceivable – a PV plant, a wind turbine, a data centre, a charging park – but none of them is written.
The observable is directly below: reasoned refusals. It stands at zero. Until it does not, everything above is a claim.
Status today
These four values are not maintained by hand. They come from the logbook: an entry may set them, the strip shows the most recent one for each. The status therefore cannot be changed without someone writing down why.
Milestones
Every criterion was written down on the day it was declared, before anything was measured. That is the point: a criterion formulated after the result can always be met.
The dates are a plan. The criterion is what decides: a milestone counts as reached when its criterion is met, never because its window has passed.
Deliverable An asset record for one real battery: data source, access path, and written permission to publish the data.
Criterion A real asset is named, with a data source and with permission to use its data. No invented data sheet, no example asset.
A pack without a concrete object is a schema without a reality. As long as this milestone is open, everything that follows is paper.
Criterion declared 10 Sept 2026
Deliverable Schema v0.1 and a runtime that answers one proposed action with yes, no-and-why, or escalate.
Criterion An agent that does not know this repository and was not built for it uses the schema to refuse an action it is not allowed to carry out — and names the layer that forbade it.
The measure of this project, literally. Not that an agent can do something, but that it reasonably does not. “Foreign” is the hard part: an agent you built yourself towards the right answer proves nothing.
Criterion declared 10 Sept 2026
Deliverable Every constraint carries a citation, plus a written check by a person who did not write the pack.
Criterion A human can verify the reasoning against its source — a citation rather than an assertion — and could refute it if it were wrong.
A refusal that cannot be refuted is not a limit, it is an opinion. This milestone is why the schema has to carry sources and not only values.
Criterion declared 10 Sept 2026
Deliverable One refusal record that names two layers and cites the source behind each of them.
Criterion A refusal that needs two layers and names both — for instance: the asset could, but it may not. Either layer alone would have let it through.
The layers on their own are occupied: Layer 1 is the core of every energy management system, Layer 3 the core of every trading system. If something of our own emerges here, it is the connection — not the layers. This milestone is the first at which that would show.
Criterion declared 10 Sept 2026
Deliverable A machine-readable mandate for one real asset: owner, every other party holding a write path, what the agent may do, its bounds, what it must escalate, and who can revoke it — plus one logged refusal that cites it.
Criterion A refusal that comes neither from physics, nor from law, nor from price: the action is executable, permitted and worthwhile — and the agent still may not carry it out. The refusal names the mandate that forbids it and who can lift it.
This is the one no existing system can produce. Layer 1 is what every energy management system already does; Layer 3 is what every trading system already does. A refusal on authority is unoccupied — and it is the difference between an agent that is careful and an agent that is allowed.
It is declared without a date on purpose. Every other milestone has a target window; this one has a predecessor instead. Putting a date on it would be a guess dressed as a plan.
Criterion declared 14 Sept 2026
Logbook
Newest first. Entries are appended, never edited; anything that turns out to be wrong is corrected by a later entry and still stays where it is. Failures appear here at the same size as progress – where it does not go forward is the more interesting part.
There is nothing to fill in and nothing to request. Anyone who wants to know how far this goes can read along.
Two layers of the five are not written: the digital one and the institutional one. Until today the page labelled them “the two layers the pack does not contain”, which reads as out of scope. It is not what was meant, and it contradicted the paragraph below it, which calls Layer 4 “the next stage, not a footnote”. The label now says still to be written.
Both are in scope. They are in different states, and the difference matters.
It was marked presupposed until yesterday because the original version addressed a direct marketer, who already has EMS access, schedule data and market data. For a private operator that inverts. The data connection is not a precondition, it is the first thing that becomes real — before the schema, before any runtime.
What is still missing is the description: where the data comes from, at what rate, with what latency, who else holds a write path, and how an agent tells a refusal from an unreachable source. That last one is not a detail. From the outside the two look identical: nothing happens.
On 2026-09-10 this log said the mandate question was trivial for a private setup — the battery belongs to the household, so there is only one principal. That is wrong. There are at least five, and every one of them is real at 5 kWh:
The structure is identical to the 20 MW case. Only the amounts are smaller: an agent acts on an asset, and the consequences land on a person who is not in the loop.
A consumer device keeps a write path for its manufacturer that the owner cannot cancel without losing functions.
The owner is not the only principal, and cannot revoke the other one.
On a grid-scale asset the same question would be settled in a contract. Here it stands bare. For testing the idea, that makes this the better case, not the weaker one.
Milestone 05 is declared: a refusal that comes neither from physics, nor from law, nor from price — the action is executable, permitted and worthwhile, and the agent still may not carry it out.
It is deliberately declared without a target date. Every other milestone has a window; this one has a predecessor instead. A date here would be a guess dressed as a plan.
Nothing is built. The status above is unchanged.
The project gets its own address: reality-pack.dev. Not
battery-reality-pack.energy, and not a subdomain of an existing one.
.energyIt would have contradicted a sentence in the first screen of this page: “It is the first case, not the goal.” An address that commits to energy says the opposite of that. Energy is where the layers happen to be hardest and best documented — it is not what the schema is about.
It would also have named Pack #001 only. A PV plant, a wind turbine, a data centre and a charging park are already named in the repository as conceivable further packs; each would have needed its own domain. And the renewal price is roughly five times that of the alternative, every year, for a project with no revenue that is meant to stay readable after the twelve months are over.
It would have been free and, today, perfectly honest. The argument against it is the measure of this project: a foreign agent. A schema that lives under a company’s domain reads as that company’s schema. For something meant to be consumed by people who were not involved in writing it, a neutral address is a small but real signal.
The obvious objection to buying an address for a schema that does not exist is that it is an advance on something unbuilt. It is not, and the distinction is worth keeping: the build rule binds what is delivered, not what is reserved. An unused domain makes no claim to anyone. A page that asserts things it cannot hold does.
The site configuration still carries no site value, so the delivered pages
emit neither a canonical URL nor an og:url. Checked today: the domain has no
DNS delegation and answers nothing. Pointing metadata at an address that does
not respond would be exactly the kind of claim this project keeps removing from
its own files.
That one line gets set when the domain actually answers. Deployment itself is a separate question and not decided here.
While there is only one pack, the page stays at the root. A /battery path gets
created when there is a second pack for it to be distinguished from — an index
page above a single entry is structure without content.
The status above is unchanged. An address is not a schema.
Going through the layers turned up three errors. All three sat in our own files, all three were found by reading, none ever produced an error message.
The README said: “a machine-readable representation of a real asset across five layers”, followed by a table with five rows. The page has always said: Layer 0 presupposed, Layer 4 roadmap, three are included.
Both texts were written on the same day. Nobody lied — the sentence had two jobs and did only one: it defines what a Reality Pack is in general, and directly below it says that this one is Pack #001. A reader concludes five.
That is exactly the promise the build rule forbids — stated in the same file, two sections further down. A promise may only say what the artefact holds. Corrected: the table now has a third column, and the sentence before it says three.
The digital layer was marked presupposed. Choosing a real device showed it is nothing of the sort: it is decided by the purchase. The same manufacturer’s range offers devices with an open local interface and devices without. With the latter the access belongs to someone else’s cloud, which can throttle or withdraw it.
And the consequence the word “presupposed” conceals: an agent that cannot tell a refusal from an outage does not have this layer. From outside the two look identical — nothing happens.
The state is now called procured. Not a nicer word, a truer one.
An entry dated 2026-09-10 said “Layer 5” twice. The stack counts from zero: 0 digital, 1 physical, 2 normative, 3 economic, 4 institutional. There is no fifth. Corrected.
Three errors, found by laying our own files side by side. None of them would have surfaced through a build, a test or a checksum, because none of them is a technical defect. All three are statements that had stopped being true.
That is the same class as the button which pointed at a deleted section two days ago, and the same as the duplicated status display in the footer: rules only bind going forward. Backwards, someone has to look. The build rule was in the README from the start — it did not prevent the promise two paragraphs above it.
The status above stays unchanged. A corrected sentence is not progress, it is the withdrawal of an advance.
The test object includes solar panels, up to 2,000 Wp — the ceiling of the simplified rules. The inverter is already inside the battery, which brings four MPPT trackers and accepts up to 5,000 Wp of solar input. So what is missing is panels and a mount, nothing else.
They are not part of this experiment because they generate electricity. They are part of it because without them two layers stay empty.
A battery on its own is fully known: state of charge, power limits, temperature — all measured, nothing forecast. There would be no uncertainty in the interface that could be declared.
That would remove the sentence the whole undertaking hangs on: an agent does not need the answer, it needs the declared boundary of the answer. With panels a quantity enters that the agent cannot command and has to predict. Only against that does it show whether the schema is more than a data sheet in JSON.
The battery alone would still be subject to registration: every stationary battery storage system must be entered in the Marktstammdatenregister — regardless of size, commissioning date and actual feed-in. So Layer 2 would not be zero.
But everything concrete hangs on generation: the 800 W feed-in limit, the 2,000 Wp module ceiling, the aggregation per metering point. That is precisely the rule set on which the idea of two devices failed — this project’s best evidence so far. Without panels it would not be there.
Whether a battery that charges exclusively from the grid and never feeds in fits cleanly into that rule set could not be settled. German renewables law defines a plug-in solar device as modules, inverter, connection cable and plug — a battery does not appear in the definition.
That is itself a Layer 2 finding: a real, purchasable configuration for which the rule has no clear category. Not “forbidden”, not “permitted”, but not provided for. An agent that depends on an unambiguous answer has none here — and must be able to say so instead of picking one.
| Configuration | The question put to the agent |
|---|---|
| Battery only, dynamic tariff | Is electricity cheap right now? One source, one decision. |
| With panels | Do I take the cheap hour now — or keep capacity free for tomorrow’s sun? |
Only then do two uncertain sources compete for the same storage. That is the question stated further down this page: which permissible action is worth most?
The status above stays unchanged. Still nothing bought, nothing built and nothing refused.
Decided: the test object will be an Anker SOLIX Solarbank 4 E5000 Pro (5 kWh, ~€1,600), together with a dynamic electricity tariff. It has not been bought, so milestone 1 stays open.
It has what matters for a foreign program: an official local Modbus TCP interface, read and write, enabled in the manufacturer’s app. No cloud detour, no reverse engineering. For home batteries, local write access is the exception, not the rule.
The dynamic tariff is not an add-on, it is the reason. It makes Layer 3 real: hourly exchange prices are a limit nobody in this house set. A refusal that goes back to a rule we wrote ourselves proves nothing.
What is not real about this asset belongs here just as much: Layer 2 is simulated. The device feeds in at 800 W and is therefore far below the 4.2 kW at which § 14a EnWG applies — it will never receive a control command from a grid operator. The rules it has to be measured against are ones we write ourselves. That is a weakness of the test object, not of the schema, and it stands here so that nobody later mistakes it for a success.
The idea was to stack two devices to get past 4.2 kW and make Layer 2 bite. It failed on three independent grounds, and the failure is the best evidence so far for this project’s thesis:
An agent that reads only the data sheet plans 2,500 W of feed-in and is wrong. An agent that adds the second box arrives at 5,000 W and is wrong not by a factor but categorically: the result is not a larger number, it is a different legal regime. The same hardware may do different things depending on how it is installed and registered. A schema that describes only the device cannot answer that question in principle.
Two observations for the schema design fell out along the way:
Noted, not decided. It is the only asset at household scale I know of where Layers 2 and 3 come from outside:
The price has to be named honestly: roughly €8,000 to €12,000 installed instead of €2,000, against a profit share of up to €100 a year. Financially that is not a business. It would be an expenditure for insight, and it only pays once the schema carries at the small asset at all. Before that it would be an advance.
The status above is not changed by this entry: no schema, no runtime, no users, zero reasoned refusals. A purchase decision is not a refusal.
The experiment starts today, and today none of it exists: no schema, no runtime, no users. That is why the four values above stand where they stand.
What does exist is a question, a frame of twelve months, and four milestones whose criteria are fixed today — before anything was measured, so that they cannot move later.
What happened today: the technical description of the Battery Reality Pack was moved here from the Agent Reality Stack. It stood there as a product page. This project sells nothing, so the commercial layer has been removed — pilot, annual licence, operating contract, channels, the booking prompt. The technical content stands unchanged below it.
Three promises stood on that page which it did not keep. Two are corrected: an unfilled placeholder for a monetary amount sat in the middle of the delivered text and is gone — not repeated here, so the string does not end up back in the text nobody wants it in. The “refusal demo” executes nothing and now says so itself: there are five hand-written examples. The third promise was that the page answered with 404 on the production host while being advertised. That is no longer a promise either: this page is deployed nowhere.
The beginning of an experiment is a beginning. Anyone reading back in twelve months should be able to see what was not there on this day.
The stack, applied
One agent, one asset, one day.
“Maximise today’s economic return from this asset.”
A digital agent can already read the market, forecast prices, and produce a trading strategy. That strategy is not yet an action it is allowed to take.
What information is available?
Rich. This part is solved.
Can the action actually be executed?
Half the strategies the model produced are physically impossible.
Is the action permissible?
Some of what remains is legal. Not all of it.
Which permissible action is worth most?
Only now is optimisation a well-posed problem.
Is this agent authorised to do it?
Nothing in the previous four layers answers this.
Only the intersection of these five realities defines the set of actions this agent can legitimately execute. Every layer it cannot represent is a layer where it is not deciding — it is assuming.
Built on the Agent Reality Stack
The reality layer between a trading agent and a grid‑scale battery. It says No before it gets expensive.
The pack
The customer’s agent – their own, a third party’s, or one we operate – asks the pack before every action: is this action, for this asset, right now, executable, permitted and sound in the Bilanzkreis? Layer 0 is not a fourth test – it is the ground the other three stand on, and the reason an answer can be trusted at all.
Do we know — and how old is what we know?
Not an assumption but a procurement: with an open local interface the access belongs to the operator, otherwise to someone else’s cloud, which can withdraw it.
SourceEMS/BMS exports, schedule data, market data — via exports in a pilot, live in operation
Can the asset do it right now?
SourceData sheet, EMS/BMS, operator settings
Is it allowed to?
SourceGrid operator, prequalification notices, rulebooks
Is it worth it, and what does deviating cost?
SourceTrading systems, Bilanzkreis settlement, exchange data
The pack is meant to answer Yes, No with a reason, or escalation to a human – every answer logged and traceable to the layer that produced it.
The five verdicts below are hand-written examples, not the output of a running system. There is no running system today. They show the shape a refusal has to have for a human to be able to check it.
The agent acts, the pack bounds. That is the difference from an optimiser: the pack does not replace the trading decision – it is what makes the decision permissible.
Origin
The Agent Reality Stack describes five layers an AI agent must master in order to act in the real world and not only in the digital one. The Battery Reality Pack instantiates four of them for one asset class.
The four it describes are set out above, each with its own state. The fifth is here. It is in scope and it is not written – it is milestone 05. Anything that promises reality must not stay silent about its own limits.
The layer still to be written
The open layer
The pack answers whether an action is permitted. It does not yet answer whether this agent may carry it out for this principal. That is exactly where the step from shadow operation to live operation fails today – and why Layer 4 is the next stage, not a footnote.
Evidence Every answer is logged and reasoned. That covers the evidence half of Reality + Authority + Evidence. The authority half arrives with Layer 4.