This is the multi-page printable view of this section. .
Design Records
- 1: Pinned TLS Parameters and Handshake Compatibility
- 2: Bucket Configuration Replication: Source Times, Deletions, and Deterministic Convergence
- 3: An Unsigned Header Is Not Part of the Request
- 4: Durable IAM Revocations
- 5: Federated CopyObject: Preserve the Destination Contract
- 6: Multi-Pool Object Consistency
- 7: Object Lock Replication Ordering
- 8: SSE-C Replica Integrity
- 9: Replication Reliability: Delete Completion, MRF Visibility, and Resync Cancellation
- 10: Replicated Tag Ordering: Revision Timestamps, Tombstones, and Resurrection
- 11: SILO Server 20260903 Pre-release Review
- 12: Replica Metadata Normalization: What a Trusted Copy May Not Re-Inject
- 13: Request-Header Deadlines: Absolute Header Limits and Rolling Body Idle Timeouts
- 14: Conditional DELETE: One Condition, One Logical Object
- 15: DSN-Only Database Notifications: A Compatibility Boundary for #53
- 16: Preview Text, Never Execute It: SILO Console Text Preview PRD
- 17: Go 1.27 TLS Defaults and OIDC Discovery Failure Modes
- 18: No I/O Before Auth, No Privilege From Headers
- 19: One Endpoint, Two Privileges: Separating User and Group Status
- 20: Config Environment Files Are Not Shell Scripts
- 21: Two SSE-C Keys, One CopyObject Response
- 22: Why CompleteMultipartUpload Must Return ChecksumType: Review of PR #57
- 23: When the Total Is Unknown: Folder Download Progress
- 24: A Listing Must Not Drop a Null Version That Still Has Quorum
- 25: A ListObjects Shortcut Must Not Turn a Missing Bucket into an Empty One
- 26: Read-Only Checksum Audit and Reliable CLI Output
- 27: Optional Checksums, Mandatory Failure: Repairing UploadPart and UploadPartCopy Compatibility
- 28: BadDigest, InvalidRequest, and the CompleteMultipartUpload Checksum Contract
- 29: ListMultipartUploads: Implementation, Upgrade Contract and Design History
Design records capture the reasoning behind SILO maintenance decisions: the problem being solved, the compatibility boundary, rejected alternatives, implementation requirements, and the evidence required before release.
Read the date and status first: a proposal describes work awaiting implementation or evidence; merged on main establishes source inclusion; released links an actual tag or release record. Historical decisions and deferrals do not automatically describe current behavior. Prefer fixed commits, PRs, stable tests and release records; distinguish local checks, remote CI, artifacts and production deployment.
1 - Pinned TLS Parameters and Handshake Compatibility
Status, 2026-09-17. The Go TLS default repair
48e184652shipped in Server 20260916. The reporter of issue #154 retested on that release and reports the sameconnection reset by peeron OIDC discovery. No root cause is established for that deployment. #154 was closed on 2026-09-11 on the strength of the merge rather than a retest; that was wrong, and this record states what the repair covers, what it cannot cover, and what has to change.
Update, 2026-09-17. An upstream Go issue reports the same symptoms from a different product, and a second one carries a full diagnosis of the mechanism. Both are summarized in The known upstream instance. The practical consequence: the first question to ask an affected operator is what inspects that traffic, not what to set in SILO.
Evidence class. Go behavior below is read from the Go 1.27.1 standard library source and the official release notes. Handshake byte counts and branch reproductions come from the synthetic fixtures of the #154 investigation, recorded in Go 1.27 TLS and OIDC. No forensic capture from the affected deployment has been obtained.
In plain terms
A TLS connection opens with a ClientHello listing the algorithms the client supports. Many corporate networks put a device in the path — a WAF, a TLS inspection appliance, a load balancer — that parses that message. Some of those devices, on seeing an algorithm identifier they do not recognize, do not ignore it: they reset the connection.
Go 1.27 added three new identifiers to that message — the ML-DSA post-quantum
signature schemes — for a total of 12 extra bytes. Rebuilt on Go 1.27, SILO
says something slightly different on the wire than it did on Go 1.26. Something
in front of the reporter’s identity provider does not accept it and sends a TCP
reset. curl from the same container succeeds because curl uses OpenSSL, which
does not offer those identifiers.
Go anticipated this class of breakage and provides an escape hatch,
GODEBUG=tlsmlkem=0. That switch has a precondition: it only applies when the
application has not specified its own algorithm list. SILO carried an explicit
curve list inherited from upstream, and that list happened to contain a
post-quantum entry — so the switch had no effect for us. The 2026-09-09 repair
removes those explicit lists, which makes the switch work again.
The limitation is the whole story: the switch disables post-quantum key exchange (ML-KEM); it cannot disable post-quantum signatures (ML-DSA), and the 12 extra bytes are ML-DSA. Go exposes no switch for those. If a deployment is failing on the ML-DSA offer, the repair does nothing for it — which is consistent with the retest result.
What Go is actually solving
Post-quantum key exchange is a deadline, not a preference. The threat model
is harvest-now-decrypt-later: an adversary records encrypted traffic today and
decrypts it years later once a cryptographically relevant quantum computer
exists. Confidentiality therefore has to migrate before such a machine exists,
not after. NIST finalized ML-KEM (FIPS 203)
and ML-DSA (FIPS 204) in 2024, and
the deployed form is hybrid: X25519MLKEM768 runs classical X25519 and ML-KEM
together and combines both secrets, so a flaw in ML-KEM leaves the connection no
weaker than X25519 alone. Browsers, CDNs and SSH implementations have been
enabling hybrid key exchange by default since 2024; Go followed in 1.24.
Signatures are not on the same clock. A signature cannot be forged
retroactively — a quantum computer built in 2035 does not invalidate a handshake
signed in 2026. So ML-DSA in TLS 1.3 today is an advertisement of support, not
a live dependency. Go 1.27 added MLDSA44, MLDSA65 and MLDSA87 to the
default signature algorithm list, which is where the 12 bytes come from:
Those identifiers are defined for TLS 1.3 only. isDisabledSignatureAlgorithm
drops them when the configuration cannot reach TLS 1.3, which — as noted below —
is the only lever an application has over them.
The switch that was never meant to cover explicit lists
The change that actually broke us is narrower and, read in context, defensible. From the Go 1.27 release notes:
Post-quantum hybrid key exchanges can now be explicitly enabled in
Config.CurvePreferenceseven if thetlsmlkem=0ortlssecpmlkem=0GODEBUG options are used. Those options were always meant to only apply to the default set used whenConfig.CurvePreferencesis nil.
The standard library has documented the two paths as alternatives since Go 1.24:
In code it is one branch: an explicit list is used verbatim; only the empty case consults the defaults, and the GODEBUG lives in the default path.
The principle is that explicit configuration outranks an environment variable — otherwise an operator could silently override a security policy the developer wrote in code. Go 1.26 and earlier filtered explicit lists through the GODEBUG as well; SILO’s working behavior on Go 1.26 was therefore relying on something Go considers to have been a bug.
Worth noting for future upgrades: Go ties GODEBUG defaults to the main module’s
go directive, so a module that still declares an older version keeps the older
behavior until it moves. SILO’s go.mod declares go 1.27.1, which is an
explicit statement that we take the 1.27 behavior.
Where SILO’s debt was — and where it still is
The list in question was
{X25519MLKEM768, CurveP256, X25519, CurveP384, CurveP521}, inherited from
upstream and later extended with the post-quantum entry. Pinning it produced the
worst of both worlds: the deployment could not benefit from the library’s
evolving defaults, and could not use the library’s compatibility switch.
48e184652 removes that assignment at all eight Server TLS configuration points —
the general external HTTP transport, the replication transport, the
client-certificate cloud transport, the internode transport, both grid links,
etcd, and the inbound S3/Console listener — retires the TLSCurveIDs helper, and
adds wire-level regression tests across five outbound constructors, two peer TLS
versions and both GODEBUG settings. Certificate and hostname verification,
cipher-suite policy, proxy handling and HTTP/2 selection are unchanged, and no
post-failure downgrade or retry was introduced. The earlier OIDC-only candidate
was withdrawn because the same transport also serves identity plugins,
notification and Lambda reachability checks, audit webhooks and S3 cloud tiers.
The same debt remains one layer down. Cipher suites are still pinned:
crypto.TLSCiphers() and crypto.TLSCiphersBackwardCompatible() are assigned in
the outbound transports, LDAP, etcd, the grid links and the inbound listener. The
listener at least has MINIO_API_SECURE_CIPHERS to choose between the two sets;
no outbound path has any switch at all. The next time Go changes the suite
defaults, the same script runs again.
The general guidance this record adopts: set MinVersion and leave the rest to
the standard library. Where an incompatible endpoint genuinely requires a
narrower profile, expose it as product configuration that is visible in
mc admin config and in a support bundle — never as a value pinned in source.
Protocol ossification is an old problem: TLS 1.3 had to disguise itself as TLS
1.2 on the wire, and GREASE exists precisely to force middleboxes to tolerate
identifiers they do not know.
What is actually different on the wire
Measured against the same synthetic IdP fixture, with default settings:
| Build | ClientHello | ML-KEM (4588) | ML-DSA 0x0904-6 | User-Agent |
|---|---|---|---|---|
| 20260804, Go 1.26.5 — works | 1497 B | present | absent | MinIO (…) |
| 20260903, Go 1.27.1 — fails | 1509 B | present | present | Silo (…) |
| 20260916, Go 1.27.1 — repaired | 1509 B | present (default) | present | Silo (…) |
20260804 with tlsmlkem=0 |
275 B | absent | absent | MinIO (…) |
20260903 with tlsmlkem=0 |
1509 B | still present | present | Silo (…) |
Two readings matter and are easy to miss.
ML-KEM was already being offered by the working release. Unless the operator
had set tlsmlkem=0 before the upgrade, ML-KEM is not a variable between the
working and failing versions, and restoring the switch cannot by itself explain
or repair their failure.
Under default settings, exactly two things changed between the two releases:
the three ML-DSA identifiers, and the HTTP User-Agent, which the rebrand changed
from MinIO (…) to Silo (…). The second only matters if the reset happens
after the handshake completes.
Blast radius and why the symptom misleads
The pinned list covered both directions of every TLS path, but the consequences differ sharply:
| Path | Consequence of an incompatible ingress | Visibility |
|---|---|---|
| OIDC discovery / JWKS | Identity init blocks, Console never starts, the node serves nothing | Immediate, fatal |
| KMS/KES, LDAP(S) | Encryption or directory authentication unavailable. Both already used default curves, so the GODEBUG always worked for them | Immediate |
| Replication targets | Replication silently backs up; visible only in metrics | Easy to miss |
| Audit/notification webhooks, Lambda checks | Audit records lost, events undelivered | Easy to miss |
| S3 tiers / cloud backends | Transition failures, remote objects unreadable | Easy to miss |
| etcd, internode grid | Internal traffic, rarely crosses such a device | Low |
| Inbound S3/Console listener | Clients cannot connect; determined by the peer’s hello | Depends on client |
Only the identity path blocks startup, which is why it was reported first. The others degrade quietly, so a deployment can be affected without anyone filing an issue.
Which deployments are exposed at all:
| Situation | Why |
|---|---|
| An SSL-inspection NGFW, SASE, SWG or IDS/IPS sits on the outbound path | The reported case, and the one with a confirmed vendor defect |
| Egress through an enterprise VPN that inspects TLS | Same mechanism, applied to the whole egress |
| The environment allow-lists JA3/JA4 client fingerprints | The fingerprint changes whenever the handshake changes, so every toolchain upgrade can trip it, with the same silent reset |
| An old TLS terminator or load balancer in front of the endpoint | Intolerant of unknown extensions, or mishandles a hello split across segments |
| Migrating from an upstream MinIO release built with Go ≤ 1.22 | That handshake carried neither post-quantum key exchange nor ML-DSA — roughly 275 bytes against roughly 1509. This cohort makes the largest single jump and is the most exposed |
| Regulated environments that disallow post-quantum algorithms | They need the classical profile for policy reasons, not interoperability ones |
Deployments whose outbound paths carry no inspection device are unaffected, and should not set any of the compatibility options below.
Four properties combine to make the symptom nearly undiagnosable in the field:
identity initialization retries at randomized 0–3 second intervals; the discovery
and JWKS fetches have no total timeout, so startup can hang indefinitely;
/minio/health/live and /minio/health/ready both stay 200 while identity is
offline, so a Kubernetes readiness probe reports success (only
/minio/health/cluster returns 503 with X-Minio-Server-Status: iam-offline);
and the error text — read: connection reset by peer — does not say whether the
TLS handshake completed. A curl control test from the same image then succeeds,
because curl uses a different TLS stack and negotiates HTTP/2.
What 20260916 itself changed
Removing a pinned list does not restore the previous wire format; it adopts the current default, which is larger:
| Release | supported_groups actually offered |
|---|---|
| 20260903 and earlier (pinned) | X25519MLKEM768, X25519, P256, P384, P521 |
| 20260916 (Go defaults) | X25519MLKEM768, SecP256r1MLKEM768, SecP384r1MLKEM1024, X25519, P256, P384, P521 |
For an operator who sets nothing, 20260916 therefore offers two additional
post-quantum identifiers at every TLS endpoint, inbound and outbound. This is
harmless for almost every deployment, but an ingress that allow-lists group
identifiers could in principle be broken by 20260916 where 20260903 worked.
GODEBUG=tlssecpmlkem=0 disables exactly those two while retaining
X25519MLKEM768. For an operator who does set tlsmlkem=0, 20260916 produces a
smaller, more conservative handshake than any prior release — which is the point
of the repair.
Three mechanisms, one discriminating fact
The investigation reproduced three mechanisms that each produce the reported error text. The repair addresses one of them.
| Branch | Mechanism | Covered by 20260916 |
|---|---|---|
| A | The ingress rejects ML-KEM, and the operator had tlsmlkem=0 set before the upgrade |
Yes — this is what the repair restores |
| B | The ingress rejects the ML-DSA signature identifiers | No. Only a TLS 1.2 cap suppresses the offer, and no such setting is exposed |
| C | An HTTP-layer rule rejects the changed Silo User-Agent after a successful handshake |
No, and no TLS change is relevant to it |
A fourth observation explains why the curl control test misleads but not why the upgrade regressed: the discovery transport disables HTTP/2 and sends no ALPN, while curl negotiated h2. That difference exists in both the working and failing releases.
One fact separates A/B from C, and the fixture results separate B from A: whether the reset arrives immediately after the ClientHello or only after the GET is written. A single packet capture in the failing network position, or the ingress’s own reject reason for the same second, settles it. That evidence was never requested from the reporter — which is the process failure this record exists to fix.
The known upstream instance
Branch B is not hypothetical. Two upstream Go issues describe it:
- golang/go#81199, “add a GODEBUG
to disable advertising ML-DSA signature algorithms in the ClientHello (Go 1.27
regression against TLS-inspecting middleboxes)”, open since 2026-08-28.
The reported symptoms are identical to #154:
read: connection reset by peeron every TLS 1.3 connection through a TLS-inspecting firewall, while curl, OpenSSL 3.6 and Node.js 24 succeed from the same host, a Go 1.26 build of the same program succeeds, andMaxVersion: tls.VersionTLS12works. The product named there is Palo Alto Prisma Access with PAN-OS 10.2.10-h37, and the Go team’s reply notes that a patched version is available. - golang/go#79626, closed, carries
the diagnosis. A network administrator traced it with their own firewall team:
a Palo Alto IDS/IPS threat signature for OpenSSL CVE-2020-1967 — reported
there as PA threat ID 58033, last updated 2022-07-12 — inspects the
signature_algorithms_certextension and resets connections whose algorithm list it does not expect. Palo Alto shipped a content update disabling that signature on 2026-06-23.
CVE-2020-1967 was a null-pointer dereference reachable through a malformed
signature_algorithms_cert; the vendor’s detector for it has now been firing on
legitimate modern handshakes for years. This also explains the 12-byte delta
precisely: Go 1.27 adds three ML-DSA identifiers to both
signature_algorithms and signature_algorithms_cert, which is 3 × 2 bytes in
each of two extensions.
Go will not provide a switch. The security lead’s position on #81199 is that
they cannot GODEBUG every low-level ClientHello change, that GODEBUGs have been
reserved for changes that alter negotiated parameters, and that advertising a
signature algorithm is not supposed to change anything. The proposed tlsmldsa
setting (golang/go#81307) has not
landed. Chrome is also expected to start GREASEing signature_algorithms_cert,
which will break the remaining affected devices harder and faster.
Two consequences for SILO. First, no plan may depend on an upstream knob appearing: the explicit compatibility setting in the requirements below is mandatory, not a fallback. Second, the fastest diagnostic question for any affected operator is what inspects this traffic — a vendor content update may close the case in one step, and that is a real fix rather than a workaround.
The option space
Grouped by who has to act. Each row solves specific branches; before the discriminating evidence exists, any of them is a guess.
Ingress side — the only complete fix.
| Action | Branches | Cost |
|---|---|---|
| Update the middlebox — threat-signature content as well as firmware — so it tolerates unknown algorithm identifiers | A, B, C | Requires the device’s owner. For the identified product a content update already exists, so this can be the shortest path, not the longest |
Route around it: point MINIO_IDENTITY_OPENID_CONFIG_URL at an endpoint that does not traverse the device, or use split-horizon DNS. The discovery document’s issuer must not change |
A, B, C | Certificate hostname and issuer consistency must hold |
| Terminate TLS in a sidecar (stunnel, Envoy) that connects to the IdP itself | A, B, C | An extra component and its own trust chain |
Send the request through HTTPS_PROXY |
none | A CONNECT tunnel forwards the same ClientHello; only a proxy that terminates TLS changes anything |
Deployment side — available today.
| Action | Branches | Cost |
|---|---|---|
GODEBUG=tlsmlkem=0 |
A | Disables hybrid key exchange process-wide, including internode links and the inbound listener; certificate verification unaffected |
GODEBUG=tlssecpmlkem=0 |
the two groups added by 20260916 | Narrower; retains X25519MLKEM768 |
GODEBUG=fips140=on |
A and B | Not recommended. The FIPS allow-lists exclude ML-KEM and ML-DSA, but also replace cipher suites and curves and exclude Ed25519/X25519 for the entire process |
| Remain on 20260804 | A, B, C | Forfeits every subsequent security fix; emergency measure only |
Allow the Silo User-Agent at the ingress |
C | One rule change |
There is no way to disable ML-DSA today. Go provides no GODEBUG for
signature algorithms — the crypto/tls entries in internal/godebugs/table.go are
fips140ems, tlsmaxrsasize, tlsmlkem, tlssecpmlkem and tlssha1 — and
tls.Config has no signature-algorithm field. The only lever is the version
gate: the identifiers are TLS 1.3-only, so capping MaxVersion at TLS 1.2
suppresses them. SILO does not expose that anywhere. That is the gap.
Why the toolchain is not rolled back
Rebuilding on Go 1.26 would restore the exact handshake of the working release, and upstream reports confirm that downgrading works. It is still the wrong instrument, for five reasons.
- It pays for someone else’s defect with everyone’s toolchain. The device at fault has a vendor fix available.
- It expires. Go supports the two most recent releases. Once Go 1.28 ships, a 1.26 build runs on a standard library that no longer receives security fixes, and the same decision returns with worse options.
- Upstream is not going back. Go has declined to add a switch, and Chrome intends to GREASE the same extension. A rollback defers the problem rather than solving it.
- It is not a compiler swap. Server, Console, mcli and silo-pkg all declare
go 1.27.1; building the current source with Go 1.26.7 is rejected by the module requirement, so a downgrade is a coordinated change across four repositories plus whatever transitive modules already require 1.27. - It only covers one branch. If the reset is the HTTP-layer rule of branch C, the toolchain is irrelevant.
One adjacent idea does not work either: lowering the go directive in go.mod
while still compiling with 1.27. That mechanism only governs behaviors that have
a GODEBUG, and ML-DSA advertisement has none — it would change unrelated
defaults such as the macOS root-certificate behavior and leave the handshake
exactly as it is.
Rolling back the release image is a different decision and a legitimate one. An affected operator staying on 20260804 while their appliance is patched is sound emergency practice; the product freezing its compiler for the same reason is not.
Requirements taken from this
- Phase-labeled diagnostics on the discovery and JWKS fetches. Record, via
httptrace, connection established → handshake started → handshake completed (with negotiated version, suite and group) → request written → first response byte, and name the last phase reached in the returned error. No configuration, no protocol change, and it separates branches A, B and C from a log line instead of a packet capture. This is the highest-value item. - An explicit outbound TLS compatibility setting, covering at least the
identity provider: classical curves, and optionally a TLS 1.2 cap. It must be
opt-in, warn loudly at startup, and be visible in
mc admin configand support bundles.MINIO_API_SECURE_CIPHERSis the existing precedent. This is the only in-process remedy for branch B, and since Go has declined to add a signature-algorithm switch, it is mandatory rather than a fallback. - Startup robustness: a total deadline and context cancellation for the discovery and JWKS fetches, and a decision on whether readiness should reflect identity state, which currently it does not.
- A supported diagnostic subcommand derived from the investigation probe: handshake phases, negotiated parameters, and classical / TLS 1.2 / User-Agent controls, so an operator can localize this class of failure without us.
- Retire the remaining pinned cipher suites, or at minimum give outbound paths the switch the listener already has.
Two options are explicitly rejected. Shipping a lowered default — a
godebug directive in go.mod, or ENV GODEBUG=… in the image — silently
weakens post-quantum protection for every connection and still does not address
branch B. Automatic downgrade-and-retry after a reset hands an active
attacker a trivial way to force TLS 1.2 with classical curves by injecting an
RST; browsers removed exactly this kind of fallback for that reason. If it is
ever built, it belongs behind the explicit opt-in of requirement 2.
Release and communication gates
The repair’s compatibility requirement was documented — in the Server README’s TLS section, in Go 1.27 TLS and OIDC, and in the 20260916 release notes. Three things were not done, and each contributed directly to the outcome:
- The reporter was never told, in the issue, that the repair requires an operator-set GODEBUG to take effect, nor which branch it addresses.
- The requirement never entered an upgrade checklist. It was filed as a record of what changed rather than as an action the operator must take.
- The issue was closed on the merge rather than on a retest by the reporter.
The underlying misclassification is the lesson: this was treated as a correctness repair (restore the library default) when it is also a compatibility contract change — its effect depends on an operator action. Classified correctly, it would have gone through the release-gate and user-communication paths rather than the documentation path alone.
Three gates follow, and apply to every future change of this shape:
- Any repair whose effect depends on an operator action is listed as a required action in the release notes and in the upgrade checklist, with the same treatment as a coordinated-upgrade item.
- An issue reported by an external user is closed only after that reporter confirms a retest. A merged patch changes status labels, not the issue state.
- When a repair does not fully resolve the reported symptom, the issue states which branch was addressed, which branches remain, and exactly what evidence is needed — rather than relying on the reporter finding a README section.
Attribution
This record extends the #154 investigation and the September Go 1.27 stack review with the post-release retest result, the mechanism analysis, the option space and the process gates. Reproduction artifacts and the full evidence chain are retained outside the documentation tree. The supportable statement remains: the merged repair restores Go key-exchange defaults at the eight affected Server configuration points, and that is verified with synthetic negative controls. It diagnoses no specific deployment, and #154 requires a retest with phase-level evidence before any root cause is claimed.
2 - Bucket Configuration Replication: Source Times, Deletions, and Deterministic Convergence
Publication update, 2026-09-17: The #77 repair recorded here shipped in Server 20260916; deletion-record export stays off by default. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
#77 is a reproduced site-replication correctness defect. A receiver replaces source time with arrival time and may then reject a genuinely newer deletion. Some configuration types stop exporting their timestamp after deletion, preventing heal from recovering a delete missed during an outage. Adding a DELETE branch alone cannot solve both problems.
Merge status (2026-09-12): the repair and research archive were merged into
mainthrough PR #180, commit 48ec10312. #77 is closed. All nine pre-merge CI checks passed.
Review boundary: the plan went through four Claude Code Opus 5 Max review rounds, passing the final two. Full implementation review, remediation review, and focused final acceptance returnedGO_WITH_NONBLOCKING_NOTESin all three rounds. Final blockers are zero; the requested full cmd and final lint checks have now passed.
Applicability: this page describes the repair now included in main. Check the specific version of downloads and running installations; full deletion recovery still requires every node to be upgraded and deletion export to be enabled consistently.
Existing work and scope
The earlier release notes, security hardening record, and Server compatibility page already document the #77 deletion limitation. They do not provide a complete record of its state model, alternatives, or verification boundaries. This page supplies that reasoning.
The source repository archive retains the original reproduction, plan versions, final reports and invocation identities for all seven review rounds, finding dispositions, executed test logs, source/binary hashes, and a rerunnable two-site driver. Raw model reasoning streams, binaries, and temporary lab volumes are excluded. Original artifacts and copies with normalized workstation paths and document links have separate hashes.
| Existing work | What it repaired | What it does not establish |
|---|---|---|
| #91 | Per-site configuration counts, Policy/Quota reporting, malformed-field isolation | Accurate counts do not prove convergence of values and source times |
| #103 | Serialized writes to the whole .metadata.bin record |
A lock cannot validate an ordering decision made before acquiring it |
| #76, #78 | Object Lock’s replication payload and existing-bucket adoption protection | They do not replace a consistent source-state comparison for six types |
| CORS replication repair | A separate per-bucket CORS deletion register and replication trust boundary | Its semantics cannot be applied blindly to every metadata type |
This repair covers Policy, Tags, SSE, Quota, Versioning, and Object Lock. Lifecycle/expiry has its own merged-payload time semantics; CORS keeps its separate mechanism. Notification, IAM, object replication, MRF, resync, and public counters are not rewritten. Object replication reliability has a separate record.
The supported release target is the maintained silo, silo-console, mc, and silo-pkg stack. Compatibility with unmodified upstream MinIO/MC remains best effort and does not require downgrading maintained components or recreating dependency forks.
What the reproduction established
The original regression reproduced on both ErasureSD and 16-disk Erasure ObjectLayers. A peer PUT originated at a past time T, but the receiving disk stored current arrival time. A subsequent DELETE carrying T+1 minute was rejected as older. RPC success alone therefore cannot establish correct final state.
| Configuration | Original PUT preserves source time | Established empty-event behavior | Original timestamp export without payload |
|---|---|---|---|
| Policy | No | Delete | Yes |
| Tags | No | Delete | No |
| SSE | No | Delete | No |
| Quota | No | Delete; zero-value JSON has separate semantics | No |
| Versioning | No | No operation | Not used as deletion |
| Object Lock | No | No operation | Not used as deletion |
Three further entry-point defects matter: an old bulk event can overwrite a newer field; a request that checks state before queuing for the lock can act on an obsolete decision; and remote Tag heal omits UpdatedAt. Each requires a repair at the actual entry point, beyond changing source selection in heal.
The facts a field needs
Reuse the existing payload, field UpdatedAt, and bucket Created. No disk format, SDK field, or persisted deployment ID is added.
| State | Conditions | Eligible source |
|---|---|---|
| Unknown or invalid | Unknown creation time, invalid payload, or field time before creation | No; missing information cannot mean deletion |
| Empty baseline | Empty payload at zero time or Created | No |
| Live baseline | Valid nonempty payload at zero time or Created | Yes, for historical configuration initialization |
| Real update | Valid nonempty payload later than Created | Yes |
| Real deletion | Empty payload for a deletable field, later than Created | Yes; this timestamped empty value is a tombstone |
Empty Versioning and Object Lock events remain no-ops. Sharing a helper must not give these types deletion semantics. Zero field times use Created as a comparison baseline; this does not turn historical emptiness into a new delete.
Ordering first gives real states priority over baselines, then compares source time between real states. At the same time, a deletion wins over a live value. Conflicting live values at the same rank use a stable content key, with the lexicographically greater key winning. For timestamped peer events and heal, an identical effective state causes neither a save nor a notification. A local write allocates a new monotonic revision even when the visible payload is unchanged. Live baselines also use content-key ordering; a later creation default cannot outrank a real change.
This is deterministic conflict resolution. It does not mean a lexicographically greater configuration better expresses business intent. Operators must still choose and resubmit the intended value after conflicting concurrent changes.
Comparison must match persistence
Policy sets are map-backed, so ordinary JSON encoding can depend on iteration order. The Server sorts the complete JSON tree of an already validated policy, including Statement, Action/NotAction, Resource/NotResource, Principal, and Condition. Sid and numeric precision are retained. This adds no syntax unsupported by the existing parser.
That parser accepts NotAction and NotResource, but the original structure’s required Action/Resource encoding can fail on the corresponding empty sets. Explicit field encoding addresses this, and Policy GET/admin export use it too. The purpose is to keep an accepted policy readable, beyond making its comparison key deterministic. GET/export now sort Statement arrays that formerly followed stored order, and sort set arrays whose map-backed order could vary between calls. Authorization is unchanged; upgrading does not rewrite every stored policy.
Quota keys use the existing parsed representation encoded as JSON. {}, JSON null, and valid zero-quota documents remain live documents, not implicit deletion events. Older senders converted a zero-quota PUT to deletion for peers while keeping a local document. mcli quota clear sends a zero quota; upgraded peers retain that live zero document, with no capacity enforcement either way. An empty Policy instead follows the established peer deletion interpretation: a valid empty-policy PUT succeeds, and a subsequent GET returns the existing NotFound response.
XML configurations use the bytes of a valid document. There is no general XML canonicalization layer. Versioning needs one exception: apply the existing Object Lock constraint before comparing the effective document that will actually be persisted. Otherwise comparison can accept a value that Save rewrites, causing heal to send it again next time.
Lock the decision as well as the save
The six fields share one .metadata.bin record. Loading raw state, validation, comparison, modification, and saving must all happen under the existing metadata.lock. Comparing outside the lock still permits stale decisions; separate field locks would allow whole-record read/modify/write operations to overwrite one another.
A local write allocates its time inside the lock:
This lets a local correction advance beyond an already stored future field time. Timestamped peer events retain their original time. Identical or older timestamped peer states return without a write; this is not a no-op guarantee for local PUT/Delete.
Bulk processing handles only explicitly supplied fields, validates them, and saves at most once. Omission preserves a field; explicit null follows its type’s semantics. An invalid field cannot leave half the bulk update persisted. Import assigns a common time for the selected six-type fields under the final per-bucket commit lock. Disk state and outbound events use that commit’s final snapshot, not a later read combined with an earlier time. A normalized empty Policy uses the existing dedicated deletion event so bulk omitempty cannot lose it.
The save helper returns its normalized snapshot. Public Update/Delete signatures stay unchanged, while other metadata retains its existing processing and notification behavior.
Adoption must not manufacture a deletion
Adoption can change Created. Moving it earlier while retaining an empty field’s old default timestamp makes an empty baseline appear to be a real deletion. Moving it later can invalidate historical initial values.
The narrow repair, under the existing adoption lock, rebases only these six fields whose previous time was zero or equal to old Created. Actual update and deletion times remain unchanged. An empty payload alone is not proof of a default, and existing-bucket configuration protection is not redesigned. Genuinely different bucket generations still require operator intervention.
One rule across the real entry points
| Entry point | Required behavior |
|---|---|
| Local S3/Admin writes | Monotonic time under the lock; use the committed snapshot for outbound events where needed |
| Typed peer events | Compare and persist original source time under one lock; retain existing types and legacy Object Lock payload compatibility |
| Bulk/import | Explicit field presence and atomic save; allocate import time at final commit |
| Initial synchronization | Preserve historical live baselines; include real deletions after opt-in |
| Local/remote heal | Use the same comparison for selection and apply, with complete source times; unknown IDs and failed peers do not block healthy targets |
| Status export | Value and time belong to one record; gate newly exposed deletion times during rollout |
Initial synchronization retains its existing five-type send path. Versioning is still initialized through MakeBucketHook and aligned through heal. No extra initialization path is added merely to make the table symmetric.
Heal filters valid candidates before selecting the maximum, instead of seeding from the first map entry and only then filtering defaults. Public mismatch counters cannot be the sole gate: equal content with different source times still needs synchronization. Conversely, equal effective states need no further write or RPC.
Why rollout needs a default-off switch
The startup setting MINIO_SITE_REPLICATION_METADATA_TOMBSTONES defaults to off. It controls visibility of newly exposed deletion information and does not detect remote capability.
| Behavior | off | on |
|---|---|---|
| Source-time ordering and atomic apply | Active | Active |
| Ordinary deletion events | Still replicated | Still replicated |
| Existing Policy deletion-time export | Preserved | Preserved |
| Real deletion-time export for absent Tags/SSE/Quota | Hidden | Exported |
| Additional real deletions in initial synchronization | Existing behavior | Include all four deletable types |
Old implementations cannot safely consume all newly exposed deletion information. For example, old Quota heal can clear the payload while retaining an already parsed cache value. An instruction to upgrade does not itself isolate this rolling-upgrade window, so the default stays off.
Upgrade every node at every participating site to a build containing the repair, ensure consistent settings within each site, and drain old requests. Then set on consistently and restart. Before a downgrade, first set off and restart every fixed node, then roll back the software. The old software’s original defects return with it.
While off, hidden Tags/SSE/Quota tombstones can cause repeated stale heal RPCs that a fixed receiver rejects. A subsequent heal with zero RPCs is expected only when complete state is visible and stable.
Retained and rejected alternatives
| Decision | Reason |
|---|---|
| Retain one internal six-field helper | The same source-time defect was reproduced across all six; shared ordering prevents entry-point drift while preserving type-specific deletion behavior |
| Keep the whole-bucket lock, persistence fields, and heal interval | They provide atomicity, durable deletion state, and missed-event recovery without another coordination service |
| Do more than replace UTCNow with source time | That alone leaves stale lock-external decisions, invisible deletes, equal-time conflicts, and initialization gaps |
| Do more than export tombstone times | Old receivers and Quota cache handling remain unsafe without controlling the rollout window |
| Do not use deployment ID as a tie-breaker or add an HLC/schema | Existing time and content keys suffice for the bounded contract; cross-site causal ordering is not claimed |
| Do not reject all zero-time typed events | Old Tag heal really omits time; preserve inexpensive protocol compatibility with an explicit limitation |
| Do not impose one deletion rule on all metadata | That would delete Versioning/Object Lock or violate separate Lifecycle/CORS rules |
The initial implementation adds 729 and removes 692 production Go lines, a net increase of 37, chiefly replacing duplicated apply and heal branches. Line count does not establish minimality. Necessity must connect each mechanism to a concrete failure; sufficiency must cover every real entry point; minimality asks which failure returns if a mechanism is removed.
Validation and its limits
The environment is local go1.27.1 darwin/arm64. Four groups of pre-implementation audit tests failed on the unfixed baseline. The resulting regression suite passes on both ObjectLayers; its main entry points are in cmd/site-replication-metadata{,-heal,-gate}_test.go. The baseline audit log retains the original failures alongside the passing existing tests.
| Validation | Observation |
|---|---|
| Source time, four deletions, duplicate/reordered events, queuing before the lock, different-field writes | Regressions and targeted race checks pass |
| Both equal-time arrival orders, deletion priority, Policy keys and negative-set GET, bulk/import | Boundary regressions and supplemental race checks pass |
| Full cmd package and internal/S3 Select race | Full cmd passes on final production code fcbb93e89 (492.776 seconds); internal/S3 Select race passed at 62cf066ff |
| Build, vet, lint, generated files, compatibility checks | Final production build/vet pass; lint reports zero issues at 461e9a721. Generated-file and compatibility checks passed at 62cf066ff. Optional typos is unavailable and skipped according to the Makefile |
| Linux/Darwin/Windows × amd64/arm64 | Six cross-compiles pass at 62cf066ff; this is not runtime acceptance on six platforms |
| Two real site processes, four data directories per site | Initial synchronization preserves Created for six historical live configurations; normal 30-second heal restores consistency after dropping a real PUT’s outbound RPC and injecting reordered events |
| Missed deletions and restart | Four deletion states persist across source-process restart and converge after reconnection |
| Quiescence and diagnostics | Two separate 65-second observations contain no metadata RPC across two normal heal cycles; repeated exceptional events deduplicate by bucket/field/reason |
| Fixed and pinned old implementation together | Tags PUT/DELETE smoke passes with all switches off; this does not prove complete mixed-version correctness |
| Review repairs: creation-time recovery, policy status key, heal diagnostics | Real ObjectLayer creation-time and legacy-order policy tests fail with old production code overlaid and pass after repair. Full cmd, internal packages, S3 Select race, lint, generated files, branding checks, and six cross-compiles were rerun at 62cf066ff |
| Final cleanup regressions | Targeted diagnostic, initial-sync, physical-time boundary, adoption, and CORS race tests pass at fcbb93e89; related race tests pass again after test-style changes in 461e9a721 |
| Two-site acceptance repeated on the final binary | The same run passes on binaries built from clean 62cf066ff and clean fcbb93e89; final 461e9a721 only changes test style and has identical production code |
Local evidence contains baseline failures, test logs, a rerunnable two-site driver, metadata snapshots, and source/binary SHA-256 identities. The first two-site binary reports 01aaef2b5 + dirty; its actual Go build-info and SHA-256 have now been recorded. After review repairs and cleanup, runs were repeated on binaries built from clean 62cf066ff and fcbb93e89. Actual --version, Go build metadata, and SHA-256 identities are retained with the baseline binary built from 5c5765816. Final 461e9a721 differs from fcbb93e89 only in test formatting and equivalent conditional syntax; that diff is recorded separately.
The table above records local tests and isolated-process observations. Remote integration evidence is separate: all nine checks on PR #180 passed, comprising six Go CI jobs, DCO, vulnerability analysis, and release-pipeline validation. The merged main tree was verified to match the PR merge tree that passed these checks. This is not proof of real Linux multi-node cluster behavior, official release artifacts, or production deployment.
Adversarial review record
The plan used actual Claude Code claude-opus-5 --effort max for four rounds. The first two drove corrections to state comparison, committed snapshots, historical baselines, and import boundaries. The last two returned GO_WITH_NONBLOCKING_NOTES with zero pre-implementation blockers. Plan approval is not proof that implementation is correct.
The separate implementation review pinned 1ee64a8d8 and used the same model and effort to inspect the complete diff, production call chains, formal tests, and runtime evidence independently, challenging sufficiency, minimality, rollout safety, and evidence identity. Its verdict was GO_WITH_NONBLOCKING_NOTES: no unconditional blocker, one conditional blocker, and nine further findings. It confirmed the core convergence mechanism and found no counterexample to ordering, deletion, or duplicate suppression within the declared contract. Three findings were real defects the change had introduced and were repaired in 62cf066ff. The subsequent fcbb93e89 closes the diagnostic gap when no valid source exists and adds an initial-sync regression.
| Finding | Assessment | Disposition |
|---|---|---|
| F1: a bucket with no recorded creation time can no longer write any of the six configurations | Confirmed regression, conditionally blocking | Repaired. GetBucketInfo returns the physical probe unchanged for a metadata-free request, as ListBuckets already did; initial synchronization recovers the time and passes it to the bucket creation hook |
| F2: replication status compares statement order while heal compares the canonical key | Confirmed; a permanent false mismatch that heal can never resolve | Repaired. Status compares the same key heal compares; per-site presence counting is unchanged |
| F3: one log key for four heal conditions, at error level for a normal transient | Confirmed | Repaired. Each reason keeps its own key at warning level. Empty baselines and missing buckets stay quiet; invalid existing state is still diagnosed when no valid source exists |
| F4: three user-visible semantic changes not written down | Partly valid | The original reviewed README already documented empty Policy and zero Quota at its end; the first review missed them. The addition covers Policy GET/export ordering and negative-set behavior |
| F5: whether the Policy encoder has unnecessary callers | The second review corrected the first assessment | Comparison and status keys must agree. GET/export/peer need the encoder to read or replicate negative-set policies that can already be persisted. PUT/import normalization is optional for comparison, but removing it adds branches and representation differences, so it stays |
| F6: an orphaned helper and a stale ordering comment | Confirmed nits | Comment corrected. The unused isBucketMetadataEqual and the obsolete test of that helper have been removed |
| F7: with the switch off, missed Tags/SSE/Quota deletions do not converge | A correct reading of the plan’s trade-off | No change; stated in rollout and in the limits below |
| F8: evidence gaps - a stubbed recovery test, no legacy-order policy case, unverifiable binary identity | Confirmed | Recovery now runs on the real ObjectLayer; a permuted legacy policy case was added; the two-site acceptance was repeated on a binary built from the clean final tree, with identities recorded |
| F9: bucket generation conflicts | Declared out of scope, and not a regression | No change; see the limits below |
| F10: adoption moving Created later than a real field time | A coverage gap, not a defect | The case now fixes both sides: preserve historical time, exclude earlier-generation state as a source, and accept a valid adopted-generation input as a target |
A stub is worth calling out separately. The original recovery test injected an object layer whose creation probe returned the expected time, so it passed against code that could never behave that way in production. The replacement stamps the bucket directory on every local drive and drives the real object layer, and it fails on the unrepaired code.
The second review pinned 62cf066ff, again using actual claude-opus-5 --effort max. It returned GO_WITH_NONBLOCKING_NOTES with zero conditional or unconditional blockers. It retraced production paths, checked the author’s F1/F2 reproduction results from formal tests with old production code overlaid, and revised the first review’s assessment of the Policy encoder. The reviewer did not execute the tests.
| Follow-up finding | Final disposition |
|---|---|
| NB-1: physical Created is an approximation | Retain the generation boundary and document directory-mtime limits. A real-drive test fixes the behavior: an earlier peer event is skipped and a local correction succeeds. One source timestamp cannot lower bucket identity |
| NB-2 and NB-8: silence without a source; lost source time in recovery-error logs | Repaired. Invalid existing states remain visible and recovery failures retain the event time. No source means no RPC; empty baselines stay quiet |
| NB-3: outage logs multiply by bucket and field | Accept and document the existing bucket/field/reason granularity; a site-wide log aggregation framework is outside this correctness repair |
| NB-4: draft logs are not product-failure evidence | Mark findings-before.log and findings-after-1.log as superseded fixture failures. Formal tests with old production code overlaid reproduce the physical-time, policy-order, initial-sync, and diagnostic defects separately |
| NB-5 and NB-6: unused helper; incomplete recovery-path description | Remove the helper and both tests that only exercised it. Document initial sync’s persisted recovery and retain real CORS-path tests |
| NB-7: adoption covered only the source side | Add the target side: valid adopted-generation input replaces invalid earlier-generation state |
Diagnostic regressions now use the existing logger target to check warning level, distinct reasons, repeated calls, and zero RPCs. The initial-sync test drives a real source ObjectLayer and the complete outbound sequence; its peer only acknowledges RPCs. That proves outbound content. The real two-process experiment supplies a separate level of evidence, and the two must not be conflated.
The third focused acceptance pinned final production code fcbb93e89 and again returned GO_WITH_NONBLOCKING_NOTES, with zero blockers. It also inspected the test-only style diff in 461e9a721 and confirmed equivalent semantics. Full cmd and lint were still running when the reviewer read their logs; both subsequently exited with code 0. Test formatting caused lint to fail at fcbb93e89 itself; corrected 461e9a721 is the delivery baseline that satisfies the required checks.
Three nonblocking improvements remain outside this repair: heal diagnostics can show a zero source time after decode/parse failure; the initial-sync unit test does not execute the local peer branch and therefore does not prove that recovered Created is persisted locally (that production path was reviewed); and strict system-log capture could need isolation if concurrent background logging is introduced. No additional production accessor or test hook was added for these points. A counterfactual failure proves the assertion actually reached, such as empty-baseline noise. Warning-level and per-reason deduplication assertions pass after repair; that is not a claim that each was separately demonstrated failing on old code.
Operational boundaries that remain
- Historical timestamp pollution is not reconstructable. Arrival-time replacement and old untimestamped events have lost source facts. Installing the repair cannot recover their true historical order. Inspect every site and resubmit the intended configuration or deletion at the authoritative site.
- Zero-time typed events remain compatible. They use monotonic local time and emit
legacy-zero; the existing zero-time constraint for bulk remains. These events are outside the timestamped-source convergence guarantee. - Bucket identity conflicts are not merged automatically. Resolve generation differences first. Events before target Created are not applied; unknown creation time is recovered only from a real physical bucket. Unknown or missing buckets are not written. An existing bucket whose physical creation time is still zero produces an InternalError (500) and aborts initial synchronization; a missing bucket instead retains NoSuchBucket (404). An invalid existing field can also block a typed write with InternalError; heal may replace it using a valid peer source. The recovered value is a physical approximation from bucket-directory modification time. It can change with top-level entries, differ across drives, and be later than actual creation. Older source events are still skipped; one event is insufficient to lower the bucket identity. A successful configuration write or initial site sync records the recovered value. Until then, status reports the unknown time and periodic healing skips the bucket in both directions.
- A physical clock is not a causal clock. Local correction can advance beyond an already known future field time, but cannot infer the business intent of all concurrent writes.
- Diagnostics are bounded; success does not mean applied.
legacy-zero,before-created,indeterminate,unreachable, andpeer-errorreuse LogOnceIf with stable keys and error text, details in attributes, and existing hourly cleanup. Each reason keeps its own key, so one condition cannot deduplicate another away. Empty baselines and peers that do not yet have the bucket stay quiet. Invalid existing states emitindeterminateeven when no valid source can be selected, without causing an RPC. Normal duplicates and older events remain quiet. This is a per-bucket/field/reason bound: both unreachable-peer warnings and indeterminate-state warnings can grow with bucket and populated-field count, not a fixed site-wide limit. - Code, integration, and release require separate acceptance. Issue state, Server version, images, packages, published documentation, and production settings each need their own evidence. The main-branch merge recorded here does not establish that release artifacts or production installations have been upgraded.
3 - An Unsigned Header Is Not Part of the Request
Publication update, 2026-09-17: The September source repairs discussed here shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
This record describes the unsigned-header coverage repair committed to SILO as 123325430 and merged through PR #173, tracked as SN-2026-011. It was reported by Oren Yomtov against a released build and reproduced locally on both signature paths.
Status on 2026-09-11: the original repair is pushed and merged through PR #173. The follow-up signing and payload-verification fixes described below are also merged through PR #177, with all eight PR checks passing. Source validation and published releases are separate: the currently published September 3 Server release does not contain these fixes.
Scope: SigV4 header coverage, consistent policy inputs and body-checksum verification. S3 field names, object and bucket metadata formats, replication protocols, encryption formats and client commands are unchanged.
Security property for ordinary signed and presigned SigV4: unsigned client-suppliedx-amz-*operation headers cannot change an authorized request; policy evaluation and body verification use the effective signed inputs.
Too Long; Didn’t Read (TL;DR)
A presigned PUT URL can sign only the host header. SILO confirmed that each header named in the signed-headers list had arrived, but it never walked the headers that actually arrived, so an x-amz-* header outside that list was accepted and used. cmd/api-router.go routes any PUT carrying x-amz-copy-source to CopyObjectHandler on that header alone. Together these turned a write grant for one object into a server-side copy that reads any object the signing key can reach, executed as the signer — a confused deputy. The Authorization-header path behaved the same way when the header was left out of SignedHeaders.
The repair states one invariant:
This matches AWS S3, which refuses the same request with AccessDenied (“There were headers present in the request which were not signed”). The affected signing code was inherited from upstream minio/minio. The released SILO baseline predates this repair; source provenance alone does not establish the status of every upstream build or other fork.
Failure: the coverage gap
extractSignedHeaders in cmd/signature-v4-utils.go iterates the signed-headers list and, for each name, pulls the value from the request (or the query string). It proves that every promised header is present. It never asks the opposite question — is every x-amz-* header that arrived actually in the list?
The one place that walked the arriving headers, checkMetaHeaders, matched only the X-Amz-Meta- prefix and was called only from the presigned path (doesPresignedSignatureMatch). The Authorization-header verifier (doesSignatureMatch) called nothing equivalent. So an unsigned x-amz-copy-source — or any other operation-shaping x-amz-* header — sailed through on both paths:
Reproduced locally on RELEASE-style builds: the control PUT returns 200 with an empty body; the same URL plus the one unsigned header returns 200 with a CopyObjectResult whose ETag is the md5 of the victim object, and the destination reads back the victim’s bytes. Where the destination bucket already allows anonymous GetObject, the copied private bytes are then readable with no credentials at all.
Provenance
The gap is inherited from upstream MinIO; SILO did not introduce it. The SigV4 verifier in cmd/signature-v4-utils.go and the Authorization-header path doesSignatureMatch in cmd/signature-v4.go are original MinIO code dating to 2016, and the header-driven CopyObject dispatch in cmd/api-router.go traces to 2019. The only routine that ever walked the arriving headers, checkMetaHeaders, was added upstream on 2023-07-27 in minio/minio#17737 (535f97ba6). Upstream therefore recognized the class — an unsigned header must match the signed set — but scoped the check to the X-Amz-Meta- prefix and to the presigned path, leaving x-amz-copy-source and the whole Authorization-header path uncovered. The inherited verifier predates the SILO fork.
Before the unsigned-header repair, SILO’s change to cmd/signature-v4-utils.go was the one-line dependency-path migration in 9b11dc946, moving the policy import to pgsty/silo-pkg/v3. The vulnerable verification behavior came from upstream. The original repair (123325430) and the follow-ups in PR #177 change that boundary. The vulnerable code predates SILO’s fork baseline, the upstream 2025-12-03 maintenance-mode commit from which the first SILO release was cut.
This record establishes the repair in the maintained SILO source graph. It does not claim that SILO is the only implementation with a fix, or establish the current maintenance status of other projects.
The repair
checkMetaHeaders becomes checkUnsignedHeaders, is broadened from the X-Amz-Meta- prefix to all of X-Amz-, and is called on both the presigned and Authorization-header paths. A header that is not covered by the signed set is refused with ErrUnsignedHeaders before any handler logic runs.
Four decisions shaped the exact boundary. Each had a plausible alternative that was rejected for a concrete reason.
Membership, not value equality
The inherited check compared signedHeadersMap.Get(k) == val[0]. For a header absent from the signed map, Get returns the empty string, so a header whose first value is empty compared equal and passed. A multi-value header such as X-Amz-Copy-Source: ["", "/src/secret"] could therefore smuggle an unsigned copy-source past a value-equality check. The repair tests membership in the signed set instead. A signed header’s value is already bound by the signature, so value equality was never the property that mattered; presence in the list is.
Exempt X-Amz-Content-Sha256
X-Amz-Content-Sha256 can be omitted from SignedHeaders because the effective payload hash is bound separately in the canonical request. For presigned requests, the query value takes precedence, with a header fallback when the query value is absent. An explicit UNSIGNED-PAYLOAD remains valid. PR #177 aligns the policy condition with this effective value while preserving header-presence semantics, and checks header-only presigned body hashes in the generic authentication path as well as upload paths. The exception does not permit policy evaluation or body verification to use a different value.
Derive signature age from the signed date
The original repair exempted an internal x-amz-signature-age scratch header written after verification. That was too late for PUT and UploadPart authorization, which runs before signature verification. PR #177 instead derives s3:signatureAge directly from the signed X-Amz-Date, and removes the scratch header, its constant and its exemption. A forged date fails signature verification; an unsigned client header under the old name is rejected. Verification remains idempotent without mutating request headers.
Inject X-Amz-Tagging after authentication
PutObjectTaggingHandler derives an X-Amz-Tagging header from the request body so that policy conditions can read it, and it previously did so before authenticateRequest. With the broadened check, that server-synthesized header — which the client never signs — would be refused as unsigned. The injection now happens after signature verification and before authorization, which still has it for policy conditions. Rejected alternative: blanket-exempt X-Amz-Tagging the way content-sha256 is exempt. That would let a client set object tags through an unsigned header on any signed or presigned write, reopening a smaller version of the same class of bug.
Scope across signature modes
- Authorization-header (signed) and presigned SigV4: both now enforced. These are the reachable paths.
- Streaming SigV4: the seed verifier does not call
checkUnsignedHeaders. The copy handlers reject streaming authentication through their ordinary authentication dispatch, but that does not establish complete header coverage for streaming PUT/UploadPart paths. No dedicated streaming-copy rejection test is claimed here. A universal coverage guarantee requires separate implementation and regression evidence. - SigV2: unaffected. V2 canonicalization folds the
x-amz-*headers into the string-to-sign by construction, so an addedx-amz-*header changes the computed signature and is rejected as a signature mismatch.
Status code: 400 versus 403
AWS returns 403 Forbidden for an unsigned header; SILO returns 400 AccessDenied (ErrUnsignedHeaders), inherited from upstream. The attack is refused either way, and the error Code string is identical; only the HTTP status differs. Raising it to 403 is a one-line change to cmd/api-errors.go that also shifts the pre-existing meta-header rejection. It is left as a deliberate, reversible election rather than folded silently into a security fix, because it is a behavior change for the existing unsigned-meta-header path and is not required to close the vulnerability.
Tests
Several existing tests built a signed request and then set x-amz-copy-source, x-amz-copy-source-range, or x-amz-metadata-directive after signing — that is, they depended on the very behavior this fix removes. They now re-sign with signRequestV4 after setting those headers, as required for these operation-shaping headers by the repaired verifier. signRequestV4 excludes the Authorization header from its own signed set, so re-signing is safe. Current coverage includes checkUnsignedHeaders unit cases for empty first values, the payload-hash exception and rejection of the obsolete unsigned signature-age header and TestPresignedVerifyIdempotent, which verifies the same presigned request twice.
Evidence
- A built server reproduced the confused deputy on both the presigned and Authorization-header paths, then refused both after the fix while the control
PUT, a realminio-goCopyObject,PutObjectwith user metadata and tags, and body-basedPutObjectTaggingall continued to work. go test ./cmd/passes on the fix tree;gofmt,gofumpt, andvetare clean.- Adversarial review (round one) independently surfaced three defects in the first draft — non-idempotent verification via the scratch header, the empty-first-value bypass, and over-rejection of an unsigned
x-amz-content-sha256— each of which is addressed above and confirmed by re-running the reviewer’s own adversarial test suite against the final tree. - Adversarial review (round two, against the committed fix) found no regression and confirmed the re-signed tests keep their original intent: an invalid access key still returns
InvalidAccessKeyId, and a wrong SSE-C key still returns403after the signature validates. It surfaced three adjacent, pre-existing gaps that also fail on the parent commit and are out of this change’s scope; they are recorded under follow-ups below.
Compatibility and operations
- Ordinary clients: no request change. Conforming ordinary signers include the nonexempt
x-amz-*headers they send. Custom signers must verify this contract; streaming modes have the separate boundary above. - Unsigned
x-amz-*headers: now refused withAccessDenied, as on AWS. A client that added such a header without signing it was already outside the SigV4 contract. - Rolling upgrade: wire and storage formats are unchanged. Upgraded nodes enforce the boundary; nodes still running an older build remain exposed until upgraded, so behavior can differ by node during the rolling window.
- Rollback: data written by the fixed version stays readable by the previous version, but rollback reopens the confused deputy.
Residual risks and follow-ups
- Release delivery: a source fix and a public engineering record do not establish that a published binary or image contains the fix. Verify the selected release and artifact separately.
- CVE: the reporter requested one; the finding carries the stable fork-local
SN-2026-011identifier until a CVE is assigned. - Status code election: the
400-versus-403choice above is open. - Adjacent signing fixes: PR #177 addresses repeated copy-source ambiguity, signature-age authorization ordering, and the effective payload-hash policy value. It also closes the separately reproduced header-only presigned body-checksum gap. The regression set covers signed and presigned requests, policy enforcement before upload verification, and real HTTP bucket-policy tampering. These follow-ups are distinct from the original
SN-2026-011finding; their merge and release status is recorded above. - The general question: this repair covers
x-amz-*request headers. Any future control that lets request syntax select an operation must answer the same question this one did — is this value covered by the signature before it is allowed to mean anything? The repeated-header gap above is the same question in a different guise: the value the signature binds and the value the handler consumes must be the one and the same.
Conclusion
The signature is the request. Everything an x-amz-* header claims is a claim until the signature covers it:
Confirming that the promised headers arrived is not the same as confirming that the arrived headers were promised. Refuse any unsigned
x-amz-*header before the handler runs, on the ordinary signed and presigned paths covered by this repair. Streaming coverage remains a separate boundary.
The ledger separately tracks the payload-verification repair as SN-2026-012. The signing change alone does not change storage formats, but the same main candidate also contains IAM changes requiring coordinated upgrade. Do not use this record as approval for a rolling upgrade of that entire candidate.
4 - Durable IAM Revocations
Publication update, 2026-09-17: The September source repairs discussed here shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
Source status, 2026-09-16: #191 and #192 are merged into main, but are absent from published Server 20260903. They change persistent IAM state and require a coordinated upgrade. Follow the upgrade and recovery runbook and SN-2026-013.
SILO retains the version of a deleted IAM record so an offline site cannot restore an older identity or grant when it reconnects. This covers built-in users, service accounts, groups, policy documents, and policy mappings in their actual user/STS-parent/group namespaces. It also revokes the deleted built-in user’s older service accounts, STS credentials and group grants across deliberate same-name recreation.
Ordering and persistence
Deletion records occupy the original IAM configuration paths. They contain the
originating timestamp, Deleted, and, where required, RevokedBefore; identity
deletion records contain no secret key or session token. Normal IAM listings and
authorization hide deleted records. Object storage and etcd both serialize each
path’s version comparison and write with a distributed lock. The source timestamp
is persisted without replacing it with the receiving node’s clock.
Older events cannot overwrite a newer revision. At an identical timestamp, a deletion wins over a live record. A local deliberate recreation receives a version newer than the stored deletion. An older user/group deletion arriving after recreation retains its revocation boundary while preserving the newer live record. Receiving an already-applied tombstone does not rewrite or advance it. These rules also apply after cold loading persistent IAM state.
The user or group revision is the commit point of deletion. Cleanup of mappings and children follows that commit; cleanup failure cannot undo it. The API returns the cleanup error and still notifies sibling nodes to reload the committed state. Such an error does not mean the identity is still active. Retry the intended revocation only while IAM writes and same-name recreation are paused. When the API returns a committed-cleanup error, the admin handler does not send its immediate cross-site hook; cross-site propagation relies on the normal deletion-healing retry. Sibling deletion notifications reload current shared state, so a delayed notification cannot delete a subsequently recreated identity. Etcd siblings also receive persistent changes through watches. Failed notifications/watches remain eventual propagation, not a distributed instantaneous revocation transaction.
Keep node clocks synchronized and monitor offsets. Ordering uses source wall clock timestamps, with monotonic advancement for local writes to the same path. It does not establish causal order between concurrent writers at different sites or resolve conflicting live updates with identical timestamps deterministically.
Users, child credentials and groups
A recreated user retains RevokedBefore. New service accounts and built-in STS
issuances carry a signed siloParentRevocation claim identifying the parent
boundary known at issuance. Editing or replaying an old child does not update
this claim. Old children remain invalid even when their own update timestamp is
newer than the parent deletion; children issued for the recreated parent remain
valid. Claims are read from verified tokens on credential load/write.
A newer replicated service-account snapshot can replace an older service with the same access key, including a deliberate change of owner or secret. Local duplicate creates remain rejected. Snapshots preserve disabled status and the service’s own revocation boundary, so earlier mappings cannot attach to the new service. Equal-version retries reload the committed identity without rewriting it. Periodic live healing compares source versions even when the public status summary is unchanged, and includes disabled identities as healing sources. Service snapshots also preserve their absolute expiration. The receiving site does not reapply the local minimum lifetime for a newly issued credential. A newer already-expired snapshot still supersedes the old key, is denied by authentication, and is collected into a durable service tombstone by normal loading. Cache/claims loading failures are returned for retry, not acknowledged; after a committed replacement the stale cached secret is evicted immediately. Collisions with an existing built-in or cached STS identity report an error and require an explicit administrative resolution; replication cannot change its credential kind. Concurrent conflicting service updates with exactly the same timestamp can retain different winners at different sites; the status summary does not resolve that case.
Each group member has its own MemberGrants timestamp. Changing another member
or the group’s enabled status does not reissue everyone else’s grants. Effective
membership requires the grant to be newer than both the user’s and the group’s
retained boundaries. Listings and policy evaluation use the same effective
membership. Peer snapshots preserve grant times, including unknown legacy grant
times; they cannot treat a recent snapshot time as a fresh grant to a revoked
identity. A new explicit administrative group grant can restore access.
This does not implement a general conflict-resolution protocol for all group membership edits. In particular, the inherited live-group snapshot merge adds members and does not reconcile a missed ordinary member removal. Removing a member from a live group during a site outage is a separate known limitation; do not infer that this change resolves it. User/group deletion boundaries and same-name recreation are covered here.
Retention and expiration
Permanent identities, groups, policy documents and mappings have no automatic tombstone TTL. A disconnected peer or an old backup may return arbitrarily late. Successful replay acknowledgements are an optimization, not permission to garbage-collect this history.
Natural expiration of an immutable STS token physically removes its token-key record and any legacy token-key mapping, without generating a permanent tombstone. An early STS revocation is retained until that token’s expiration plus the existing clock-skew allowance. Replaying the same revoked token with a later event timestamp cannot recreate it. Etcd uses an expiration lease; object storage collects expired STS tombstones during its existing credential loading/purge. A record with unknown expiration is retained conservatively. Cleanup writes are best effort and use a short lock budget. If one fails, that load stops optional reclamation, reports the error and still loads healthy users; expired credentials stay denied and retain their existing durable version. The next load retries. Healthy cleanup has no fixed record quota. The reusable STS parent policy mapping is not assigned the token’s TTL by deletion cleanup. External-IDP disablement is an early revocation, not natural token expiration, and its cached STS and service accounts are included in cleanup. Expiring service accounts retain a durable revision because their access keys are reusable and an older version might have no expiration.
Direct per-token RevokeTokens delivery between sites is not a new guarantee of
this change. The guarantee for built-in parent deletion follows from the durable
parent boundary, including children not currently present in the deleting node’s
cache.
Healing, failures and operational cost
Each process starts one healing loop. Losing its distributed leadership lease pauses work until leadership is reacquired; it does not permanently terminate healing. Configuration reloads do not create extra loops. The 30-second interval starts after leadership is acquired and after each completed pass. Initial lock retries and endpoint recovery can add further delay; it is not a convergence SLA.
The normal IAM loaders maintain an in-memory index of deletion records and
retained boundaries, without secrets. Healing uses this index; it does not add a
second full walk of config/iam/ every cycle. Existing full IAM loading still
scans persistent records, including tombstones, at startup and on refresh.
Each peer receives batches of at most 128 records. The sender remembers which path/version each peer acknowledged. Unrelated new changes at either site do not reset that progress. A failed batch remains pending while later independent batches can progress; a lost response may cause safe idempotent replay. A pass has a bounded duration, and its successful acknowledgements survive that timeout. The protocol reports each node name and process instance. Switching between known node instances behind a load balancer preserves acknowledgements. A new node instance conservatively invalidates prior acknowledgements once; repeated switches among those known instances do not reset progress. Restore persistent state only with the affected processes stopped, so a restore cannot reuse an old process acknowledgement.
Steady-state healing still traverses/sorts the retained in-memory set and checks the peer’s protocol status. It suppresses repeated deletion PUTs once acknowledged. Memory use scales with retained paths and peers; startup storage reads scale with history. This release does not provide general history compaction. After upgrading, older live records whose receiving sites originally assigned different timestamps can require an initial reconciliation wave. Allow for its storage writes and sibling notifications when planning the maintenance window.
The cluster IAM metrics include revocation_records,
revocation_heal_failures, revocation_heal_duration_millis, and
revocation_heal_last_success_timestamp_seconds. Errors are also logged. A
nominal 30-second scheduler interval is not a convergence deadline: outages,
large backlogs, lock contention and failed requests can require more passes.
Cached credential lookup checks the in-memory parent revision index and performs no additional storage read. STS issuance and cold credential loading still consult the persistent parent revision. Revision I/O and distributed lock waits release the IAM cache lock while a separate local writer mutex preserves write order. These operations have bounded contexts, including etcd lock and lease cleanup.
Protocol and supported upgrade
The server-owned versioned route is
/minio/admin/v3/site-replication/peer/iam-revisions. It carries source versions,
member grant times and distinct user/group revocation items without changing the
admin client SDK or S3 API. Older servers reject this route. The sender reports
the failure and does not fall back to a route that would discard the metadata.
The existing legacy IAM route remains readable for best-effort compatibility;
this does not confer the new guarantees on an older peer.
All participating servers must be upgraded for the guarantee in this document. Mixed old/new nodes sharing an IAM backend and rolling downgrade are unsupported: older binaries do not interpret tombstones or signed parent boundaries correctly. Use a maintenance window for coordinated upgrade:
- Pause IAM changes and isolate any offline site or backup whose state is unknown.
- Back up each site’s complete IAM storage and required encryption material. A live IAM admin export omits deletion history and is not an adequate backup.
- Stop all nodes sharing each site’s IAM backend, replace their binaries, and restart them on the upgraded version. Complete this for every participating site before relying on the new revocation semantics.
- Check IAM loading, site-replication errors and revocation convergence. Verify representative old credentials are denied and deliberately reissued ones work.
- Resolve pre-upgrade revocations explicitly. Absence cannot reconstruct an already-lost deletion version: remove surviving old records on the sites that still have them, and rebuild stale offline peers from approved state before admitting them. Do not reconnect an unknown old snapshot just to discover its deleted credentials.
Credentials issued by an older server for a recreated parent lack the required signed boundary and must be reissued by an upgraded server. Parents without any retained revocation history preserve existing credential behavior.
A pristine built-in policy remains protected from local deletion. If an administrator explicitly overrides that policy and later deletes the override, the durable deletion now suppresses automatic recreation of the built-in policy on reload. This prevents reload from undoing the deletion. Restore the policy by an explicit policy-create operation if desired. Local deletion of a nonexistent policy remains idempotent and does not create a new tombstone; replicated unknown deletions retain their version.
For rollback, stop and isolate the affected sites and assess changes since the backup before restoring compatible state. Restoring an older backup can itself lose later revocations and requires reconciliation/rekeying before access is reopened. Do not delete tombstones online or convert only live IAM records to make an older binary start. Server, client, Console, package and deployment acceptance remain separate delivery gates of the maintained PGSTY stack.
Regression and performance checks
Focused coverage is in iam-revocation_test.go, iam-revision_test.go,
iam-revision-lock_test.go, iam-revision-boundary_test.go,
iam-replication-protocol_test.go, iam-credential-retention_test.go,
iam-peer-reload_test.go, and iam-replay_test.go. Set
SILO_TEST_IAM_REVOCATION_ETCD to a disposable etcd endpoint to include backend
lifecycle/locking/boundary tests; they use isolated key namespaces.
BenchmarkIAMCachedCredential and BenchmarkIAMSetTempUser can be run against
the pre-change source for a comparable local baseline.
BenchmarkIAMRevisionConvergedHealing covers 1,000 and 10,000 retained records;
it measures steady-state index/network work and asserts zero repeated PUTs. Its
fake peer does not measure durable catch-up throughput.
BenchmarkIAMColdLoadExpiredServices measures loading and cleanup with 100 or
1,000 expired reusable credentials; run it with -benchtime=1x. Use actual multi-site
signed S3/STS tests and deployment-specific latency/scale measurements in
addition to these component tests.
Errors and observability
An admin delete can return HTTP 500 after the authoritative revocation has committed but dependent cleanup failed. It is not evidence that the old identity remains usable. Retry the intended revocation while IAM changes and same-name recreation are paused, then verify old and reissued credentials at every site. When a persistent parent revision cannot be read or issuance cannot be committed, STS fails closed through an internal-error path, including STSInternalError; it must not mint a credential by guessing a missing boundary.
Scrape each process’s authenticated /minio/metrics/v3/cluster/iam endpoint:
| Metric | Interpretation |
|---|---|
minio_cluster_iam_revocation_records |
Retained deletion records and parent boundaries in this process’s index |
minio_cluster_iam_revocation_heal_failures |
Failed convergence passes since process start |
minio_cluster_iam_revocation_heal_duration_millis |
Duration of the last pass |
minio_cluster_iam_revocation_heal_last_success_timestamp_seconds |
Unix time of the last successful pass |
The index is per process; summing it across siblings double-counts shared state. Healing requires site replication and a leadership lease. Shared-backend deployments without site replication do not run the site healing pass; zero healing metrics are not evidence of a failure there. A healthy counter on one leader does not prove every peer’s credential checks pass.
Why simpler alternatives fail
Physical deletion alone loses the ordering evidence an offline peer needs. Comparing only the child’s update timestamp is also insufficient: editing an old child after a parent deletion would appear newer and could restore access. The signed issuance boundary remains unchanged by such an edit. Replacing a received source revision with local receive time breaks source order and can turn a retry into a new mutation. The implementation preserves source timestamps and separately advances only deliberate local writes. Tombstone TTL or restoring only live export records discards the very history that prevents replay.
The versioned source and metrics definitions define this contract. This record replaces the former in-repository IAM design; it is not a production recovery certificate.
5 - Federated CopyObject: Preserve the Destination Contract
Publication update, 2026-09-17: The federated CopyObject repairs recorded here shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
Source status, 2026-09-16: this records the main-branch CopyObject fixes in #157, #159, #163, #177 and #179. They are absent from Server 20260903. The subject is the legacy etcd bucket-federation path that forwards a copy to another deployment as PutObject, not the bucket/site replication scheduler.
Forward bytes once, encrypt at the destination
The proxy reads logical source bytes: it decrypts and decompresses as needed, then declares the logical length to the destination. It must not locally encrypt/compress and ask the remote to do it again. SSE-C reads still require the source key and secure transport. The destination’s explicit SSE headers or its own encryption defaults choose how the new object is stored; the proxy must not inject its local defaults when the caller selected none. SSE-KMS context is forwarded in the expected JSON-object form.
Trusted raw SSE-C replica CopyObject is explicitly rejected with 501 NotImplemented in this federation path, before creating the destination. That combination is not made safe merely by passing replication markers. Ordinary key-authorized SSE-C copies and the dedicated replica path are different operations.
Checksum and metadata rules
Remove every reserved internal metadata prefix, case-insensitively, from the ordinary forwarded write. Forward public metadata and tags through supported fields, not internal storage encoding. The checksum must describe the logical full object actually written:
- A requested algorithm or a multipart-composite source requires a full-object checksum at the destination.
- For a nonempty stream, the client uses a trailing checksum, including
x-amz-trailer; the destination must consume and validate that trailer. - Empty content uses the ordinary checksum header because there is no streamed checksum trailer to carry its digest.
- A stored full-object source checksum can be forwarded as an ordinary checksum header for validation.
- A required remote checksum that is missing, malformed or has a multipart
-Nsuffix is an error; it must not be reported as a valid full-object result.
The remote write may already have committed when a malformed success response is detected. A resulting error does not prove destination absence. Inspect the written version before blindly retrying a versioned copy.
Object Lock is not user metadata
Forward legal hold through the typed LegalHold option so it becomes x-amz-object-lock-legal-hold, not x-amz-meta-*. Retain the full precision of the retention timestamp; routing it through a whole-second conversion would weaken the requested value. Authentication and the destination’s Object Lock rules still apply.
Return the committed write
The response and ObjectCreated:Copy event describe the destination key, logical size, ETag and exact version ID. Modification time is obtained from the destination write, not invented at the proxy or obtained through a later unversioned HEAD that could observe another writer. The internal write-time response is bound to the authenticated federation request. Source-version response headers still identify the selected source when one was provided.
Evidence and deployment boundary
See the handler, write-time transport, and object-copy-federation*_test.go / object-federation-time_test.go. Tests cover ordinary, empty, multipart, compressed and encrypted sources, remote checksum failures, destination defaults, legal hold, response identity and events. These component fixtures do not certify a production etcd federation or every S3-compatible destination.
Verify both forwarding and destination builds from the component matrix. Existing objects are not rewritten by upgrading. Related records cover SSE-C replica integrity and Object Lock ordering.
6 - Multi-Pool Object Consistency
Publication update, 2026-09-17: The September source repairs discussed here shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
Source status, 2026-09-16: #178 repairs multi-pool mutation ordering; #207 extends current-object selection to conditional PUT. These changes are merged into main and absent from Server 20260903. They close the source defects tracked by #133 and #144; this does not establish acceptance of a new release artifact.
One key can have several physical copies
A pool expansion, decommission or interrupted cleanup can leave copies of one key in different pools. A per-set lock does not serialize an operation with a writer choosing another pool. Nor does reading the first copy prove that its ETag, tags or retention represent the logical object.
The relevant failures include a stale ETag satisfying a conditional write or delete, a newer retention/hold disappearing behind older object metadata, and an older copy becoming visible after deleting the selected copy. Reading each pool independently and combining only its success status does not solve these races.
Shared selection and lock boundary
The pool coordinator holds one namespace object lock across selection and mutation. objectPoolInfos reads the addressed version in every pool, including draining pools. Unknown/unreadable state is an error, not evidence that the object is absent. Copies are ordered by modification time with pool index as the equal-time tie-breaker.
An unqualified metadata request first resolves the logical current version, then gathers copies of that version. It must not merge metadata from unrelated object versions. Object Lock retention, legal hold and tags have independent source timestamps; mergedPoolObjectInfo combines those fields by their own order rather than treating the object’s modification time as every field’s revision. An ordered removal is state too.
Metadata writers, relevant object writers and healing use the same coordinating lock. The lock is not a transaction that can undo completed disk writes on every pool after a later failure.
Conditions and deletion
For conditional DELETE, evaluate once before cleanup and clear the callback before invoking lower pools. An explicit versionId compares that version. For conditional PUT, select the latest logical representation across eligible pools while holding the outer lock, including the delete-marker state, before installing the replacement. #207 addresses the case where a stale destination-pool copy made If-Match/If-None-Match disagree with the current object.
Cleanup and metadata propagation can still partially modify physical copies before a storage failure is reported. A quorum/cleanup failure must not be interpreted as an unchanged object or as proof every old copy disappeared. Verify state and retry after recovery. Batch XML ETag conditions remain unsupported; see conditional DELETE.
Object Lock and replication
Replica writes reconcile destination retention and legal-hold state against authoritative same-version copies across pools. A stale incoming replica must not shorten a newer retention or turn off a newer hold. Tags similarly travel with their timestamp, including an empty deletion value. See Object Lock ordering and replicated tags.
Tags across pool migration
Rebalance and decommission read each version as a FileInfo, convert it to an ObjectInfo — which moves the tag header out of the ordinary metadata map into UserTags — and rewrite the object in the destination pool. Since their introduction in 2022 the migration writers copied only the ordinary metadata map, so tags were dropped for ordinary and multipart objects on both entry points; the 2020 field split itself is correct. Commit fced86303 gives the four migration writers one metadata assembly that clones the map, restores UserTags and keeps the tag revision fields exactly as stored, without inventing timestamps, under the existing coordination locks. Regressions cover ordinary and multipart objects, initial tags, updated and cleared tags, and versioned and legacy readers.
The repair only prevents future loss. Tags already dropped by an earlier migration are not recovered; restoring them needs a trusted source of the old values. The post-repair eight-node rebalance and decommission acceptance was not completed at the time of writing.
Evidence and operator impact
The coordinator implementation, pool entry points, and the two merged PRs identify the source contract. Coverage includes multiple pools, stale copies, delete markers, explicit versions, concurrent writes and failure paths. It is not a claim of atomic rollback or arbitrary fault tolerance.
Upgrade all participants before depending on the shared-lock guarantees. Inventory historical copies, tags and Object Lock state when a deployment may already have encountered the defect; the source repair does not prove historical cleanup. Use the replica audit runbook and component matrix. The multi-pool repair is separate from the default limitations and scan cost of multipart upload listing.
7 - Object Lock Replication Ordering
Publication update, 2026-09-17: The later repairs in #129, #134 and #178 shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
Release boundary, 2026-09-16: f4c1286c9 shipped in Server 20260903. The later timestamp-only removal, SSE-C retransmission and cross-pool repairs in #129, #134 and #178 are on main and not in that release. The advisory ledger records the original ordering defect.
Preserve the old state before rebuilding metadata
A replica COPY used to rebuild metadata from the incoming request before comparing retention and legal-hold timestamps. The old timestamps were therefore lost, so a stale request could appear authoritative. The legal-hold timestamp also went into the retention timestamp key. The first fix captures the stored state before rebuilding the destination map and keeps each field’s revision under its own key.
Retention and legal hold are independent registers. A later retention does not make an older legal hold authoritative, and the object’s modification time is not a substitute for either field’s revision. Apply an incoming field only when its timestamp is newer than that field’s stored timestamp; retain the stored value for a stale or equal update.
Removal is an ordered value
Removing retention or a hold can leave no live value but still carries a source timestamp. Dropping that timestamp would let a delayed old value return. The later repair recognizes timestamp-only removals during receive, comparison and resend. An absent value with no ordering evidence is different from a recorded removal.
The receiver must preserve these distinctions through ordinary metadata copies and full SSE-C replica retransmission. A retransmitted object cannot erase a newer destination hold or resurrect an older retention just because its body is being rewritten.
Cross-pool authority
If one version has copies in multiple pools, one local set cannot decide the latest field state. The multi-pool coordinator gathers same-version copies under the shared object lock and merges retention and hold independently by their timestamps. Unknown pool state is not an empty value. This closes the source-level boundary originally left open as #133.
Authorization and limitations
Ordering does not grant permission to change Object Lock. The replication trust checks and the relevant S3/admin permissions still apply. It is not a new user API for bypassing governance or compliance retention. It also does not create causal ordering between unsynchronized source clocks or a distributed rollback after partial storage failure.
A historical stale update may already have changed stored state. Upgrading only prevents the repaired paths from accepting the same error again; it does not reconstruct missing retention history. Verify exact-version retention and legal hold at each site against authoritative records before declaring recovery.
Evidence
See the Object Lock merge implementation, the replica write paths in erasure-object.go, and the linked PRs. The tested cases include stale/newer field updates, empty-value removals, distinct field timestamps, SSE-C retransmission and multiple pools. These source tests do not prove every historical replica has converged.
Use the replica audit runbook, SSE-C integrity record and release matrix together. Upgrade all participating nodes before relying on the combined ordering contract.
8 - SSE-C Replica Integrity
Publication update, 2026-09-17: The September source repairs discussed here shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
Source status, 2026-09-16: the repairs in #122, #123, #124, #126 and #134 are on main, not published Server 20260903. Earlier zero-byte/read-attribute authentication and destination-key checksum repairs did ship in 20260903. Do not treat all SSE-C fixes as one release.
Ciphertext and trust
An ordinary SSE-C request supplies the customer key and operates on plaintext. An authorized replica transfer can carry raw ciphertext with the sealed object-key metadata needed to preserve the original object. The destination must store those bytes verbatim; encrypting them again produces an unreadable double-encrypted object. The internal marker alone is not authorization: the exact replication marker and the required replication permission remain mandatory.
This path does not reveal the customer key or make a normal keyless read permissible. Ordinary GET/HEAD and GetObjectAttributes keep their key/permission checks. Federated raw SSE-C replica COPY is explicitly unsupported.
Multipart sizes and checksums
Encrypted part size and logical plaintext part size are different. Replica multipart records preserve the actual logical size of each part so partNumber reads, ranges and GetObjectAttributes agree with the source after overwrite. Checksums must retain the correct encryption context and logical meaning; a checksum response cannot be decrypted with the source key after a copy committed under a different destination key.
For ordinary SSE-C key rotation, any explicitly requested checksum algorithm forces a complete rewrite on current main, even if it names the existing algorithm. A multipart source becomes a single-part destination; the ETag can change and replication retransmits object bytes. An eligible metadata-only rotation without that request preserves the prior checksum state, including absence. See the operator procedure.
Retransmission and Object Lock
A keyless target HEAD can report an existing SSE-C object as inaccessible rather than absent. The sender must distinguish this from NoSuchKey; it cannot assume an ordinary metadata-only COPY will repair the replica. Existing SSE-C replicas use object retransmission. A destination with undecodable old replica state can be replaced through the repaired retransmit path, while preserving the newer destination Object Lock state and ordering removal timestamps correctly.
A failed old replica is not automatically certified repaired after a Server upgrade. Re-read the exact source and replica versions with approved key access, compare the logical bytes and part boundaries, and check retention/legal hold separately. Never infer integrity from a successful metadata-only HEAD alone.
Compression and historical objects
Current main excludes every new SSE-C write from compression, including normal PUT, multipart initiation, COPY and Snowball. SSE-S3 and SSE-KMS continue to follow their encrypted-compression setting. This avoids a replica format that transports ciphertext without the required compression metadata.
Historical compressed SSE-C objects are not automatically rewritten. Preserve their keys and exact version identities, inventory affected data, and rehearse a supported rewrite/recovery path. The change is preventive, not a background migration or a promise that arbitrary old ciphertext can be recovered.
Verification boundary
The compression decision, replication-trust-ssec-replica_test.go, erasure-multipart-ssec-replica_test.go, replication-ssec-retransmit_test.go and compression-ssec_test.go preserve the relevant source and regression contracts. Tests include incorrect-key/unauthorized controls, multipart byte comparisons and retransmission cases. They do not replace a deployment’s historical-state inventory.
See Object Lock ordering, multi-pool consistency and replica recovery. Upgrade every participating Server before relying on the combined behavior.
9 - Replication Reliability: Delete Completion, MRF Visibility, and Resync Cancellation
Publication update, 2026-09-17: The repairs recorded here (PR #162, the second-round PR #196 and the third-round purge repairs) shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
This page records the analysis, design choices, review, and implementation of #153, #152, and #137. They belong to the same replication reliability series, but affect operation classification, recovery visibility, and task lifecycle respectively. One general retry patch cannot repair all three.
As of 2026-09-09: PR #162 is merged as
d1105bbb, and all three issues are closed. All eight checks on the tested PR head, followed by main Go CI and VulnCheck, passed.
Second round, 2026-09-16: PR #196 (fix0c61128d2, verification recordaea3882c9) repaired the replication worker’s own delete exits and the persisted MRF marker-recovery path — a different surface from the first round. It is on main and not in Server 20260903; see the second-round section.
Third round, 2026-09-16: commitseb4f5e5b3,254b19ac0and358ab38fbrepair what an exact-version purge does to drives and what a queued creation does after a purge. On main, not in Server 20260903; see the third-round section.
Review: the plan was discussed with the installed Claude Code Fable 5.1 Max, followed by a review of the implementation. The final verdict was GO.
Delivery boundary: this work completed code, tests, and main integration. It did not create a Server tag or formal release, and does not establish that existing packages, images, or production deployments contain the fixes.
Overall decision and the surrounding series
The selection rule was to repair a reproduced invariant at the smallest boundary that owns it, retain existing recovery mechanisms, and use deterministic tests to prove that work can finish. Additional complexity needs a concrete counterexample.
| Issue | Confirmed defect in current SILO | Selected repair |
|---|---|---|
| #153: delete-marker purge | Single-object DELETE classified a permanent deletion as marker replication; the remote marker disappeared while source purge remained PENDING | Classify by purge status, matching bulk delete, scanner/heal, and resync |
| #152: invisible MRF drops | The queue was already bounded, but internal drop counters did not reach administration or monitoring; object and delete worker arguments were reversed | Expose existing counters, warn at actual drop sites with deduplication, and align worker routing |
| #137: unreliable resync cancellation | One shared token could not stop multiple runs; blocking phases missed cancellation; stale runs could overwrite terminal state | Give each run an owned context, cancel by resync ID, and constrain registration, finalization, and state updates |
The surrounding series had already separated three concepts that are easy to confuse:
- #136 / PR #138 repaired counter completeness: receive and apply the final result before persisting terminal state, without waiting for the one-minute periodic flush.
- #139 repaired outcome accuracy: an existing destination object does not prove that this update succeeded. Success and failure must come from the target’s actual replication result.
- #137 repairs cancellation and resource lifecycle: the task must stop, its walker, workers, and result consumer must exit, and an old run must not pollute a new run’s state. This change preserves the first two contracts rather than introducing another accounting mechanism.
Authorization to use internal replication semantics belongs to the earlier CORS and replication trust record. Authoritative Object Lock state across pools remains tracked separately in #133, open at the original record date and subsequently repaired on main by #178 (not in Server 20260903); see multi-pool consistency. Closing these three issues does not mean every replication concern is resolved.
The release-gating target is the maintained pgsty/silo stack with Console, mcli, and silo-pkg. Compatibility with upstream MinIO/MC remains best effort. Neither a mechanism proposed for upstream nor an external experiment can automatically be treated as an observation on current SILO.
#153: distinguish marker replication from permanent version deletion
The report and the reproduction differ
The original issue described a sustained HTTP 405 storm involving ILM and replication, and proposed treating every delete-marker 405 probe as completion. Current SILO source and measurements do not justify adopting that explanation and patch directly.
The baseline was 450dcb848. The experiment used two locally built SILO servers, separate disposable data directories, and maintained mcli / minio-go clients. A critical observation was that mcli uses bulk DELETE even for a single key; that route was already correct. A direct single-object S3 DELETE was therefore necessary to exercise the faulty entry point.
| Observation | Single-object DELETE before the fix | After the fix |
|---|---|---|
| Matching delete-marker version at the target | Removed | Removed |
Purge state for that version in source xl.meta |
Still PENDING, although ordinary replication status was complete | Cleanup completed on the first replication attempt |
| Original data version | Retained | Retained |
| Scanner needed to finish this cleanup | Required later recovery | No scanner assistance needed in this experiment |
Checking only that the remote version disappeared would falsely declare success. Source metadata must be inspected too. The reproduction establishes incorrect initial purge classification and completion state, plus an unnecessary probe. It did not reproduce the external report’s sustained 405 storm or request-volume figures. Existing scanner/heal and resync producers already select the proper purge path; they cannot be described as necessarily repeating that erroneous probe every cycle.
Why one classification condition is sufficient
When deleting an existing marker, the object layer can return both DeleteMarker=true and a nonempty VersionPurgeStatus. The former describes the version being operated on; the latter describes the operation now required. They are compatible facts.
The old single-delete handler checked only DeleteMarker and populated DeleteMarkerVersionID. Completion then updated ordinary ReplicationStatus instead of completing the purge. The final condition is:
This matches the classification already used by other producers. The fix belongs in the DeleteObject handler. It requires no storage-format change, relaxation of replica deletion protection, or new recovery task. Legacy PENDING entries remain recoverable through the existing scanner/heal purge path.
Why a 405 is not unconditional success
A 405 from a versioned HEAD can establish that the marker exists at the target. For replicating a delete marker, that can mean idempotent completion. For permanently deleting that version, it means there is still work to do.
Writing VersionPurgeComplete merely because HEAD returned 405 could let the source remove its metadata while leaving the unwanted target version behind. The implementation retains existing 405 semantics. Real remote failures such as 403, 405, and 503 must not be disguised as successful permanent deletion.
Source history traces marker-based classification to an upstream 2020-11-19 commit. The 2023-07-10 optimization changed scheduling from dsc.ReplicateAny() to the returned object’s replication/PENDING purge state while retaining classification based only on DeleteMarker. At this location, the 2025-04-02 commit merely moved Pending to replication.VersionPurgePending; it did not first introduce that scheduling condition. This identifies lineage, not a bisect across every historical release, and does not establish when the entire externally reported storm was introduced.
#152: expose drops from an already bounded queue
MRF, or Most Recent Failures, records recent failed replication work for background processing. It already had limits: mrfSaveCh has capacity 100000, and mrfRetryLimit is 3. The code drops a queue entry when RetryCount > mrfRetryLimit or the save channel is full.
The source object and its pending replication state remain. Dropping a queue entry does not mean losing source data. However, subsequent repair depends on the scanner; a prompt retry can turn into a wait for a scan, so this mechanism does not establish a fixed recovery deadline.
The actual defect was that TotalDroppedCount / TotalDroppedBytes increased internally but were omitted from both public statistics snapshots and from Prometheus v2/v3. A deterministic capacity-one fixture demonstrates the failure: one entry is admitted, a 20-byte overflow entry and a 30-byte retry-exhausted entry are dropped, and internal totals become 2 / 50 while administration reports 0 / 0.
The selected minimum change
- Atomically read the existing counters into both administration snapshots.
- Register and load cumulative counters in metrics v2 and v3, and document their meaning.
- Warn at the two actual drop sites: retry exhaustion and a full MRF channel. Fixed messages and deduplication keys prevent changing counter values from defeating deduplication. Handing ordinary worker overflow to MRF is not itself a drop, so it does not gain a warning here.
- Route ordinary objects, healing, and deletion consistently by
(bucket, objectName)to restore worker affinity for the same object.
| Interface | New metric |
|---|---|
| Prometheus v2 | minio_node_replication_mrf_dropped_operations_total |
| Prometheus v2 | minio_node_replication_mrf_dropped_bytes_total |
| Prometheus v3 | minio_replication_mrf_dropped_operations_total |
| Prometheus v3 | minio_replication_mrf_dropped_bytes_total |
These counters accumulate since Server startup and reset on restart. Operations count entries, not unique objects; one object can contribute repeatedly. Bytes cover known sizes, with deletion entries contributing zero bytes. They measure neither data loss nor the complete backlog. Operators should examine increases alongside replication backlog and target health.
This work does not enlarge the queue, increase retry limits, add a persistent retry scheduler, or introduce another backoff framework. Existing limits continue to bound memory use, and the scanner remains the eventual repair mechanism. The repair makes an invisible condition observable; it does not promise a recovery deadline under arbitrary failures.
#137: make cancellation part of a run’s lifecycle
One token cannot cancel a group
Previously, cancellation placed one unkeyed token into a shared channel. One site resync can cover several buckets, with up to 10 concurrent bucket runs, while dispatchers and workers compete to consume the token. Only one bucket might stop; an unrelated task might consume it; or a leftover token might affect a later task.
Two blocking phases had independent defects: a bare receive from Walk output did not observe cancellation, and sending to a full worker channel did not observe it either. Once a worker exited, the dispatcher could block forever on a channel with no receiver. Walk also inherited the parent context, so an early function return could not stop its own walker.
State handling had related faults: site updateState modified a local value without writing it back to the map; bucket Canceled handling was incomplete; and an old finalizer could overwrite canceled state with Completed.
Registration, cancellation, and status share a write boundary
Each resyncBucket creates an owned context.WithCancelCause and registers before waiting for a concurrency slot. Registration and cancelResyncID share the resyncer’s state lock:
- Registration validates that the target still exists and the resync ID still matches, and reads its current cancellation state.
- Cancellation marks matching Pending / Started states Canceled before canceling all matching registered contexts.
- A registered run still waiting for a slot receives cancellation; a run registered after cancellation sees the canceled state.
- Walk, dispatch, workers, and result sends observe that run’s context. Unrelated IDs are unaffected, and there is no buffered token for a later task to consume.
A dedicated operation mutex serializes site start/cancel configuration setup. Canceling running work still executes if part of the target-configuration loop fails. Bucket finalizers and counter updates check the current target and resync ID, ignoring late results from removed, replaced, or canceled runs. Site state is actually written back, and a late Completed result cannot overwrite Canceled.
Success and failure require different finalization order
The result-consumer contract from #136 must remain intact: persist terminal state only after the consumer finishes.
Canceling before joining workers on a normal completion path could discard pending work and turn a successful run into Failed. The final code calls finish exactly once from one defer, using defer order to complete cleanup. It therefore needs no additional sync.Once.
WithCancelCause distinguishes an explicit user cancellation from parent-context interruption. User cancellation becomes Canceled; parent interruption observed during finalization must not leave an interrupted run marked Completed. Shutdown while a run is still waiting for a slot preserves the existing Pending state for restart recovery.
Protect terminal state and resumed work
Cancellation can still arrive between the context check and the terminal save. Under its lock, markStatus must persist an existing Canceled state even if the finalizer previously computed Completed. An old resync ID must also be unable to overwrite a new run.
This is not a transaction across metadata files. If a bucket completed and persisted before cancellation, a subsequent site cancellation can leave different records reading “bucket Completed, site Canceled.” Completion happened first. The guarantee is that late completion cannot turn an already canceled run back into Completed, not that cancellation erases work completed before it.
The recovery loader, loadResync, previously launched goroutines and then immediately executed defer cancel(). Inspection of SILO’s actual shared-lock implementation confirmed that this cancel ends the merged leader context; it is not a no-op. A WaitGroup now retains that context until resumed runs exit. Losing leadership still cancels the existing context; replacing it with a global context would bypass the leadership constraint. Loading disk state must also preserve newer in-memory start/cancel state.
Fable review and the complexity decisions
The review used the installed Claude Code with model argument claude-fable-5-1[1m] and --effort max. Baseline reproductions and the minimal proposal produced agreement on all three repairs. The final patch and validation were then reviewed again, yielding GO. The conclusion rests on code and evidence, not model agreement alone.
| Proposal or review point | Final decision |
|---|---|
| Treat a purge’s 405 probe as completion | Rejected: marker existence does not prove permanent deletion |
| Add bounded retries, backoff, and a persistent MRF scheduler | Not introduced: the existing queue and scanner already provide recovery; the demonstrated gap is visibility |
| Send more shared cancellation tokens | Rejected: this still cannot guarantee identity routing, broadcast, or isolation from later tasks |
| Add a separate cancellation tombstone registry | Unnecessary: existing target status and resync ID under one lock close the registration race |
| Give each run an owned context and cancelable blocking operations | Retained: full-queue deadlock and walker leaks have deterministic reproductions |
Protect finish with sync.Once |
Initial review required protection against double close; the final implementation has one deferred call site, and re-review accepted omitting Once |
| Wait for resumed runs before releasing leader context | Initial review required checking necessity; SILO’s cancel is effective, so the WaitGroup stays, with leadership-loss coverage |
The final review accepted two further implementation boundaries. The infrequent resync start operation holds the status lock across configuration reads and writes, consistent with existing terminal saves. The active registry uses resyncOpts, including resyncBefore, as its key; current callers reuse the same in-memory values, and no identity mismatch was reproduced. Future changes that reconstruct time values or restore task identity should revisit equivalence, rather than adding another registry without evidence now.
Validation and reproducible evidence
Regressions were added against unchanged production code first, and failed before the fixes. New tests exercise the production handler, real erasure storage, actual metric registration, and resyncBucket, rather than only testing helper logic that repeats the implementation.
| Validation area | Result and evidence |
|---|---|
| Single-delete purge | ErasureSD and Erasure source/target cleanup; legacy PENDING recovery; marker idempotency; real 403/405/503 failure semantics |
| MRF | Actual capacity-one overflow and retry exhaustion; administration JSON; registered v2/v3 counter types and values; object/delete worker affinity |
| Resync | testing/synctest coverage for blocked Walk, full worker queues, user and queued cancellation, unrelated IDs, later tasks, terminal races, stale IDs, slot release, and leader recovery |
| Complete local suite | go test ./... -count=1 -timeout=30m passed, with 50 tested packages |
| Concurrency and repetition | Focused replication race tests passed; cancellation regressions passed 100 repetitions |
| Tooling and contracts | Local build, vet, lint, generated-file checks, rebrand compatibility guard, and git diff --check passed; dependencies and compatibility baseline unchanged |
| Two native servers | Built from local source; direct single-object DELETE removed the target marker and source xl.meta entry while retaining the original data version; no downloaded Server Docker image |
| Remote integration | All eight PR checks passed, followed by all six main Go CI jobs and VulnCheck |
Tests pinned to the merged revision: delete markers, MRF visibility, and cancellation lifecycle. With that revision and Go 1.27.1, the relevant checks can be repeated with:
After complete local validation, only new-test formatting and fixtures changed: an existing ARN was reused, and the collector supplied the metric prefix, preventing the compatibility scanner from treating test strings as new protocol identifiers. The guard was not weakened. Relevant tests, lint, and compatibility checks were rerun after those adjustments, and remote CI checked the final commit.
| Evidence point | Exact identity |
|---|---|
| Baseline | 450dcb8484bc1337deba0cf608cc893a6691d794 |
| Final PR head | 66fe61ff65c83d68b74baa637a11623015c7aa21 |
| Merged main | d1105bbb3d4a0afa33b3a4ac11b821235038ed0e, with the same source tree as the final PR head |
| PR Go CI | 34320440012 |
| Main Go CI | 34321319278 |
| Main VulnCheck | 34321319274 |
Maintenance and release decisions
Future changes must continue to establish operation identity, visible failure, truthful per-object outcomes, complete terminal counters, and cancellation that releases its own resources. Neither an API returning Completed nor an object existing at the destination can replace those checks.
The code verdict is GO, and the delivery facts are main integration and passing CI. A formal release still requires selecting a tag, verifying packages and images, and establishing that deployments contain the repair. The scanner-dependent MRF recovery delay, the open cross-pool issue, and the externally reported 405 storm not reproduced on current SILO remain part of this decision record.
Second round (2026-09-16): worker-side purge classification and persisted MRF recovery
The first round classified deletes at the handler entry (#153), exposed MRF queue drops (#152), and completed resync cancellation (#137). The second round — PR #196, fix 0c61128d2, integration verification aea3882c9 — repairs a different surface: the replication worker’s own exits, the outer aggregation, and the persisted MRF recovery path for delete markers. The two rounds are complementary; neither subsumes the other.
Status: on verified main
40220bd836cb, not in Server 20260903. All evidence is synthetic (real single-drive and 16-drive storage, signed DELETEs, controlled HTTP targets, real persisted-MRF disk files replayed through a fresh worker pool). The externally reported 405 storm is not reproduced and not attributed to these paths.
What was still broken
- Legacy-shape tasks were skipped entirely. A task with an empty
VersionID, a non-emptyDeleteMarkerVersionID, a COMPLETED creation target, and a PENDING purge target never issued its DELETE — the per-target creation early-return suppressed it. - Fixing only the target function made aggregation lie. The outer status selection keys on
VersionIDand creation state, so a failed legacy-shape purge aggregated as COMPLETED, emittedObjectReplicationComplete, and skipped queueing to persisted MRF — worse than the baseline. - Persisted MRF dropped every marker 405. Replaying a disk MRF entry for a delete-marker version fetched real marker metadata plus
MethodNotAllowed, and the error path discarded it — the marker MRF recovery route was a dead end. - Failed purges overwrote successful resync markers, completed purges were re-sent, and a not-yet-ready HEAD unconditionally overwrote creation state.
- A multi-target empty-state regex misparse could corrupt creation state once. A disk marker carrying only creation metadata, replayed against a task with purge state, let two empty target states parse as a bogus
Pending; the deleted-flag guard then rewrote the whole creation block as empty entries with fresh timestamps — a one-shot corruption that requires disk/task divergence to reach.
Root causes: the worker had no task-level purge classification (the per-target predicate in the community’s #184 was itself wrong — a composite purge status can never reach COMPLETE through it; its investigation and proposed fix nevertheless shaped this follow-up, with thanks to Julien Laurenceau); MRF recovery had no valid-405 identity gate; and the delete task carried no retry count, so the existing mrfRetryLimit drop was unreachable on the delete path.
The repair
- One classification, every exit.
isVersionPurge()(non-emptyVersionID, or a non-emptyDeleteMarkerVersionIDwith a composite purge status) drives both the inner target function and the outer aggregation. Purge exits write onlyVersionPurgeStatusand leaveReplicationStatusempty — the storage layer’s “do not update” signal. - A three-field clear guard empties the composite status on the purge path, making the multi-target misparse unreachable at disk writes.
- Purges send the canonical permanent-delete request — explicit
versionId,ReplicationDeleteMarker=false, no HEAD/readiness probe (authorization is the DELETE’s own). This also prevents a lost-response retry from re-creating a marker at an unversioned fallback. - A valid-405 gate for MRF recovery. A
MethodNotAllowedschedules recovery only when the returned object is a delete marker, the bucket/object/version identity matches, and the modification time is non-zero. - A bounded retry budget. Delete tasks carry a retry counter, incremented at all three persisted-MRF entry points (aggregation failure, lock failure, queue-full fallback), respecting the existing limit; after exhaustion, the scanner can still re-raise healing.
- Audit and event status map
COMPLETEtoCOMPLETEDat the statistics/event boundary only, reusing the existing legacy constant.
What 405 means, precisely
For creation (replicating a delete marker), a HEAD 405 on the target marker version means already created — idempotent completion. For purge (permanently deleting a version), a 405 from MRF identity probing with full marker identity means work remains; the purge’s own success is decided by the DELETE alone, and a DELETE 403/405/503 is always a real failure. Empty or null version markers return ObjectNotFound, not 405, and sit outside the gate.
Boundaries that remain
The legacy in-memory task shape does not serialize across restart and has no current producer — its handling is robustness, not an active repair. Targets without a configured client still only log. Target-level resync replacing a purge subset is a pre-existing defect this round neither caused nor fixed. Persistence was driven directly in tests; timer-based flush and process-crash durability are not claimed. Multi-process site-replication meshes and cross-region acceptance are out of scope. For the tag-ordering and replica-metadata repairs in the same reliability series, see Replicated Tag Ordering and Replica Metadata Normalization.
Third round (2026-09-16): exact-version purges, marker healing and stale creations
The second round made the worker’s purge exits truthful. The third round — commits eb4f5e5b3, 254b19ac0 and 358ab38fb — repairs what a purge does to drives and what a queued creation does after a purge. It was driven by three-site container runs (three sites, four processes and sixteen drives each, EC 12+4) in which a purged delete marker came back on both target sites while the source still read the data version.
Status: on main after the published Server 20260903. Unit and race coverage on single-drive and 16-drive fixtures. The three-site rounds ran on a combined build (
6f27ee6) containing the same purge repair, not on the final main. One adversarial review round (Codex, gpt-6-astra) returned REQUEST CHANGES with a single should-fix, corrected in358ab38fb.
What was still broken
- A purge could create the marker it was removing. The lookup marks a version under purge as deleted for visibility, and the purge write-back reused that flag as a disk instruction. On drives that lacked the marker, the write-back added it; a half-purged set flip-flopped instead of converging.
- A missing version was acknowledged too early. The read path reports a version as absent once half of the drives say so, so a retried purge returned 204 while up to half of the drives still held the marker.
- Healing a marker discarded its metadata. The heal path rebuilt delete markers with an empty metadata map. A healed marker lost its replication and purge state, looked pending again, and could be re-replicated as a fresh creation.
- A queued creation ran after the purge. Every GET/HEAD/LIST of a pending marker, the scanner and MRF queue a creation task carrying a snapshot of the marker, and tasks from several frontends serialize on the per-object replication lock. In the reproduced sequence the user’s purge was acknowledged by both targets, and about 85 ms later a queued creation for the same VersionID and original modification time reached both; the source ended with 0/16 marker copies and each target with 16/16. A target that no longer holds the marker creates it: nothing in the request distinguishes a first creation from a stale one.
The repair
- Purge intent is classified once (
isVersionPurge): a valid explicit VersionID; no marker-creation, prefix, movement, free-version, expiry or transition options; and either a completed purge status or no replication state at all, or a trusted replica request. A purge never sets the marker-creation flags, and its result merges physically removed and reliably absent drive replies. The response still reports a delete marker when a stored marker was removed, and only then; a data version whose earlier purge is still pending is reported as a data version. - Absence needs a write-quorum majority. A retried purge whose lookup reports the version missing re-reads every drive. Fewer than
N/2+1absent replies fail with insufficient write quorum, and a set with any remaining copy is scheduled for MRF healing. Offline, corrupt or unreadable drives never count as absent. The pools layer applies the same check to pools omitted by lookup, and replica receivers resolve purges by version across pools. - Healed markers keep their stored metadata. Healing copies the marker’s persisted map and drops only operation-time healing, movement and tier keys. Typed state is used only for creation writers without stored metadata, and no timestamp is invented.
- Creations are re-checked under the lock. Before sending, the delete replication worker re-reads the source version. A missing version, a non-marker version or a version under purge makes the task stale, and it is dropped without touching the source. A read that cannot confirm either way is retried through MRF instead of being treated as absence.
Evidence and boundaries
Two three-site batched-recovery rounds (source restarted and not restarted, one target site returned in two halves) converged within about 15 s and stayed stable for the rest of a 450 s window. The stale-creation sequence reproduced on the combined build; the new regression fails on unmodified main and passes with the repair on both fixtures.
The re-check cannot intercept a creation already on the wire or one replayed by another site. Closing that needs a receiver-side persistent purge record or sequence, tracked in #217. After a majority-acknowledged purge followed by the loss of every process, the minority copies on the drives that had been offline remained through 450 s and seven scanner cycles per site. All frontends read the data version, but no persistent owner of that cleanup is proven. These two boundaries are inherited behavior and were not made a condition of the 2026-09-16 release.
10 - Replicated Tag Ordering: Revision Timestamps, Tombstones, and Resurrection
Publication update, 2026-09-17: The September source repairs discussed here shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
This page records the analysis and repair of two defects in how object tags
keep their ordering across replication, merged into Server main as
PR #193 (fix
03027727d) and
PR #196 (fix
680eac66e).
As of 2026-09-16: both fixes are on verified main
40220bd836cb. They are not in the published Server 20260903; use the linked PRs to identify a build containing them.
Delivery boundary: source acceptance (regression tests plus the R4–R8 integration run whose PR #196 checks passed) only. No tag, package, image, or production rollout is established by this record.
Evidence class: synthetic signed HTTP tests use real single-drive, 16-drive and multi-pool erasure backends; sender, wire-shape and precondition checks also include function-level tests. No customer incident is attributed to these paths.
The shared model: a tag value and its revision are one state
Tags replicate with an internal revision timestamp
(wire header X-Minio-Source-Tagging-Timestamp, stored as x-minio-internal-tagging-timestamp). The ordering rule on the receiving
side is simple: a replicated tagging state wins only when its revision is
newer than the stored one. Both defects in this family break that rule by
making the revision — not the value — the part that gets lost:
- R4 dropped the timestamp on the way into an SSE-KMS destination, so newer tag updates lost to stale stored state inside the storage-layer reconciliation.
- R5 meant an empty value (a deletion) carried no revision at all, so the protocol could not express “deleted at time T” — and a delayed event could resurrect what a client had already deleted.
R4: SSE-KMS destination copies dropped tag revision timestamps
Failure form. A trusted replication COPY to an SSE-KMS encrypted
destination returned 200, the object was correctly encrypted, plaintext GETs
worked — and the destination’s tags and their timestamps stayed at the old
values. Because the HTTP request succeeded, nothing surfaced the loss. The
storage-layer reconciliation (reconcileStoredObjectTags) compares revisions
under the object write lock; without the incoming timestamp, the old stored
state won.
Trigger surface. Not only explicit SSE-KMS headers. Bucket-default KMS and global automatic encryption hit the same code path, so the defect could fire with no KMS header in the request at all.
Root cause. The option builder for PUT-like requests parsed the trusted
source-tagging timestamp into ObjectOptions, then the SSE-KMS branch
constructed and returned a different ObjectOptions carrying mtime, ETag,
replication trust, and two Object Lock timestamps — but not
ReplicationSourceTaggingTimestamp. The omission dated back to upstream
c4373ef290 (2021); a 2026 Object Lock repair added two more timestamps to
that literal and still missed this one. Before R5, the consuming path was COPY ordering. R5 adds timestamp persistence for replica PUT and multipart initiation, as well as the duplicate-precondition consumer. Those SSE-KMS paths also depend on R4 preserving the option, so backports must consider the pair together.
Fix. One field added to the existing SSE-KMS ObjectOptions literal
(03027727d), nothing
else. Regression tests cover every destination encryption (none, SSE-S3,
SSE-KMS with and without key context, SSE-C) crossed with trusted/untrusted
source, missing/valid/malformed timestamps, and 50 signed ordered COPY+GET
cycles across two single-pool backends (single-drive and 16-drive erasure)
with 1–3 ns event spacing. The KMS cases use a test stub.
No backfill. A lost source-tagging timestamp cannot be reconstructed at the destination. After upgrading, new tag events replicate in order; replay of old events still follows the timestamp comparison: an incoming event must be strictly newer; the stored value wins ties and rejects older events.
Deferred observation. The same SSE-KMS literal also omits the proxy and speedtest option fields; the speedtest flag is read on the storage path, so global auto-encryption would drop it on a speedtest PUT. Recorded here as a separate follow-up, without asserting a public issue exists; deliberately not bundled into this repair.
R5: empty tag values had no revision, so deletions could resurrect
Failure form. Nine baseline regressions across real-storage and function-level tests:
- A successful
DeleteObjectTaggingnever minted a new revision, so a later-arriving trusted metadata COPY with an older view of the tags re-instated them. - A newer COPY carrying an empty tagging state was ignored — empty meant “nothing to say” instead of “deleted”.
- The first replica PUT parsed the source tag timestamp and never persisted it.
- Equal visible tag values collapsed to “no replication needed”. HEAD does not expose a tag revision, so a newer deletion or re-addition was invisible and ordering was not restored.
- A queued replication event’s completion callback wrote the old tags from its snapshot back over an already-committed deletion — and without a timestamp, the resurrected set inherited the deletion’s newer revision, which is worse than the originally-reported symptom.
Root cause. A tag value and its timestamp (including the empty value’s timestamp) constitute one state. The old protocol could only represent non-empty states: DELETE-tagging never minted a revision, the sender only attached timestamps in the non-empty branch, and the receiver only made decisions in the non-empty branch.
Fix (680eac66e):
- Every tagging mutation mints one revision. PUT/DELETE tagging handlers unconditionally stamp a single UTC RFC3339Nano revision, whether or not replication selects the object. The storage layer enforces monotonicity under the write lock: a local revision that is not strictly newer is advanced to stored+1 ns, and multi-pool backends compute one value that strictly exceeds every pool’s copy.
- The sender transmits tombstones. Empty values carry their recorded revision; empty-without-revision is not fabricated; a malformed recorded timestamp fails closed rather than being silently repaired.
- The receiver accepts tombstones. Replication COPY captures the stored tag pair before rebuilding, so an empty value with a timestamp enters the existing reconciliation as a state that can win. Duplicate suppression is relaxed only for a strictly newer source revision. Multipart completion orders tags from the persisted upload metadata. The delete-acknowledgement path no longer writes snapshot tags back. Ordinary SSE-C rotation drops the old tag timestamp from copied encryption metadata so it cannot overwrite the new local revision.
Operator-visible changes:
- Equal timestamps now resolve to stored-wins on unqualified and explicit null COPY requests. The old handler let the incoming state win ties on the unqualified path; the change is a compatibility-visible tightening in the correct direction.
- An ordinary COPY with
tagging REPLACEand no tags now genuinely clears destination tags. The old default-metadata path carried source tags over — an S3 consistency improvement, but a behavior change. - Every ordinary COPY (including key rotations that change nothing) records one new local revision.
- Both ends of a replication pair must upgrade together. An old peer still drops empty-value revisions, so a new sender’s tombstones are invisible to it.
- Objects with recorded revisions incur one extra metadata COPY per object during explicit resync/heal. When the destination has bucket-default or automatic KMS encryption, that “metadata” COPY rewrites the object data — budget accordingly for bulk resync.
A nonempty legacy tag set without a revision uses object ModTime as its sender fallback; an empty set without a revision does not acquire a fabricated tombstone. A strictly newer trusted tag revision bypasses only the internal ETag/version duplicate guard (otherwise 412 PreconditionFailed); client If-Match and If-None-Match remain enforced.
The tag repair adds no wire field or storage format and has no capability negotiation. Both endpoints are needed for ordered deletions; an old hop retains the old behavior. This statement does not authorize rolling upgrade or downgrade of the entire September candidate, which also contains the separate IAM migration. Reconcile already-damaged tags from an authoritative source by a new explicit tag change; the lost historical order cannot be reconstructed.
Limits, stated as limits:
- No migration: history with missing or wrong deletion timestamps cannot be reconstructed, and no historical tombstones are fabricated.
- Replication rules with tag filters evaluate target eligibility against the post-deletion (empty) tagging state, so a rule filtered on the deleted tag never sees the deletion. This is a pre-existing scoping decision, unchanged here.
- Arbitrary clock skew is not a total order. The repair establishes per-hop ordering, not multi-site causality.
- A malformed recorded timestamp makes the sender’s construction fail permanently and the event retry through MRF until an explicit, correct tag change replaces it — deliberately fail-closed.
Rejected alternatives
- Synthesize a tombstone from ModTime for empty-without-revision objects. Every object that never carried tags would gain a revision; combined with forced metadata replication, every object would take the metadata-COPY path on every hop.
- Transmit only when the value is empty. Withdrawn by its own proposer
during review: the equal-value re-add sequence (set X at T1, delete at T2,
re-add X at T3) loses the re-add’s revision under that condition. A
regression test (
TestTaggingRepeatedValueNeedsRevisionDelivery) pins the counterexample. - A new HEAD revision protocol. The pinned
minio-gometadata extractor discards internal response headers, so this needs a new wire contract for marginal benefit; the worst case is still a forced metadata COPY. - A distributed causal clock, or regenerating timestamps at commit. The wall-clock model plus the in-lock monotonic guard is the minimal correct fix.
- Per-pool monotonic guards only. Ordinary reads return single-pool object info, so a sender could emit a stale primary-pool revision; the unified multi-pool value is required.
Verification and what it does not prove
R4 regressions cover the encryption × trust × timestamp matrix and the ordered sequence on two single-pool backends (single-drive and 16-drive erasure), with a stub KMS. R5 separately has a multi-pool tagging-deletion regression. The R4–R8 integration run repeated targeted tests on the merged tree, including the isolated R5 multi-pool test; all eleven PR #196 checks passed. These establish the tested ordering cases. They do not establish multi-site scheduling under real clock skew, cross-region failover, or behavior of deployments that upgrade one end of a pair only.
The upgrade summary and release boundary live in the component matrix; the sibling repair that stops normalized replica metadata from being re-injected is recorded separately in Replica Metadata Normalization.
11 - SILO Server 20260903 Pre-release Review
This is the durable pre-release engineering record behind SILO 20260903. It explains why an earlier “all issues are solved” assessment was not accepted at face value, what the independent review found, how the fixes were narrowed, and which gates still remained at the review point. The linked release note records the later publication result.
Decision: the source candidate at
6e112d1856d4f3655f30fc81ee47e9f43d50d8f3is a code-level GO for remote review. Production release remains a conditional GO until remote CI, Test Release, tag and artifact verification, signing, container publication, and public pull checks complete.
Baseline:RELEASE.2026-08-06T00-00-00Zat3be10fcc1a44f6620ded0bd303461f9d688cca23.
Scope: SILO Server behavior and its embedded/pinned runtime components. Documentation, the standalone Console, mcli, package repositories, images, and the deployed site are separate deliverables.
Publication closure: the later final tree9b11dc9469e650815b775cb47b039610644f5da4was published asRELEASE.2026-09-03T13-18-01Zon 2026-09-04 after the remote, package, provenance, container, and public-download gates below completed. The conditional decision in this page remains the historical review criterion, not the current release state.
2026-09-09 follow-up: the repairs and main validation for #153, #152, and #137 are recorded separately in Replication Reliability. That record preserves #136’s accounting contract, explains cancellation lifecycle decisions, and keeps #133 separately tracked. It does not change the historical assessment of the 0903 release candidate below.
Why the second review was necessary
The first implementation pass had strong test results and resolved most reported defects. Its conclusion was nevertheless too broad: it treated green tests and a clean worktree as proof that every security invariant had been closed.
An adversarial review asked different questions:
- Can the same invariant be bypassed by a different valid wire representation?
- Does a pre-authentication fast path still perform I/O or acquire state?
- What happens when metadata exists but cannot be loaded?
- Do two individually correct read-modify-write paths share the same serialization boundary?
- Does request sanitization preserve all SigV4 streaming state?
- Does a validation claim describe the final tree or an earlier one?
- Is a complex mechanism protecting a reproduced failure, or only a hypothetical future?
That pass found real defects after the initial “ready” claim. The correct response was not to distrust all prior work, but to narrow every assertion to an invariant and an observed tree.
Review result by area
| Area | Adversarial finding | Final resolution | Status |
|---|---|---|---|
| Bucket metadata | Independent config locks could lose updates to the shared .metadata.bin record (#102) |
One bounded metadata.lock surrounds every whole-record writer, migration, import, adoption, and healing path; changed-field replication avoids stale whole-record replacement |
Closed in candidate |
| Bucket creation | ForceCreate and site adoption could replace existing config with defaults |
Preserve existing records and update only creation/adoption state; add regression tests for clobbering | Closed in candidate |
| Object Lock | Comparing lock-document bytes with one canonical XML document missed valid configurations carrying a Default Retention rule | Parse Object Lock first, then derive the versioning invariant from the parsed enabled state; verify update, read-back, and disk reload | Closed in candidate |
| Pre-auth CORS | Arbitrary path segments could cause metadata reads and cache growth | CORS lookup reads resident metadata only and does no object-layer I/O | Closed in candidate |
| CORS startup | A nonresident name could fall back to global CORS before metadata initialization | Preserve an explicit fail-closed startup state | Closed in candidate |
| CORS load failure | Forgetting that a real bucket failed to load made it indistinguishable from a nonexistent bucket and exposed the global fallback to pre-signed requests | Maintain a bounded failed-bucket set, clear it on every successful load/remove/refresh path, and keep those buckets fail-closed | Closed in candidate |
| CORS recovery | A successful on-demand GetConfig reload did not initially clear the load-failure bit |
One-line final fix 84e1580a4 plus targeted race coverage |
Closed in candidate |
| Replication trust | Presence of client-controlled internal headers enabled privileged behavior in multiple handlers | Authenticate first; require an exact marker plus s3:ReplicateObject or s3:ReplicateDelete; carry a private context decision; sanitize untrusted headers afterward |
Closed in candidate |
| Streaming uploads | The sanitized request clone did not initially share the original trailer map | Preserve the trailer map so late-arriving streaming checksums remain visible | Closed in candidate |
| Snowball | A request-wide trust bit could leak between extracted entries | Derive and isolate trust per entry; preserve request defaults across workers | Closed in candidate |
| SSE-C | Zero-byte reads and GetObjectAttributes could skip customer-key authentication |
Require a successfully unsealed key, with a separate authorized-replica exception | Closed in candidate |
| Delete authorization | Explicit version deletes checked the ordinary delete action instead of requiring s3:DeleteObjectVersion |
Align single and multi-delete authorization, keep replication deletes on s3:ReplicateDelete, and preserve auth/audit context |
Closed in candidate |
| Admin authorization | User/group status changes always checked the enable action | Check the action that matches the target state | Closed in candidate |
| Checksums | Multipart and copy paths omitted fields, accepted invalid combinations, or computed over the wrong representation | Complete algorithm/type validation, server-side part calculation, federated propagation, AWS errors, and CopyObject transform ordering | Closed in candidate |
| Release evidence | Full acceptance initially described a tree that changed afterward | Record full acceptance at ebac0ca73 and current-tree targeted gates separately |
Closed as an evidence defect |
The invariants that now define the candidate
Trust is derived once, after authentication
An internal-looking header is still client input. The request must first pass the existing authentication path in its original signed form. Only then can the handler combine:
- an exact, single replication marker;
- a non-anonymous authenticated identity;
s3:ReplicateObjectors3:ReplicateDeleteon the addressed resource;- replica status where the narrower replica-only semantics require it.
The result lives in private request context. Header stripping is defense in depth for legacy consumers, not the source of authority.
This ordering matters because SigV4 may sign the headers. Sanitizing first would reject legitimate replication with SignatureDoesNotMatch. The sanitized clone also has to share the request trailer: trailers arrive after the initial header parse and carry streaming checksums.
The complete receiver-wide model is in No I/O Before Auth, No Privilege From Headers.
A shared record has one write boundary
Policy, lifecycle, SSE, tags, quota, replication, Object Lock, versioning, and CORS are logical fields but physical members of one bucket record. A per-field mutex cannot protect a whole-record read-modify-write.
The selected repair is deliberately smaller than a new database or transaction layer:
The lock does not cover object data I/O and is bounded to a bucket-metadata operation. Migration and healing must participate because they also replace the whole record. Replication receivers merge only changed fields so an older remote snapshot cannot erase unrelated local state.
Failure is a state, not the same thing as absence
The CORS hot path must distinguish four states:
| State | Result |
|---|---|
| Metadata system not initialized | No CORS headers |
| Known real bucket whose metadata load failed | No CORS headers |
| Resident bucket with a bucket CORS document | Evaluate that document |
| No resident metadata and no known failure | Use the server-wide fallback |
The second row is why a failed-bucket set survives the simplification pass. A pre-signed URL is already authorized by its signature and may access a private object without bucket-policy evaluation. In that case the bucket CORS document is the browser-origin boundary. Losing the failure bit and using a permissive global fallback would weaken that boundary.
The set remains bounded by real bucket load attempts and is maintained through two helpers. Successful load, removal, stale-bucket cleanup, refresh, reset, and concurrent load all have tests.
Object Lock is semantic, not textual
Any valid enabled Object Lock configuration implies versioning. XML whitespace, element order, and the presence of a Default Retention rule do not change that meaning. Therefore normalization follows parsing, not a byte comparison against one canonical document.
The resulting versioning record is plain Enabled. A suspended state and an exclude-prefix extension are incompatible with the lock invariant and are removed on update, read-back, and reload.
Complexity audit
The pre-release pass explicitly looked for over-design, duplication, defensive programming without a threat model, and stale compatibility machinery.
Complexity retained because it protects a reproduced failure
- One metadata lock: retained because a deterministic cross-type lost-update test reproduced data loss.
- CORS tombstones: retained because site replication cannot distinguish deletion from “never observed” without them.
- CORS load-failure state: retained because a pre-signed URL provides an authenticated, policy-independent counterexample.
- Two replication trust levels: retained because ordinary replication and replica-ciphertext/SSE semantics do not use identical wire shapes.
- Post-authentication sanitization: retained because sanitizing before SigV4 verification breaks legitimate signed requests.
- Adversarial multi-pool/null-version tests: retained because single-pool happy paths do not exercise the state-selection failures they caught.
Complexity removed or narrowed
- CORS failure-set mutations were centralized in
noteLoadFailureandclearLoadFailure. - The replication import path now applies changed fields rather than copying an entire possibly stale record.
- Obsolete encryption helpers, dead event-target functions, and abandoned handler branches were deleted.
- The compatibility guard stopped inventorying every exported source symbol and now protects the actual served routes and frozen wire/configuration surfaces.
- The old
wait_pipelint exemption was removed;gomodguard_v2replaced deprecated configuration. - Dynamic timeout tests no longer call global
rand.Seedfrom a parallel package. - The server returned from the temporary
silo-gofork to the reviewed upstream-compatibleminio-gorevision.
Changes deliberately not introduced
- no generic metadata transaction framework;
- no second CORS cache or unbounded negative cache;
- no new public “trusted replication” request header;
- no cross-repository release gate that makes the server depend on a later Console or documentation release;
- no partial conditional-delete contract in the release candidate;
- no broad rewrite of inherited site-replication registers without dedicated convergence tests.
Deferrals and why they do not all have the same severity
| Item | Classification | Release decision |
|---|---|---|
| Conditional delete #10 | Inherited missing S3 feature; dangerous only to callers that assume unsupported If-Match / per-object ETag is enforced |
Document prominently; do not merge the incomplete PR or a single-only half contract |
| Multi-site config deletion #77 | Inherited convergence defect for policy/SSE/tags/quota; CORS has its own fixed register | Not a single-site blocker; deployment condition for users relying on those multi-site deletes |
ListMultipartUploads #79 |
Inherited listing-conformance gap | Known issue; not a data-integrity blocker for ordinary multipart workflows |
Federated CopyObject #99, #100 |
Legacy-backend checksum/inline-object gaps | Block use of the affected features, not the general server release |
| ILM relocation PR #60 and broad SSE issue #61 | New capability requests | Outside the release safety boundary |
“Inherited” does not mean harmless. It means the defect was not introduced by this change set and should be evaluated against the documented release contract. A deployment that depends on one of the affected paths inherits a deployment-specific stop condition even when the general release remains conditional GO.
Evidence
Full acceptance tree
The full local acceptance corresponds to ebac0ca73bbf251b070bb6df4d8005015841f901:
- full
cmdandinternalsuites; - complete
cmdrace suite: 365.448 seconds, pass; - lint: 0 issues;
- rebrand/compatibility and generated-file guards;
govulncheckwith no reachable vulnerability;- six
make verifydeployment shapes: 174 PASS / 0 FAIL.
The first two make verify attempts encountered environment/setup failures while obtaining mcli, not test failures. The successful run used the locally checksum-pinned mcli, retained the outbound proxy for GitHub downloads, bypassed it for localhost, and placed GNU userland tools first in PATH. That distinction is part of the evidence rather than something to hide.
Post-acceptance candidate
The only code change after that full run is 84e1580a4, which clears one CORS failure-state bit after a successful on-demand metadata reload. The candidate merge adds no code; 6e112d185 changes only Helm release metadata and documentation. On the final candidate, the following pass:
git diff --check;- targeted CORS and Object Lock
go test -race; - rebrand guard;
- generated-file check;
- lint with 0 issues.
- Helm lint, default and optional renders, chart packaging, and the seven-resource legacy-upgrade identity guard.
This evidence is proportional to a one-line state-transition fix, but the remote CI and release workflows must still run against the pushed tree.
Go, no-go, and ownership of the remaining gates
Code decision: GO
No confirmed code defect from the two review rounds remains unresolved in the candidate. The fixes are covered at the layer where their invariants live, and the retained complexity corresponds to reproduced counterexamples.
Production decision: conditional GO
The server must not be described as released until all of these are facts:
- candidate commits are pushed and reviewed;
- remote CI and Test Release pass on the pushed head;
- the intended tag points at the reviewed chart 7.0.2/server 0903/client 0903 release tree;
- Draft artifacts, checksums, SBOMs, attestations, and signed RPMs verify;
- finalize and Docker release publish both classic and distroless variants;
- anonymous download and pull tests pass;
- release notes are updated from the tagged facts and the documentation site is deployed.
Any failure in steps 1–6 is a release blocker. A local green suite cannot substitute for them.
Deployment-specific stop conditions
Operators should delay even a successfully published release when they cannot yet:
- update every node in a distributed cluster within one coordinated maintenance operation;
- update every member of a site-replication group before using bucket CORS;
- revise IAM policies for
s3:DeleteObjectVersionand status-action separation; - avoid or explicitly accept the known #10, #77, #79, #99, or #100 path their workload depends on.
The final conclusion at the review point was intentionally narrower than “everything is fixed”: the reviewed candidate was ready to enter the release machinery, the remaining limitations were explicit, and production publication was gated by verifiable artifacts rather than confidence. Those gates later completed for the release linked above; these deployment conditions describe Server 20260903, the release this record reviewed.
Restart and readback verification (2026-09-11, #116)
A later acceptance (#116, run on 2026-09-11) measured bounded restart/readback behavior, including the inherited readiness gap. The durable part is the method, which any operator can reuse when validating a restart or upgrade window:
- Acknowledgement ledger. Every acknowledged PUT writes to a unique versioned key, and the acknowledgement records the VersionId, byte count, and SHA-256 immediately. Readback fetches by exact VersionId and asserts all three — an early confirmation cannot be silently replaced by a later write.
- Readback timing. Periodic re-reads happen at 15/30/60 s after the data canary succeeds, through every peer, and the final check includes writes made during a single-node outage and after its rejoin.
- No retry masking. Each canary carries a hard 60 s deadline covering setup, request, response-body read, and sleep, and SDK retries are disabled so a recovery window cannot be papered over by client-side retries.
- Driver topology. Four Linux/arm64 containers on one host with independent network identities, one drive each (EC 2+2), and tmpfs volumes kept mounted by a holder container through the full shutdown; nodes stop in parallel with a 10 s grace period, then start in parallel.
The operational finding worth remembering: admin-endpoint readiness is not data readiness. The fixed historical record tested 0806, 0903 and the then-current main repair build. The main build reached the admin gate 2.461 s after restart, then needed another 14.489 s for the data canary: different timing origins. The 0903 run took 0.191 s after its gate, which does not establish immunity to the inherited startup window. A readiness probe against admin/health says nothing about the data plane during that window, and no fixed sleep substitutes for an actual data-plane check.
Boundaries, stated as boundaries: this acceptance covers process/container restart and TCP peer reconnection on a single Linux host. It does not prove persistence across independent hosts, host reboots, or physical media failure, and the timings are individual observations, not latency guarantees. The run artifacts are retained outside the documentation tree; the method above is the part that generalizes.
September 16 source follow-up: #10 closed through the independent single-object repair #145; batch ETag deletion is still absent. #77 and #99/#100 were repaired on main. #133/#144 were subsequently addressed by #178; #79’s default listing limitations remain open. These changes must not be retroactively attributed to Server 20260903; they shipped in Server 20260916. See the current component/source matrix.
12 - Replica Metadata Normalization: What a Trusted Copy May Not Re-Inject
Publication update, 2026-09-17: The September source repairs discussed here shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
This page records the repair of how a trusted replication receiver restores
metadata for a replica, merged into Server main as
PR #194 (fix
4fcdf37ce, merged as
9f3037e941).
As of 2026-09-16: the fix is on verified main
40220bd836cb. It is not in the published Server 20260903.
Provenance: the production logic follows PR #187 by Mikhail Khadarenka; the merged change keeps that authorship and narrows it to a review-validated boundary.
Evidence class: an HTTP-level baseline of 64 leaf cases (44 controls passing, 20 defect failures before the fix) against real single-drive and 16-drive erasure backends, plus counterfactual replay of the same suite against the baseline helper. No customer incident is attributed.
What went wrong
The ordinary PUT path normalizes metadata: it strips the transport-only
aws-chunked token from Content-Encoding, and removes the
X-Amz-Meta-X-Amz-Unencrypted-Content-Length/-Md5 user-metadata keys that a
GHSA-76wf-9vgp-pj7w mitigation deliberately deletes. The trusted replication receiver,
however, restored replica metadata by re-running the same permissive
extractor with replication allowed — replaying every supported header and all
user metadata from the original request. Concretely, on a trusted replica
write the server could store and later return via GET/HEAD:
Content-Encoding: aws-chunked(a pure transport encoding that must never be stored, per the AWS SigV4 streaming rules), or the unsplitaws-chunked,gzipstring instead ofgzip;- the two GHSA-redacted user-metadata keys — a partial rollback of that mitigation, limited to trusted replica writes;
- for Snowball entries without their own PAX header: the outer archive’s content-type, cache-control, and user metadata.
The object bytes themselves were not necessarily damaged; the stored metadata
was wrong. The regression was introduced by
56fa63bfd
(2026-04-15, the replication header trust boundary hardening, CVE-2026-34204)
— whose trust protection is correct and stays.
The fix
One file (cmd/handler-utils.go). The boolean dual-mode helper is deleted:
- the ordinary extractor unconditionally skips replication-only keys;
- a new replica extractor walks only the replication-to-internal header map and restores only the six replication-scoped fields: the SSE-C sealed key material, sealed algorithm, IV, and encrypted-multipart marker (the empty marker is honored by key presence), the actual object size, and the SSE-C checksum identity mapping;
- it never re-reads ordinary supported headers or user metadata.
Expected stored encodings after the fix:
| Requested encoding | Stored Content-Encoding |
|---|---|
aws-chunked |
none |
aws-chunked,gzip |
gzip |
gzip |
gzip |
aws-chunked, gzip (note the space) still stores gzip with a leading
space, and gzip, aws-chunked stores the whole string. These are
documented status quo, asserted by tests as such — not claims of repair.
Operator-visible changes
- Trusted Snowball entries no longer inherit the outer archive’s ordinary metadata. Without PAX records, this removes outer content-type, cache-control, expires and user metadata; with PAX records it removes fields the entry did not restate. Per-entry
minio.metadata.*still applies, the six replication-scoped fields still apply to authorized entries, and the archive’s storage class remains inherited. The in-tree batch producer callsPutObjectsSnowball, whose SDK emits the auto-extract marker, but it does not mark the outer request as a trusted replica; it did not use the affected inheritance path. - The GHSA-redacted keys are no longer written back on replica restore — matching what every ordinary PUT already did.
- Authentication, permission gating, and replication trust semantics are unchanged; the ordinary extraction path is byte-for-byte equivalent.
- Rolling back the code reopens the injection path but does not repair already-stored metadata.
Upgrade does not fix stored objects
An upgrade stops new pollution; it does not scan or rewrite existing objects. Two consequences matter:
- Objects whose authoritative source remains polluted can be selected repeatedly for metadata replication when comparison with a normalized replica detects a difference. Fix the authoritative source first, then let the copies converge.
- Ordinary S3 self-COPY is not a general remediation API: it can create new versions or shift timestamps rather than rewriting one version’s metadata in place.
The read-only audit runbook now provides an executable inventory tool and classification rules. It does not authorize or perform repairs.
The stored-metadata remediation proposal — status
A design for a future operation exists: build an inventory (including
non-current versions, not just the latest), verify by comparing against the
trusted source version or an independent checksum — never by guessing from
the wrong response header, and never by re-decompressing gzip just because a
label says so — process the authoritative source’s exact versions first,
then converge copies; keep immutable inventories and metadata backups; use
small validation batches with a rehearsed rollback. Where no supported path
exists for an object, stop and leave it — editing xl.meta directly is not a
supported operation.
Any selected procedure must preserve the required version identity and current version relationship, Object Lock retention/legal hold, tags, replication state, and encryption context. Check for concurrent changes before writing; blocked or unverifiable versions stay untouched. A local clone must prove the chosen operation and rollback before this proposal becomes an executable runbook.
This is a design proposal awaiting separate approval, not an executed procedure. No production inventory scan, object write, version change, or deployment has been performed as part of it. Treat it as the shape of a future runbook, not as a validated one.
Known limits
- The POST-form upload path (
bucket-handlers.go) calls the low-level extractor directly and never normalized encodings; that behavior is unchanged and flagged for a separate issue. - Local verification ran with a test-only capacity accommodation (a full host disk); the merged-tree rerun in the R4–R8 integration record covers the unchanged-tree case.
- No two-site scheduler, restart, or network-failure acceptance is claimed.
Verification
Regression tests (TestExtractReplicationMetadata*,
TestAPIReplicaContentEncoding, TestAPISnowballReplicaContentEncoding,
plus race-included trust/SSE-C round-trips) cover the mapping table, the six
restored fields, and the ordinary path’s equivalence; the recorded counterfactual run against the baseline helper produced 36 expected failures (20 HTTP cases and 16 helper cases), with 44 controls passing. These are the original repair’s observations, not a new execution by this documentation update. The upgrade summary lives in the
component matrix; the
sibling tag-ordering repair is recorded in Replicated Tag
Ordering.
13 - Request-Header Deadlines: Absolute Header Limits and Rolling Body Idle Timeouts
Publication update, 2026-09-17: The September source repairs discussed here shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
This page records the repair of Server’s HTTP read deadlines, merged into
main as part of PR #196 (fix
055030ea5).
As of 2026-09-16: the fix is on verified main
40220bd836cb. It is not in the published Server 20260903.
Evidence class: synthetic — a direct TCP comparison (header limited to 100 ms, header arriving over 400 ms, standard Go refuses where old SILO returned 204), plus a real single-drive process probe where flag and environment settings both rejected a 400 ms slow header while healthy requests continued. No production incident is attributed; the issue that motivated the investigation is a slow-HTTP DoS scanner report that has not been reproduced against a deployed cluster.
Two timeout classes, one connection
- Request-header absolute deadline.
ReadHeaderTimeoutbounds the total time from the start of header reading to its completion. Trickle-feeding bytes cannot extend it. On HTTP/1 keep-alive connections, Go first usesIdleTimeoutwhile waiting for the next request’s initial bytes, then starts a fresh header deadline.ReadHeaderTimeoutalso participates in the separate TLS-handshake timeout calculation. - Body rolling idle timeout. Once headers parse and the connection enters the active phase, the existing rolling semantics return: every successful read extends the deadline, and only a stall between bytes (the configured idle timeout) kills the connection. A long upload or download that keeps making progress is not capped in total duration by this HTTP/1 repair; other protocol, proxy, and application timeouts still apply.
Before this repair, the first class did not exist in practice: the
connection-layer wrapper replaced the socket deadline with
now + idle + 250 ms before every partial read, overwriting whatever
absolute deadline net/http had set — so a slow reader could hold a
connection open indefinitely by sending one byte per idle window.
Two independent defects
- The connection layer neutralized the absolute deadline. The
DeadlineConnwrapper’s read path reset the socket deadline on every partial read, defeating the read-header deadline Go’s server sets. The direct-TCP baseline proved it in isolation: with a 100 ms header limit and a 2 s idle window, a header that finishes at 400 ms was accepted. - The configuration never reached the server. The CLI accepted
--read-header-timeoutandMINIO_READ_HEADER_TIMEOUT, parsed defaults and all — and the server-context builder copiedIdleTimeoutwhile droppingReadHeaderTimeoutentirely, so the running server always saw zero. Fixing defect 1 alone left the real process accepting slow headers; the second fix is one line next to the idle-timeout binding.
The reason this stayed invisible for so long: the flag’s default (30 s) equals the idle timeout’s default, and with the flag unwired the server fell back to exactly that same 30 s — so every observable default behaved as if configured.
Configuration
- Flag:
--read-header-timeout(Hidden: true, absent from ordinary CLI help) - Environment:
MINIO_READ_HEADER_TIMEOUT - Default: 30 s (equal to the idle timeout default)
- There is no YAML configuration field for either timeout; the value binds once at startup from flag > environment > default.
| Setting | Effect |
|---|---|
| header > 0 | Absolute cap on HTTP/1 header phases; participates in Go’s TLS-handshake read window (including the HTTP/2 handshake) |
| header = 0 (explicit) | Falls back to Go’s rule: the read timeout (= idle timeout) applies; the CLI default is 30 s |
| header < 0 | Disables the header-specific cap. Positive read/write timeouts still bound TLS handshake reads, and positive IdleTimeout still bounds the keep-alive wait. This does not disable every connection timeout. |
| idle shortened, header unset | Header phase independently uses the 30 s default — the one combination looser than a naive expectation, though still strictly tighter than the pre-fix unbounded extension |
A negative value reopens unbounded slow-header trickling; it is not a recommended compatibility setting. An incomplete header cut off by the deadline normally sees a closed connection, not a guaranteed HTTP error status. The trigger was @AEGEGE’s scanner report #183; PR #195 was integrated through #196. The experiment does not establish reproduction in that deployment.
What each protocol gets
- HTTP/1: headers and keep-alive waits are absolute; the body keeps the rolling idle timeout. The connection-state hook composes with (rather than replaces) any caller hook.
- TLS: handshake reads take the minimum of the positive header deadline and the existing read/write timeouts; after the handshake, a fresh header limit begins. The handshake’s write side remains rolling — this repair is not a complete TLS-handshake resource limit.
- HTTP/2: untouched. When h2 is negotiated, the phase switching is
skipped entirely; h2 keeps its own native per-stream read timeout, which is
absolute, and
ReadHeaderTimeoutnever enters the h2 configuration. - Internal callers: Linux internode dialing uses its own rolling semantics; grid-hijacked connections unwrap to the raw TCP connection before any of this applies.
Rejected alternatives
- Clamp all future deadlines globally. Go 1.27 sets a whole-request deadline in some paths; with the read timeout equal to the idle timeout, this would hard-cap entire HTTP/1 requests — header plus body — and kill every large upload.
- Drop the read timeout and reinterpret zero as rolling idle. Zero is
net/http’s “never time out” for background reads and hijacked connections; reinterpreting it would break long handlers, and h2 would lose its per-stream timeout. - Wrap the body reader / response controller. Full chunked/drain/EOF accounting with h2 special cases is a far larger change than the header defect requires. (A later, unmerged branch explores a body-side response controller for the same DoS family; as of this record it is not part of main and not part of this repair’s claims.)
- Reconstruct the standard library’s deadline arithmetic in the hook. Duplicates stdlib internals that drift between Go versions; remembering the value stdlib actually asked for is the robust form.
- Auto-derive strictness from value comparisons. With all three defaults equal at 30 s, “shorter than the idle window” is indistinguishable in production defaults; such logic only works in test configurations.
Verification and limits
Tests pin the connection wrapper across three consecutive update periods (no extrapolation of the absolute cap), the phase transitions over HTTP/1 keep-alive, TLS, and HTTP/2-only negotiation, and the flag/env binding on the real CLI context; a process probe exercised a live server with a 100 ms header limit rejecting a header that takes 400 ms. Known limits: the TLS handshake write side stays rolling; handler CPU/storage waits have no deadline; the absolute header cap carries no slack while the rolling idle keeps its ~250 ms update slack; and multi-node, cross-region long-transfer acceptance is future work — the integration record explicitly does not count a scripted S3 long transfer as passed for this repair.
The upgrade note (a shorter header timeout also narrows the TLS handshake window; it is not a total-duration limit for uploads or downloads) is in the component matrix.
14 - Conditional DELETE: One Condition, One Logical Object
Publication update, 2026-09-17: The September source repairs discussed here shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
Status, 2026-09-16: PR #145 added single-object conditional deletion; PR #178 subsequently repaired multi-pool serialization and reconciliation. Neither change is in published Server 20260903. Check the component matrix before relying on this behavior.
The August proposal around PR #12 was broader than the code that merged. In particular, its proposed batch rejection, extra read authorization and current-version-only rule are not implemented guarantees. This page describes the maintained source at f99ed829b.
Implemented contract
| Request | Current main behavior |
|---|---|
Single DeleteObject, nonempty If-Match: <ETag> |
Check the client-visible ETag under the deletion lock; mismatch returns 412 before deletion |
If-Match: * |
Require an existing, non-delete-marker representation |
Explicit versionId |
Evaluate the addressed version, not an unrelated current version |
No If-Match, or an empty header value |
No conditional callback is installed; ordinary deletion applies |
Batch DeleteObjects with per-item <ETag> |
The request model has no ETag field; the XML field supplies no deletion protection |
Internal recursive x-minio-force-delete |
Prefix deletion returns before the object-condition path; do not use it as conditional deletion |
These conditions do not add an s3:GetObject authorization check. Ordinary deletion authorization still applies, including s3:DeleteObjectVersion for an explicitly addressed version, Object Lock checks and the separate trusted-replication path.
Why the condition belongs above individual pools
An object can have copies of different ages in more than one pool. A request condition concerns the logical object selected for the operation. Evaluating and mutating independently in each pool can delete one copy and then return 412 from another, or delete the latest copy and expose an older one.
The multi-pool path therefore acquires the shared namespace write lock, gathers the relevant states, evaluates the condition once, clears the callback for lower layers, and reconciles the selected deletion across pools. The same lock boundary must cover concurrent writes and metadata updates. This is the purpose of the later multi-pool repair, not a claim that every physical disk is atomically updated.
Counterexamples from the early design
With old ETag A in one pool and current ETag B in another:
If-Match: Amust not delete A first and only then fail against B.If-Match: Bmust not remove B while leaving A to become visible again.
A per-pool HTTP callback also risks concurrent writes to one response writer. Consume the request condition at the coordinating layer instead.
Failure boundaries
The reconciliation path reads pool state before evaluating the condition and surfaces failures instead of treating an unknown pool as empty. Cleanup errors can still follow a partial physical mutation: a failed request is not a distributed rollback guarantee. Retrying and checking actual state remain necessary after a storage failure. See multi-pool consistency.
Selection and evaluation
Evaluate once
erasureServerPools.DeleteObject holds the outer lock. Multiple pools use deleteObjectReconciled; the single-pool path reads the selected representation and invokes CheckPrecondFn before its delete-marker shortcut. Lower layers do not reinterpret the condition per copy.
Wildcards and delete markers
The DELETE helper treats * as representation existence. A delete marker does not satisfy it. A missing key or addressed version follows the corresponding not-found path; it is not manufactured into an empty ETag match.
Authorization
The handler authorizes deletion using the effective version. It does not perform the additional s3:GetObject check described in the original proposal for a specific ETag. Do not treat the proposal’s permission matrix as implemented AWS parity.
SSE-C ETags
Conditional deletion compares the established client-visible ETag projection; it does not read or decrypt the object’s payload. SSE-C read-key authentication is a separate contract from deleting an object.
Explicit versions
An explicit versionId selects that version for comparison and deletion. A matching historical ETag can therefore allow deletion of the historical version even if the current version has another ETag. This differs from the early proposal’s current-version-only rule.
Unsupported edges
An empty If-Match installs no condition. The recursive prefix-delete extension bypasses object preconditions. Batch XML ETags are not recognized by ObjectToDelete; there is no request-wide NotImplemented guard. Callers needing compare-and-delete must use the supported single-object path with a nonempty condition and verify their selected release.
Alternatives and scope
Changing a shared ETag comparator cannot fix pool selection or mutation ordering. Aggregating errors after per-pool callbacks cannot undo deletions already performed. A new transaction framework is unnecessary for consuming one callback at the existing namespace lock, but batch execution and policy enforcement require separate implementation.
Evidence and release boundary
The source evidence is the handler, pool coordinator, and batch request model. Relevant coverage includes ETag mismatch, wildcard/delete markers, explicit versions, quorum failures and multiple pools. Earlier local review or test claims for a different implementation do not establish these missing guards. A merged implementation and a released, deployed binary remain separate facts.
Follow-up work
Batch conditions
Implementing per-item ETags requires parsing them, evaluating each logical object under the correct lock, preserving quiet mode and reporting each failed condition in the per-item response. Until then, a batch ETag is ignored and must not be used as a concurrency guard.
Policy enforcement
The maintained policy package does not define s3:if-match. Supporting it requires a package change and release, correct request condition values, a Server dependency update and authorization tests. Existing If-Match execution does not by itself provide policy enforcement.
Design cost
Single-object execution reuses the established object-selection and locking boundary. Full batch conditions and policy enforcement cross additional interfaces and remain separate work. The useful invariant is narrow: a false supported single-object condition must be decided before deletion, once for the selected logical representation.
15 - DSN-Only Database Notifications: A Compatibility Boundary for #53
Release check (2026-09-16): the original repair described here is included in Server 20260903. Dated review and test accounts below record their original evidence, not a still-pending release or acceptance of a particular production installation. Later source changes and component selections are in the version matrix.
This document is the product requirements and final design record for SILO issue #53. It records the accepted compatibility boundary, implementation, and verification for PostgreSQL and MySQL bucket-notification targets.
Decision
SILO will retain PostgreSQL and MySQL notification targets, but support exactly one current configuration form for each:
- PostgreSQL requires a complete
connection_string. - MySQL requires a complete
dsn_string.
The old five-field form — host, port, username, password, and database — remains unsupported by the current KV configuration system. SILO will not re-register those keys and will not synthesize a DSN from them during legacy migration.
The legacy migration contract is deliberately narrow:
| Legacy target | Result |
|---|---|
| Disabled | Ignore it; no target is emitted. |
Enabled with a non-empty connection_string or dsn_string |
Migrate only the canonical connection-string key and the other registered target settings. |
| Enabled with only discrete connection fields | Reject migration and abort server startup before the new configuration is activated, with an actionable error that names the subsystem and target but never prints a credential. |
This is a configuration-boundary decision, not removal of the database-notification feature.
Status: implemented in f1ba68358 and included in Server 20260903.
Owner: SILO server repository.
Tracking: pgsty/silo#53.
Target: the next SILO patch release after implementation and verification.
Context
SILO inherited two generations of database-notification configuration from MinIO.
The pre-KV JSON configuration could describe a database connection either as a complete string or as five fields:
The current KV configuration exposes only the driver-native form:
This direction is not new. MinIO deprecated the five discrete fields in RELEASE.2020-04-10T03-34-42Z and instructed operators to move to connection_string or dsn_string. SILO’s current help tables, environment-variable documentation, and examples already present the complete string as the supported interface.
SILO is a new community fork with an explicit migration step. Its compatibility contract prioritizes the S3 and Admin APIs, current MINIO_* settings, on-disk data, and current KV configuration. It does not need to perpetuate every pre-2020 configuration spelling when a supported canonical form has existed for years.
The defect
Before the fix, the legacy migration helpers, SetNotifyPostgres and SetNotifyMySQL, wrote both forms into the new KV configuration. Even when the old target already had a complete connection string, the helpers also emitted all five discrete keys, usually with empty values.
The new parser rejects those keys because neither DefaultPostgresKVS nor DefaultMySQLKVS registers them. Key validation checks key presence, not whether the corresponding value is empty. Both legacy source forms therefore fail:
The failure is amplified by notification initialization. FetchEnabledTargets is fail-fast across notification subsystems: the first invalid subsystem returns an error and a nil target list. The caller logs the error and continues starting the object server, leaving healthy Webhook, Kafka, NATS, and other targets unavailable as well.
Merely returning an error from the two migration helpers does not fix that behavior. The error propagates through readConfigWithoutMigrate and initConfig, but initConfigSubsystem currently logs non-retriable configuration errors as “some features may be missing” and returns success. The server then starts without assigning globalServerConfig; notification failure is only one consequence, because region, storage class, compression, identity, and other stored settings may also be absent. The implementation must therefore carry a typed database-migration error to the startup boundary and make that error fatal. Classifying it as retriable is also wrong because the server would retry forever without any state change that could repair the configuration.
The resulting behavior is especially dangerous because object I/O still works. Operators can see a healthy S3 service while every configured event pipeline has stopped. Targets are never constructed, so delivery or later replay of events produced during the outage must not be assumed.
There is also a diagnostic-exposure issue. The unregistered password key has no sensitivity metadata and may be copied verbatim into health or diagnostic material. The registered connection_string and dsn_string keys are already treated as sensitive values.
Why the first fix was reverted
The first repair registered the five discrete keys and taught the parser to read them. That made migrated targets pass CheckValidKeys, and it appeared attractive because the target argument structures and constructors still contain code for the old fields.
It also broke the documented connection-string path.
The shared mc admin config set tokenizer discovers field boundaries by looking for registered key names. It is not fully quote-aware. Once port became a registered key, this valid input contained what looked like a second top-level field:
The tokenizer split at the port= inside the quoted value, truncated connection_string, and handed the remainder to the port parser. The command then failed with invalid port.
Under the current tokenizer, registering common words such as host, port, and password creates a direct conflict between the connection-string grammar and the top-level KV grammar. The attempted registration fix was therefore reverted. Re-registering those keys is not an acceptable solution.
Product judgment
Database notification targets are a specialized but useful capability. They provide a direct database-backed namespace view or access journal without requiring an external event bus. That remains valuable for small deployments and for users already operating PostgreSQL or MySQL.
The legacy spelling of their connection parameters has much less value. A five-field model cannot represent the useful range of driver options: TLS modes and certificates, connection timeouts, application names, Unix sockets, multi-host PostgreSQL settings, MySQL driver parameters, and future driver capabilities. Supporting both forms also creates precedence, merging, redaction, and testing questions that do not exist with one canonical value.
The complete string is the better abstraction boundary: SILO owns notification semantics, while the database driver owns connection syntax.
The product decision is therefore to keep the capability and remove the compatibility illusion. An unsupported legacy target must be rejected clearly; it must not be accepted and transformed into a configuration that later disables unrelated targets.
Goals
- Establish
connection_stringanddsn_stringas the only supported live configuration interfaces for database notifications. - Allow a legacy JSON target that already contains the canonical string to cross the migration boundary without modification to its connection semantics.
- Reject enabled discrete-only legacy targets before a partial or invalid KV configuration is activated.
- Replace the current silent runtime failure mode of #53 — healthy targets disabled while the server appears healthy — with an explicit startup-time failure that operators must resolve before the server runs.
- Ensure no migration error, log line, health report, or diagnostic bundle exposes a database password.
- Remove the ten Postgres/MySQL exceptions from the source-level unregistered-write audit.
- Make the compatibility boundary and operator remediation explicit in release and migration documentation.
Non-goals
- Supporting both DSN and discrete database fields in the current KV interface.
- Automatically synthesizing a DSN from old discrete fields.
- Rewriting the shared KV tokenizer.
- Changing
FetchEnabledTargetsfail-fast semantics in this patch. - Silently skipping an enabled database target and continuing with partial notification coverage.
- Removing PostgreSQL or MySQL notification targets.
- Deleting the legacy struct fields needed to decode and identify unsupported input. They remain on shared target argument structs that are also used by live constructors, whose discrete-field connection-string synthesis is unreachable from current KV configuration; those fields must not become supported configuration keys.
- Correcting ignored errors from the other eight legacy notification setters. Their pre-existing silent-skip behavior remains unchanged in this narrowly scoped database-migration patch and requires a separate audit and design decision.
Functional requirements
Current configuration
notify_postgresacceptsconnection_string;notify_mysqlacceptsdsn_string.- The five discrete keys remain unregistered and rejected by current configuration commands.
- Existing full strings must continue to support the database driver’s syntax, including parameters whose names contain
host,port,user,password, ordatabase. - No new public environment variables or KV keys are introduced.
- The declared legacy variables
MINIO_NOTIFY_POSTGRES_HOST/PORT/USERNAME/PASSWORD/DATABASEand their MySQL equivalents are not wired into current parsing and remain unsupported. They must not be documented as working alternatives to the complete-string variables.
Legacy migration
SetNotifyPostgresmust return without emitting a target when the legacy target is disabled.- For an enabled target,
SetNotifyPostgresmust require a non-emptyConnectionStringand write only registered Postgres keys. If both a canonical string and discrete fields are present, the canonical string wins and every discrete value is discarded. SetNotifyMySQLmust apply the equivalent rule toDSN.- Neither helper may emit
host,port,username,password, ordatabase. - A missing canonical string must return a typed or wrapped migration error identifying the subsystem and target name.
cmd/config-migrate.gomust check and propagate both helper errors. Ignoring them is forbidden.- No partially migrated configuration may be activated or persisted after either helper fails.
- Error text may name the required key and remediation, but must not include any connection-field value.
- The propagated typed migration error must abort server startup. It must not be downgraded to the non-fatal “some features may be missing” path in
initConfigSubsystem, and it must not enter the retriable-error loop. - Validation errors for a supplied canonical string follow the same startup-fatal and secrecy rules; wrapping must add target context without repeating the DSN or its components.
Original proposed error shape (illustrative, not the shipped literal):
Operator remediation
An operator encountering the error must choose an explicit remediation path. This applies both before an initial switch to SILO and when upgrading a deployment that is already running SILO: legacy migration output is not persisted, so the same old JSON source can re-enter migration on every start. A deployment that currently starts with notifications silently broken can therefore fail to start after this repair until the source configuration is corrected.
- On a compatible intermediate MinIO release, replace the old fields with
connection_stringordsn_string, verify the target, and then migrate to SILO. - Disable or remove the legacy database target, migrate the server, and recreate the target with the canonical string afterward.
- For a fresh SILO installation, create the target directly with the canonical string; no legacy migration is involved.
- For an existing SILO deployment that still reads a legacy JSON file, stop on the previous working release, back up the source configuration, then convert, disable, or remove the database target before starting the fixed release. Do not delete or rewrite unrelated configuration.
Documentation must not suggest that a discrete-only target will be converted automatically.
Availability trade-off
This decision intentionally turns one unsupported configuration from a degraded startup into a hard startup failure. The immediate availability cost is real: a server that previously served objects while all notifications were silently dead may refuse to start after the repair.
That cost is accepted because an object server that appears healthy while configured event sinks are absent creates silent, potentially unrecoverable downstream data loss. SILO is a new fork with an explicit migration boundary, and the discrete form has been deprecated since 2020. A fatal, actionable precondition is preferable to an upgrade that reports success with reduced notification coverage. The release note must make this startup behavior prominent; it must not be buried as an internal migration cleanup.
Security requirements
- The unsupported-input error must never format the legacy argument structure or its values.
- Tests must use a sentinel password and assert that it is absent from returned errors and captured logs.
- Migrated output must contain the registered sensitive connection-string key and no standalone password key.
- If a diagnostic bundle was exported from an affected deployment before this repair, operators should treat the database password as potentially disclosed and rotate it.
Alternatives considered
Register and parse the discrete fields
Benefit: preserves the old source form and uses already existing argument fields.
Rejected because: registration makes common field names visible to the shared tokenizer and corrupts quoted connection strings. It also expands the supported public configuration surface after the fields were deprecated in 2020.
Synthesize a canonical string during migration
Benefit: preserves discrete-only legacy installations.
Rejected because: it creates permanent code and test ownership for an obsolete input form, including PostgreSQL quoting, MySQL DSN formatting, socket and IPv6 behavior, defaults, and future driver drift. For a new fork with an explicit migration boundary, the benefit does not justify the continuing surface.
Skip only the unsupported target
Benefit: keeps the object server and other notification targets running.
Rejected because: silently discarding a configured event sink can cause unobservable and unrecoverable event loss. A clear migration failure is safer than an apparently successful upgrade with reduced notification coverage.
Change global notification fail-fast behavior
Benefit: limits the blast radius of future invalid targets.
Rejected for this change because: it neither repairs the database target nor closes the credential-exposure path, and it changes system-wide error semantics. It may be evaluated independently with its own operational contract.
Remove database notification targets
Benefit: removes the complete database-specific maintenance surface.
Rejected because: the targets remain useful and self-contained. The defect belongs to an obsolete configuration form, not to the notification capability itself.
Implementation scope
The server change should remain narrow:
- Update
internal/config/notify/legacy.goso the two database setters emit only canonical registered keys and reject enabled targets without a canonical string. - Update
cmd/config-migrate.goto propagate the two database-helper errors with subsystem and target context. - Define a typed database-migration error and update
cmd/server-main.gosoinitConfigSubsystemreturns it as fatal instead of logging and ignoring it. It must remain non-retriable. - Leave ignored errors from the other eight legacy notification setters unchanged in this patch; record them for a separate audit rather than expanding #53 implicitly.
- Remove all ten Postgres/MySQL entries from
knownUnregisteredWrites; the ratchet should become empty unless another independently justified legacy exception exists. - Add focused migration, startup, validation, secrecy, and coexistence tests.
- Update database-notification and migration documentation in
silo.pgsty.com.
The patch must not register the old keys, change the generic tokenizer, or refactor unrelated notification targets.
Acceptance criteria
The implementation is complete only when all of the following are demonstrated:
-
A legacy PostgreSQL target with a complete connection string migrates, passes
CheckValidKeys, and is returned byGetNotifyPostgresunchanged. -
A legacy MySQL target with a complete DSN does the equivalent.
-
Discrete-only enabled targets for both databases fail before target initialization with an actionable error containing the subsystem and target name, and server startup aborts.
-
Missing-string and malformed-string errors contain none of the sentinel host, username, password, database, or DSN values.
-
Disabled discrete legacy targets do not create configuration entries and do not block migration.
-
Migrated KVS output contains none of the ten discrete keys, including empty ones.
-
When a legacy target contains both a canonical string and conflicting discrete values, only the canonical string is migrated and no discrete sentinel appears in any output KVS value.
-
A
SetKVSregression test using the realDefaultPostgresKVSandDefaultMySQLKVSkey sets accepts a quoted connection string containingport=,host=, orpassword=. -
A configuration containing healthy Webhook, Kafka, or NATS targets cannot reach
FetchEnabledTargetswith an invalid migrated database target becausereadConfigWithoutMigratefails without yielding, persisting, or activating a partial configuration, and startup aborts on that typed error. -
initConfigSubsystemreturns the typed migration error; it neither logs-and-continues nor enters the retriable loop. -
knownUnregisteredWritesno longer contains Postgres or MySQL exceptions. -
The following verification passes:
The verbose
cmdoutput must show that tests with both prefixes actually ran; a zero-match warning is a failed acceptance check. The normal server CI suite must also pass. In the documentation checkout, runmake check.
Implementation result
Server commit f1ba68358 implements the accepted design without expanding the public configuration surface:
- the two legacy database setters emit only
connection_stringordsn_stringplus registered target settings; - disabled targets remain ignored, while enabled targets without a canonical string return a value-free
LegacyDatabaseTargetError; - only the two database migration errors are newly propagated;
- the typed error is non-retriable, escapes
initConfigSubsystem, and is classified as fatal byserverMainbeforelogger.FatalIfexits the process; - the ten Postgres/MySQL exceptions were removed from
knownUnregisteredWrites; - focused tests cover complete-string round trips, canonical precedence, discarded discrete values, secrecy, failed-migration atomicity, startup classification, and the real tokenizer key sets.
The final local Claude Code review used Claude Fable 5 at max effort and returned GO with high confidence and no blocking findings. Verification included the focused package set, race tests, go vet ./cmd, and the complete go test ./cmd -count=1 suite. The review authorized only the six-file server commit; publication remains a separate gate.
Cross-repository review found no implementation changes are required in pgsty/mc, pgsty/silo-pkg, or pgsty/silo-console: the client forwards configuration text, the package repository owns no notification schema, and Console already serializes its form into the canonical connection_string or dsn_string. The public reference and compatibility documentation is updated with this record.
Release and compatibility statement
The release note must describe this as an enforced compatibility boundary:
SILO database notification targets require
connection_stringfor PostgreSQL anddsn_stringfor MySQL. The pre-2020 discretehost/port/username/password/databaseform is not migrated. Convert or recreate such targets before switching the deployment to SILO.
Deployments already running SILO with an old-format source configuration are equally affected: after this release the server will not start until each enabled legacy database target is converted, disabled, or removed.
The issue should close only after the repair is present in a published server tag. A merged patch, a local site build, and a published release are separate completion gates.
Review record
Claude Fable 5 reviewed the first draft at xhigh effort on 2026-08-23 and returned approve with required changes. The required calibration was incorporated: startup-fatal propagation now extends through initConfigSubsystem; already-running SILO deployments are covered; the availability trade-off is explicit; canonical-string precedence, dead legacy environment variables, other ignored helper errors, and executable tests are specified.
The same model then completed a final source-backed verification pass. Final verdict: approve, with no blocking findings. It confirmed that the English and Chinese records are aligned, the requirements are implementable against the current server tree, and the acceptance criteria cover the startup, migration, parser-regression, and secrecy boundaries.
After implementation, a separate local Claude Code review using Claude Fable 5 at max effort traced the path through ExitFunc(1), inspected driver error behavior, ran the focused, race, vet, and full cmd suites, and returned GO with high confidence and no blocking findings.
16 - Preview Text, Never Execute It: SILO Console Text Preview PRD
Status: shipped in SILO Console 2.2.0 · Owner: pgsty/silo-console · Tracking: pgsty/silo#17 · Review: consensus of product, security, and frontend architecture reviews
SILO Console can preview images, PDFs, audio, and video, but not the small logs, text files, JSON documents, and XML documents that operators inspect every day. A correctly stored Content-Type does not help: these objects are classified as unsupported before the preview renderer is selected.
Restoring the old browser-native behavior would be easy. It would also be the wrong fix. An object in storage is controlled by the user who uploaded it. Loading that object as a same-origin HTML or XML document would turn a convenience feature into an execution boundary.
The accepted design therefore makes a stronger promise:
SILO previews eligible objects as bounded UTF-8 text. It never asks the browser to interpret their markup, MIME type, or file contents as a document.
This record fixes the product boundary, the resource limit, the security invariants, the implementation shape, and the evidence required before the feature can ship.
Decision
The first release will add a dedicated text preview type and a PreviewText component.
The contract is:
- Preserve every existing image, PDF, audio, and video classification.
- Only when the existing classifier returns
none, consider a text fallback. - Admit the four target extensions or four exact passive text MIME types.
- Fetch bytes through the ordinary authenticated download path, without
preview=true. - Enforce a hard application read limit of 1 MiB.
- Decode only strict UTF-8 and reject binary-looking content.
- Render one React text node inside a scrollable
<pre>. - Never use an iframe, HTML parser, XML parser, or HTML injection API.
- Show the complete object or no object; do not show a truncated JSON or XML document.
- Keep download available for files that are too large, invalidly encoded, or otherwise unavailable.
No new API route or S3 operation is added, and the backend inline MIME allowlist is unchanged. Delivery did change Console download responses: zero-byte Range requests return an empty 200, unsatisfiable ranges return 416 with Content-Range: bytes */N, and size is always emitted in object JSON. See Console 2.2.0.
Current behavior
The defect is present in SILO Console v2.1.1, the version currently pinned by SILO when this design was written.
The frontend preview union contains only:
Its extension table contains media formats, but not .log, .txt, .json, or .xml. Its MIME classifier likewise ignores text/plain, application/json, application/xml, and text/xml.
Runtime verification produced this split:
| Object | Frontend result | Console download response |
|---|---|---|
.log / text/plain |
none |
inline, SAMEORIGIN |
.txt / text/plain |
none |
inline, SAMEORIGIN |
.json / application/json |
“Preview unavailable” | inline, SAMEORIGIN |
.xml / application/xml |
none |
attachment, DENY |
The object-detail action also uses the wrong conjunction when deciding whether Preview should be disabled. An authorized user can click Preview for an unsupported object and receive only the unavailable message; in other combinations, the UI can offer an action before the server rejects it.
The preview component still contains a generic same-origin iframe fallback. It is unreachable under the current type union, so the current defect is not an exploitable text-preview XSS. The dead branch is nevertheless hazardous: adding text to the union and letting it fall through would reactivate precisely the document-loading behavior this design rejects.
Root cause
This is contract drift across three independently evolved layers.
Classification drift
The browser code decides eligibility from filename and object metadata, but its closed type union has no text representation. Correct metadata cannot select a renderer that does not exist.
Response-policy drift
The Console server separately decides whether a response may be inline. It still treats plain text and JSON as safe passive MIME types, while XML and HTML remain attachments. That server decision is not reflected in the frontend classifier.
Renderer drift
The old generic iframe remains after the set of reachable preview types became media-only. The code therefore suggests a capability that the type system can no longer invoke.
The repair must realign the three layers without making MIME metadata a security boundary.
Why same-origin iframe preview is rejected
X-Frame-Options: SAMEORIGIN is not a sandbox. It controls who may embed a response; it does not limit what code inside a same-origin frame can do.
If uploader-controlled HTML, XHTML, SVG, or active XML were ever served as an inline same-origin document, it could act with the Console origin. An HttpOnly cookie would prevent direct cookie reads, but it would not prevent authenticated same-origin requests. A permissive or accidentally widened MIME rule would then turn stored content into stored application code.
nosniff, Content Security Policy, and Content-Disposition remain useful defense in depth, but none replaces the core invariant:
Product contract
The feature is a read-only text viewer, not a web previewer and not an online editor.
The user should be able to:
- open a small eligible object from either the list or object-detail surface;
- read whitespace-preserving source text in the existing preview modal;
- select and copy text using browser-native behavior;
- understand whether a failure is caused by size, encoding, permission, object replacement, or network error;
- download the original bytes at any time.
The user must never be led to believe that:
- formatted JSON is the stored object;
- a partial XML document is complete;
- replacement characters are original bytes;
- an unsupported encoding has been decoded faithfully;
- an active HTML/XML document has been safely “sanitized” and executed.
Goals and non-goals
Goals
- Preview small logs, text, JSON, and XML without a local download.
- Keep object content inert regardless of extension, MIME, or payload.
- Bound retained response bytes and rendered text to 1 MiB.
- Preserve the stored text rather than silently reformatting it.
- Keep list and detail actions consistent with permissions and type eligibility.
- Support current object versions and explicitly selected historical versions.
- Preserve anonymous-access and subpath-hosting behavior.
- Ship the feature in Console first, then consume that exact Console revision in SILO.
Non-goals
- HTML or XHTML rendering.
- XML parsing, XSLT, external entities, or schema validation.
- Markdown rendering.
- JSON pretty-printing.
- YAML or CSV-specific behavior.
- Editing or saving.
- Syntax highlighting, line numbers, search, folding, ANSI rendering, or linkification.
- Head, tail, or truncated previews for large objects.
- Lossy decoding or automatic detection of GBK, UTF-16, Latin-1, or other encodings.
- A new backend text-preview endpoint.
- Changes to the existing SVG, media, PDF, download, share, or storage contracts.
An object such as notes.md may still be shown as raw text when its exact MIME type is text/plain. It does not gain Markdown semantics.
Eligibility contract
Eligibility is deliberately two-stage.
Stage 1: preserve the legacy media decision
Run the current image, PDF, audio, and video classifier unchanged. If it returns anything other than none, return that result.
This preserves historical behavior for conflicting filename and MIME combinations.
Stage 2: apply text fallback
Only after the legacy result is none:
-
Reject final extensions
.html,.htm, and.xhtml. -
Match the final filename extension case-insensitively against:
.log.txt.json.xml
-
Normalize Content-Type by removing parameters, trimming whitespace, and lowercasing it.
-
Match the normalized MIME exactly against:
text/plainapplication/jsonapplication/xmltext/xml
An allowed extension or an allowed exact MIME is sufficient. Broad matches such as text/, substring tests, and application/+json are forbidden in this release.
The resulting matrix is normative:
| Filename and MIME | Result | Reason |
|---|---|---|
report.txt + image/png |
image | Existing media decision wins. |
report.json + application/pdf |
Existing media decision wins. | |
server.LOG + application/octet-stream |
text | Allowed extension, case-insensitive. |
no extension + application/json; charset=utf-8 |
text | Exact normalized MIME. |
page.html + text/plain |
none | Explicit active-extension exclusion. |
page.txt + text/html |
text | Extension admits it; HTML source remains inert text. |
notes.md + text/plain |
text | Exact MIME admits raw text, not Markdown rendering. |
image.svg + image/svg+xml |
existing image path | No new text or iframe path. |
Filename and MIME affect product eligibility only. They never select an executable rendering mode.
Resource contract
The binary limit is:
Exactly 1 MiB is eligible. 1 MiB plus one byte is not.
Known sizes
- If the selected version has a known size greater than the limit, the initial attempt does not request its body. Explicit Retry bypasses that possibly stale listed size but retains the bounded Range and byte limit.
- A known zero size still takes the bounded request path; an empty response produces the empty-file state.
- If its known size is within the limit, begin a bounded request.
- An absent size is not the same as zero; it enters the bounded unknown-size path.
The current list-to-modal handoff must therefore preserve undefined rather than converting it to zero with a truthy fallback.
Bounded request
For a small or unknown size, request:
The extra byte is an over-limit sentinel.
The client must:
- Inspect
Content-RangeandContent-Lengthwhen present. - Read the response as a stream rather than calling
response.text()or building a complete Blob. - Retain at most the limit plus the sentinel byte.
- Cancel immediately when the sentinel byte is observed.
- Enforce the same limit when the server ignores Range and returns 200.
- Render only after end-of-stream proves that the complete object is within the limit.
An over-limit object opens an explanation state with its known size, the 1 MiB policy, and a Download action. It never shows a prefix fragment.
Request identity and cancellation
A preview request is identified by:
The request must use the existing generated API client or an equivalent base-path-safe helper so that it preserves:
- same-origin credentials;
- the current Console subpath;
version_id;- anonymous-mode
X-Anonymous: 1; - current error handling and permission boundaries.
Close, object change, version change, bucket change, and component unmount must abort the active request and clear the old content.
Abort alone is insufficient. A generation token or invalidation flag must also prevent a response that already completed reading or decoding from updating a newer preview.
An aborted request is not an error and must not produce an error toast.
Encoding and fidelity
The first release supports strict UTF-8 only:
Requirements:
- handle the UTF-8 BOM without displaying it;
- preserve Unicode text, emoji, tabs, LF, and CRLF;
- reject invalid UTF-8 rather than inserting replacement characters;
- reject decoded NUL characters as binary or unsupported content;
- do not guess another encoding;
- do not log or persist object text;
- always retain Download as the original-byte escape hatch.
The unsupported-encoding state should explain:
This object is not valid UTF-8 text or contains binary data. Download it to inspect the original bytes.
JSON and XML are displayed exactly as decoded source text. The first release must not run JSON.parse followed by JSON.stringify: that can alter unsafe integers, duplicate keys, whitespace, lexical forms, and the text users copy.
Safe renderer
The success state renders one text node:
The implementation must not use:
- iframe, object, or embed;
dangerouslySetInnerHTMLorinnerHTML;DOMParseror an XML parser;- Markdown or HTML rendering;
- an HTML data/blob URL;
- per-line or per-token spans;
- automatic links, ANSI escapes, or syntax markup.
One bounded text node keeps the DOM cost predictable and the security property inspectable.
The preformatted region uses a monospace font, preserves whitespace, defaults to no wrapping, owns both scrollbars, is keyboard focusable, and supports native selection and copy. No-wrap is intentional: it preserves aligned logs and avoids expensive layout of a single very long line.
UI states and permissions
The Preview action is enabled only when:
The object-detail conjunction bug must be fixed, and list and detail surfaces must share the same eligibility function.
An eligible over-limit object still offers Preview. The modal explains why content is not loaded; disabling the button would leave the user unable to distinguish size, permission, and type failures.
The modal distinguishes:
| State | Required behavior |
|---|---|
| Loading | Accessible busy state; no stale text. |
| Success | Scrollable raw text plus Download. |
| Empty | Explicit “File is empty” state. |
| Too large | Object size, 1 MiB limit, Download and Retry; no initial body request when size is known to exceed the limit. |
| Invalid UTF-8 / binary | Dedicated explanation and Download. |
| Forbidden | Permission-specific message; no retained text. |
| Not found / replaced | Object-change message; no retained text. |
| Network / server error | Actionable retry/download state. |
| Aborted / closed | Silent cleanup. |
HTTP error bodies must never be decoded and displayed as object content.
All new user-facing strings go through the existing translation layer and ship in English and Chinese together. The content region and controls must remain usable in light and dark themes and at narrow widths.
Functional and security requirements
Functional requirements
- FR1: Existing media and PDF classification remains unchanged.
- FR2: The text fallback follows the normative extension/MIME matrix.
- FR3: Eligible complete objects up to 1 MiB render as strict UTF-8 source.
- FR4: Over-limit objects render no partial content.
- FR5: Empty objects have a distinct successful empty state.
- FR6: Current and selected historical versions use the same version for metadata, size, and body.
- FR7: Anonymous access and subpath hosting retain their current request behavior.
- FR8: List and detail actions apply the same type and permission decision.
- FR9: Download, share, media, PDF, and storage behavior do not change.
Security requirements
- SR1: Object bytes can reach the DOM only through text content.
- SR2: Text Preview contains no document renderer or parser.
- SR3: At most 1 MiB plus one sentinel byte is retained.
- SR4: Closing or changing identity invalidates every previous response.
- SR5: Invalid UTF-8 and NUL content are not shown as faithful text.
- SR6: Errors, Redux, local storage, logs, and telemetry never retain preview text.
- SR7: Server authorization remains authoritative for direct requests.
- SR8: No CSP or backend inline MIME relaxation is introduced.
Implementation scope
Expected Console changes:
- Refactor preview classification so the current media decision is preserved and text is an explicit fallback.
- Add
textto the preview type union. - Add a dedicated
PreviewTextcomponent with streaming bounds, strict decode, request cancellation, and explicit states. - Route text objects explicitly to that component.
- Remove the unreachable generic iframe fallback.
- Fix the object-detail Preview disable expression and share eligibility logic with the list surface.
- Preserve unknown size instead of coercing it to zero.
- Add English and Chinese strings.
- Add classification, component, resource, security, permission, version, and browser tests.
Expected unchanged areas:
- Console and S3 API paths;
- the backend
safeMimeTypeslist; - Content Security Policy;
- object storage and metadata formats;
- image, PDF, audio, video, download, and share handlers;
- external frontend dependencies.
If a future product requires tailing, server-side transcoding, organization-wide policy, or reliable behavior through proxies that ignore Range, a dedicated server endpoint may be designed separately.
Rejected alternatives
Keep text preview disabled
Benefit: no new code or browser memory use.
Rejected because: logs and configuration objects are a routine object-storage workflow, and download-only inspection is an avoidable Console regression.
Reuse the same-origin iframe
Benefit: minimal code and browser-native presentation.
Rejected because: it turns uploader-controlled content and mutable MIME metadata into a same-origin document boundary. It also leaves resource use unbounded.
Add a backend preview API now
Benefit: central server-side limits and normalized text responses.
Rejected for the first release because: the user already has object-read permission, and the existing download endpoint provides versioning, authorization, and Range. A new API would duplicate contracts without establishing a new data-access boundary.
Show the first 1 MiB of a large object
Benefit: better large-log convenience.
Rejected because: partial JSON/XML is structurally misleading, UTF-8 boundaries need additional handling, and a single “preview” action would no longer mean complete content.
Decode invalid UTF-8 with replacement characters
Benefit: some damaged or legacy logs remain partially readable.
Rejected because: copied text would no longer faithfully represent the stored object. Lossy viewing and other encodings require a separate, explicit product mode.
Auto-format JSON
Benefit: more readable indentation.
Rejected because: parse/stringify can alter numbers, duplicate keys, lexical representation, and copied content. A future opt-in formatted view may sit beside, never replace, the raw default.
Add Monaco or another code editor
Benefit: line numbers, search, highlighting, and folding.
Rejected because: bundle, worker, CSP, and maintenance costs exceed the needs of a bounded read-only preview. A native <pre> is smaller and easier to audit.
Acceptance and test plan
Classification matrix
Automated tests must lock every normative matrix row, extension case handling, MIME parameter stripping, explicit HTML/XHTML denial, and unchanged media conflicts.
Resource tests
Cover:
- 0 bytes;
- 1 byte;
- exactly 1,048,576 bytes;
- 1,048,577 bytes;
- known over-limit size with zero body requests;
- unknown size;
- 206 with a revealing
Content-Range; - server ignores Range and returns 200;
- missing or false
Content-Length; - close and identity changes during streaming.
No case may retain or render more than the complete allowed object.
Encoding and fidelity tests
Cover UTF-8 Chinese, emoji, tabs, LF, CRLF, BOM, invalid byte sequences, NUL bytes, JSON unsafe integers, duplicate keys, original whitespace, XML declarations, DOCTYPE, CDATA, and stylesheet processing instructions.
The raw success view must preserve decoded text. Invalid and binary cases must show their dedicated state.
Security tests
Payloads containing <script>, event attributes, iframe tags, SVG handlers, XML stylesheets, external entities, and suspicious URLs must:
- appear literally in
<pre>.textContent; - create no corresponding DOM elements;
- execute no script or dialog;
- cause no object-content-originated request;
- encounter no iframe, object, embed, HTML parser, or XML parser in Text Preview.
Permission and race tests
Verify:
- no
GetObjectmeans no usable action and no retained body; - historical versions require their corresponding permission;
- metadata and body use the same version ID;
- a late old response cannot replace a new object’s preview;
- 401, 403, 404, 416, and 5xx bodies never become preview content;
- anonymous access and Console subpaths do not regress.
Browser regression
Use a real SILO/Console test instance to inspect both English and Chinese routes, light and dark themes, and narrow and desktop widths. Media, PDF, download, share, and version workflows require smoke coverage alongside the new text states.
Delivery and completion gates
The change belongs to pgsty/silo-console, even though the user report is tracked in the SILO server repository.
Delivery is staged:
- Merge the focused Console source and test change.
- Pass TypeScript checking, production build, automated matrices, and real-browser security regression.
- Update Console release notes and regenerate the actual embedded web assets.
- Publish a Console version; a minor release is appropriate for the new visible capability.
- Update SILO’s
github.com/minio/console => github.com/pgsty/silo-consolereplacement to the exact new pseudo-version. - Build a SILO candidate from that exact dependency and repeat integration checks.
- Publish the SILO binary and image, naming the first version that contains the feature.
These are separate states:
| Gate | Meaning |
|---|---|
| Console PR merged | Implementation exists in source. |
| Console assets/tag published | Console is independently consumable. |
| SILO dependency updated | SILO main has integrated the change. |
| SILO release published | Users can obtain the feature. |
Issue #17 should not be described as fixed for users merely because a local preview or Console source PR exists.
Trade-off summary
The accepted design favors:
- explicit scope over a generic browser viewer;
- complete small files over partial large files;
- source fidelity over automatic formatting;
- strict UTF-8 over silent lossy decoding;
- one inert text node over a full editor;
- the existing download API over a new backend contract;
- a verifiable security invariant over convenient same-origin rendering.
The cost is real: large logs and legacy encodings still require download, and the first release has no search, line numbers, wrapping toggle, or highlighting. Those omissions are deliberate. They make the feature small enough to audit and strong enough to trust.
Review record
The design was independently reviewed from three perspectives:
- product scope, delivery, and acceptance;
- security and frontend architecture;
- compatibility and current-source verification.
The reviewers initially differed on MIME-only eligibility and lossy UTF-8 fallback. After cross-review they reached a single contract:
- existing media classification wins;
- text fallback accepts the four target extensions or four exact normalized MIME types;
- HTML/XHTML extensions are explicitly excluded;
- strict UTF-8 and NUL rejection are required;
- lossy viewing is deferred to a separate proposal.
No unresolved design question remains. Implementation may proceed against this record.
Current Retry boundary: the too-large state permits reprobe of potentially stale listed size, never an unbounded download. The Range, response-header checks and at-most-1 MiB + 1-byte read limit still apply. Empty files are verified through the response path; unknown size must not be treated as zero.
17 - Go 1.27 TLS Defaults and OIDC Discovery Failure Modes
Publication update, 2026-09-17: The Go TLS default repair (
48e184652) shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
Follow-up, 2026-09-17: the #154 reporter retested on 20260916 and reports the same failure. The repair restores the effect of
GODEBUG=tlsmlkem=0; it does not change the default handshake, and it cannot address an ingress that rejects the new ML-DSA signature identifiers. The mechanism analysis, the full option space and the release-communication gates that follow from this are recorded in Pinned TLS Parameters and Handshake Compatibility.
After SILO’s toolchain moved to Go 1.27, the Server TLS repair
48e184652
(“fix(tls): honor Go key exchange defaults across transports”) removed its
explicit curve overrides. This page records the TLS changes and the diagnostic method for OIDC
discovery failures that came out of issue #154,
and the operational facts an administrator needs when identity goes missing
at startup.
Release boundary, 2026-09-16: Server 20260903 already uses Go 1.27.1, but does not contain
48e184652. That later TLS repair is on main; upgrading the compiler and adopting this repair are separate changes.
Evidence class, stated up front. Every mechanism below is verified by synthetic experiments: ClientHello captures, fresh-process CA probes, and fixture reproductions. The #154 customer’s discovery URL and ingress configuration were never obtained, so no root cause is claimed for that deployment — two locally verified mechanisms could each produce the reported symptom, and either the ingress rejecting the new handshake, or a proxy rejecting the changed User-Agent, remains plausible. #154 was closed on 2026-09-11 on the strength of the merge; the 2026-09-17 retest on 20260916 still fails, so an affected-environment retest with phase-level evidence is still owed.
What Go 1.27 changed
- Explicit curve preferences now override the ML-KEM compat switches.
GODEBUG=tlsmlkem=0removes all ML-KEM hybrids from the default set;tlssecpmlkem=0removes only the P-256/P-384 hybrids introduced in Go 1.26 and retains X25519MLKEM768. An application that configuresCurvePreferencesexplicitly keeps ML-KEM in whatever list it names — a deliberate Go 1.27 change. SILO had eight TLS configuration points setting an explicit list including X25519MLKEM768; the fix removes all eight assignments and retires the helper, so these Server configuration points follow Go’s defaults and the compat switches work again. The stack review found pkg, mcli, and Console clients already used defaults; Console’s HTTPS listener retains its separate P-256-only policy. - ClientHello now offers ML-DSA signature algorithms (identifiers
0x0904–0x0906). ML-DSA is a signature scheme and distinct from ML-KEM: disabling hybrid key exchange does not disable the ML-DSA offer, and an ingress that rejects ML-DSA is not fixed by any ML-KEM switch. - The ClientHello grew. Measured on the same source and dependencies:
Go 1.26.5 default 1497 bytes; Go 1.27.1 default 1509 bytes; the old
explicit list under
tlsmlkem=0produced a 275-byte hello with no ML-KEM, while Go 1.27.1 with an explicit list still produced 1509 bytes containing ML-KEM. Rebuilding with a different compiler alone changed the handshake. - macOS root-CA behavior flips with the module’s Go directive. A fresh
process honoring
SSL_CERT_FILE/SSL_CERT_DIRinstead of the Keychain is governed by thex509sslcertoverrideplatformGODEBUG default, which follows the main module’sgodirective:go 1.26modules ignore those variables on macOS (platform store wins),go 1.27modules honor them — and the consuming application’s directive wins even when a library module is older. Operators on macOS should know that setting either variable replaces Keychain trust wholesale with the file/directory given; a stale or incomplete path then breaks chains the Keychain would have accepted, and unsetting restores the Keychain. - Not everything changed. TLS versions, cipher suites, certificate and hostname verification, proxy handling, and HTTP/2 selection are unaffected. The standard-library drain cap (256 KiB / 50 ms) and other audited Go 1.27 changes showed no SILO dependency. Go 1.27 binaries require macOS 13 or newer. Downgrading is not a supported path: the module graph requires Go ≥ 1.27.1 across Server, Console, and mc.
Why the OIDC-only patch was withdrawn
The investigation first produced a minimal candidate: clear
CurvePreferences on the OIDC discovery transport only. It was deliberately
not shipped. The same transport serves identity plugins, notification and
lambda reachability checks, audit webhooks, and S3 cloud-backend tiers —
fixing two OIDC call sites would have left every other consumer on the
defective explicit list. The merged repair removes the explicit curves at all
eight Server configuration points so the compat switches apply there, keeps
certificate verification strict, and adds no protocol downgrade or automatic
fallback. The archived one-transport patch must not be reapplied on top of
the merged fix.
Diagnosing a discovery failure by phase
The startup chain is: server start → identity system init → fetch
.well-known/openid-configuration (discovery) → fetch the jwks_uri keys →
IAM store ready → Console initializes. Console’s own OIDC configuration
dialog validates through the same server-side transport. A failure anywhere
in the chain leaves IAM offline; a successful discovery does not clear the
JWKS fetch, and a 503 on JWKS blocks IAM just as hard.
Discriminate by where the connection dies:
- Reset right after the TLS ClientHello (
tls_startthen reset): suspect the ingress’s ClientHello handling — proxy CONNECT rules, TLS terminators, or anything keyed on hello size or contents. This is where the Go 1.27 changes land. - Reset after TLS completes (
wrote_requestthen reset): the TLS layer is fine; look at HTTP-layer policy — WAF rules, User-Agent allowlists (the server’s UA changed fromMinIOtoSilowith the rebrand), routing. Replacing certificates or key exchange here has no targeted effect. - x509 errors: compare the chain actually received, the SNI, and the trust store the process resolves (see the macOS section above).
- Always test from the same network position as the failing process — a fresh container does not inherit the failing container’s network namespace, and same-IP/same-proxy controls come first.
The health endpoint that tells the truth
/minio/health/live and /minio/health/ready both stay 200 while IAM is
offline — readiness as deployed does not cover the identity system. The
endpoint that reports it is /minio/health/cluster, which checks identity
initialization and returns 503 with the X-Minio-Server-Status: iam-offline
marker. Monitoring that should catch a broken IdP integration should probe
the cluster health, plus one authenticated operation.
Recovery is automatic: identity initialization retries at randomized 0–3 s intervals, and a recovered IdP brings IAM back without a restart (observed sub-second to ~1.4 s locally). Retrying cannot fix a persistent incompatibility — a hello the ingress rejects stays rejected.
Transport facts worth knowing
The discovery/JWKS client builds its own transport: HTTP/2 disabled (no ALPN,
HTTP/1.1), proxies taken only from HTTPS_PROXY/NO_PROXY (uppercase
preferred; ALL_PROXY unused), DNS refresh defaulting to 30 s in Kubernetes/Docker and 10 min otherwise (overridable by the DNS cache TTL setting), dialing that walks
addresses in order without shuffling, and timeouts of 5 s per TCP dial,
10 s for the TLS handshake, and 1 min to response headers. There is no
total timeout on the discovery or JWKS fetch itself — a slow IdP can hold
startup indefinitely; tightening that is a known, separately-sized follow-up.
Attribution
This record distills the issue #154 investigation and the September Go 1.27 stack review; the reproduction artifacts and the full evidence chain are retained outside the documentation tree. The supported statement is: the merged fix restores Go key-exchange defaults at the eight affected Server configuration points, verified with synthetic negative controls — it does not claim to have diagnosed any specific hidden deployment, and no root cause is established for #154 until a retest in the affected environment supplies phase-level evidence.
Go behavior is grounded in the official 1.27 release notes and the actual toolchain. ClientHello byte counts above describe the recorded fixtures, not a fixed size for every connection.
18 - No I/O Before Auth, No Privilege From Headers
Release check (2026-09-16): the original repair described here is included in Server 20260903. Dated review and test accounts below record their original evidence, not a still-pending release or acceptance of a particular production installation. Later source changes and component selections are in the version matrix.
This record describes the CORS hot-path and replication-request trust repair merged into SILO as PR #101 (938603458 through 04b097fd9).
Status on 2026-09-03: PR #101 merged into
mainon 2026-09-01 with four follow-up commits: per-entry Snowball trust isolation (ff44527a3), request defaults preserved across Snowball workers (ab3ae99ca), and replication validity probes that verify the replication permissions (c9ad74673) under the rule prefix (5db7be4ee). Implementation, focused and race tests, the complete server package suite, object-lock tests, vet, build, two rounds of Fable 5 design review, repeated Opus 5 adversarial acceptance, and a real local TLS two-site replication run are complete. The pre-release cleanup kept the resident-only lookup with its fail-closed startup and load-failure states, dropped only the internal-namespace special case, and made the header-stripped request clone share the original request trailer so streaming-checksum uploads keep working for untrusted requests. Tag, package, image, deployment, and production verification remain separate gates.
Scope: HTTP request interpretation before and inside the S3 handlers. No S3 wire field, object format, bucket metadata format, replication protocol, encryption format, or client command changes.
Security properties: pre-authentication CORS processing performs no object-layer I/O; a header never grants replication semantics by itself; SSE-C ciphertext paths and replica-only metadata require both authentication and the corresponding replication permission.
Too Long; Didn’t Read (TL;DR)
Two bugs looked unrelated:
- an
Originheader made the outermost CORS middleware treat the first URL segment as a bucket and synchronously load its metadata before authentication; X-Minio-Source-Replication-Requestmade downstream code believe a request was internal replication merely because the header existed.
They shared the same design failure: untrusted request shape was allowed to acquire expensive or privileged internal meaning before an authorization boundary.
The repair establishes two invariants:
For CORS, the outer middleware now reads only metadata already resident in memory. For replication, handlers authenticate the original signed request first, authorize the appropriate replication action, and then attach a private trust decision to the request context. Untrusted internal headers are stripped only after signature verification. The context decision—not header removal—is the authority used by option builders, encryption paths, object lock, event generation, and metadata persistence.
Failure A: pre-authentication CORS amplification
corsHandler wraps the complete server router. Any request carrying Origin reaches it before S3 authentication, request validity checks, and the normal API limiter.
The per-bucket CORS implementation originally called the normal bucket metadata getter:
When .metadata.bin did not exist, the loader intentionally searched legacy configuration files. With none found, it returned a valid empty metadata record rather than NoSuchBucket. The generic getter then inserted that record into metadataMap.
An unauthenticated client could therefore vary otherwise plausible names and obtain two effects per distinct value:
- repeated erasure/object metadata reads before the normal request limiter;
- growth of the in-memory bucket metadata map.
Name validation alone cannot repair this. An attacker can generate an effectively unbounded sequence of syntactically valid, nonexistent bucket names. Distributed deployments eventually prune stale map entries during the 15-minute metadata refresh; single-node deployments do not start that refresh loop, so their synthetic entries persist until restart.
Failure B: a marker header became authority
SILO and its MinIO-compatible clients use internal headers to preserve source state during replication. The most important marker is:
Before this repair, several paths treated header presence—or its raw string value—as proof that the request was a replication request. That affected more than metadata extraction:
GETof an SSE-C object could setNoDecryptionand return ciphertext without the customer key to a caller holding only ordinary read permission;- source ETag and modification time could replace server-generated values;
- source tagging, retention, and legal-hold timestamps could enter last-writer-wins comparisons;
- a past object-lock retention date could be accepted through a raw marker check;
- delete-marker identity and modification time could be supplied by the caller;
- successful object events could be suppressed;
- multipart actual size and encrypted checksum metadata could be injected at completion;
X-Amz-Replication-Statuscould be persisted from ordinary PUT, COPY, or POST-policy metadata extraction.
The earlier CVE-2026-34204 repair correctly stopped ordinary PUT and COPY from importing the replication SSE metadata that could make objects unreadable. It did not yet provide one authority shared by every reader of the marker, source fields, event state, object-lock exceptions, or multipart completion metadata.
Selected design
One exact marker, two trust levels
The marker is accepted only when it appears exactly once and its value is exactly lowercase true. Duplicate values, mixed case, and any other value are untrusted.
The handler then derives two related decisions:
| Decision | Requirements | Semantics it may enable |
|---|---|---|
trusted |
original request authenticated; non-anonymous principal; exact marker; s3:ReplicateObject or s3:ReplicateDelete on the addressed resource |
source ETag/MTime and source timestamps; actual size and encrypted checksum transfer; event and re-replication suppression; replication delete pool/version pinning |
replicaTrusted |
trusted, plus raw request status REPLICA or a multipart upload whose stored status is REPLICA |
replica status persistence; replication SSE sealed-key import; SSE-C ciphertext/no-decryption path; replica-only object-lock behavior |
The split is required by the real wire protocol. Not every legitimate replication request repeats X-Amz-Replication-Status: REPLICA.
The receiver follows this matrix:
| Incoming shape | Result |
|---|---|
| no marker | ordinary S3 operation |
| marker without replication permission | internal fields ignored; operation continues with ordinary semantics |
REPLICA without replication permission |
403 AccessDenied |
exact marker + replication permission, no REPLICA |
trusted only |
exact marker + replication permission + REPLICA |
trusted and replicaTrusted |
The explicit 403 for an unauthorized REPLICA request prevents a claimed replica write from being silently downgraded into a new ordinary object that may be replicated again.
Authenticate the original, then sanitize
SigV4 signs request headers. Removing an internal header before authentication would change the canonical request and turn a valid signature into SignatureDoesNotMatch.
The ordering is therefore mandatory:
The audit logger retains the original request. The effective request clone retains public S3, SSE, checksum, object-lock, copy-source, proxy, and replication-validity headers. It strips only internal source/replication controls, including source ETag/MTime/delete-marker/timestamps, replication SSE state, actual object size, encrypted checksum transfer, and the request use of X-Amz-Replication-Status.
Header stripping is defense in depth. All privileged consumers use the private context decision or an explicit Boolean; they do not infer trust by looking at the clone.
Replica status is not generic user metadata
X-Amz-Replication-Status is an S3 response header that MinIO-compatible servers also use as an internal request control. It no longer belongs to the generic supported-request-metadata list.
Ordinary PUT, COPY, multipart initiation, Snowball/PAX extraction, and POST policy cannot persist it merely by submitting the field. The receiver sets REPLICA explicitly only in a replicaTrusted branch.
This closes a subtle POST-policy path: a form field could previously store REPLICA, causing the resulting object to evade normal replication scheduling even though the POST principal never held replication permission.
Object lock receives an explicit decision
The object-lock parser used to accept past retention dates when the raw marker header was present. That package now receives allowPastRetainDate explicitly from replicaTrusted state.
The surrounding handler also uses the same decision when deciding whether an existing compliance/legal-hold version may be overwritten by a replica. This removes an internal-header dependency from the reusable object-lock package.
Actual replication wire matrix
The design was checked against the silo-go v7.3.1 emitter selected by the server’s go.mod, not inferred from comments or upstream documentation.
| Operation | Marker | REPLICA on this request |
Receiver decision |
|---|---|---|---|
regular replicated PutObject |
yes | yes | replicaTrusted |
replicated NewMultipartUpload |
yes | yes | persist trusted multipart replica provenance |
replicated PutObjectPart |
yes | no | trusted; replicaTrusted only when stored MPU status is REPLICA |
replicated CompleteMultipartUpload |
yes | no | trusted; preserve source ETag/MTime, actual size, and encrypted checksum |
| CopyObject metadata replication | yes | yes | replicaTrusted |
replicated RemoveObject |
yes | yes | replicaTrusted with s3:ReplicateDelete |
| batch replication PUT/Complete | yes | no | trusted; target credentials must hold s3:ReplicateObject |
| proxy/readiness/validity probes | separate probe headers | no marker authority | probe behavior retained; those headers are never stripped by this repair |
s3:ReplicateDelete is the trust gate, not the receiver’s only permission.
For compatibility with deployed target policies, a trusted replication delete
also requires s3:DeleteObject; an explicit deny on
s3:DeleteObjectVersion still blocks a named-version purge. Ordinary clients
do not use this compatibility path: an explicit UUID or versionId=null
requires an allow for s3:DeleteObjectVersion.
Requiring REPLICA for every trusted operation would break PutPart, multipart completion, and batch replication. Trusting every marker would recreate the vulnerability. Stored multipart provenance bridges the two requirements for encrypted raw parts.
CORS resident-only state machine
The outer CORS middleware must remain cheaper than the request it is about to route. It now calls a dedicated resident-only getter that takes one read lock and examines only in-memory state.
| Bucket metadata state | CORS result | Object-layer work |
|---|---|---|
| resident, valid per-bucket CORS | apply per-bucket rule; a failed refresh keeps the last loaded document, as for every other bucket configuration | none |
| resident, no CORS document | use global CORS fallback | none |
| resident, invalid stored CORS | fail closed; continue without CORS headers and log once | none |
| not resident while startup loading is still running | fail closed | none |
| not resident after startup: real bucket whose metadata failed to load | fail closed | none |
| not resident after startup: reserved, invalid, internal, or unknown name | global fallback | none |
The lookup consults the resident map and a bounded set of real buckets whose metadata failed to load at startup or during a refresh. That set is filled only from disk-derived bucket lists, never from a client path, never records a bucket that is already resident, and a successful load, Set, bucket removal, stale-bucket reconciliation, and subsystem reset clear it. Both non-resident states fail closed: a presigned URL is authenticated by its own signature, so the bucket’s CORS document is the only origin boundary a browser enforces for it, and answering with the global policy would let a leaked URL be used from any origin. The internal .minio.sys namespace no longer has a special case; like any reserved or invalid name it is not a bucket, gets the global fallback, and is rejected downstream.
Alternatives rejected
| Alternative | Why it was rejected |
|---|---|
| Validate bucket names before the old CORS getter | valid nonexistent names still provide an unbounded attacker-controlled key space and still trigger pre-auth I/O |
Call GetBucketInfo before loading CORS |
replaces eleven metadata reads with at least one unthrottled backend operation per attacker name |
| Cache every negative result with a TTL | bounds duration, not attacker cardinality or the initial I/O amplification |
| Strip replication headers before authentication | breaks SigV4 canonical-request verification |
| Reject every request carrying an internal marker | turns formerly ignored extra headers into broad client failures and breaks legitimate marker-only replication calls |
Require REPLICA on every trusted call |
breaks replicated PutPart, CompleteMultipartUpload, and batch replication wire behavior |
| Let every handler re-check raw headers independently | recreates inconsistent trust rules and leaves future consumers easy to miss |
Store a Boolean in ObjectOptions but leave events/object lock on headers |
produces two authorities that can disagree; the original bug class remains |
Implementation boundary
The selected change is intentionally layered:
- a small request-trust module defines exact marker parsing, replication authorization, private context state, and the post-authentication effective request;
- object option builders parse source fields only when their caller provides trusted state;
DecryptObjectInfo, event request parameters, multipart completion, delete options, and object lock consume the same decision;- handlers calculate trust immediately after their existing authentication path;
- multipart part handling combines current-request trust with stored MPU replica provenance;
- generic metadata extraction does not accept replica status;
- CORS middleware uses a separate resident-only metadata accessor and never calls the load-on-miss getter.
No object-layer API needs to infer HTTP trust. Programmatic internal callers that construct ObjectOptions{ReplicationRequest: true} remain unchanged.
Verification and adversarial review
Regression coverage includes:
- hundreds of distinct valid missing bucket names, both actual and preflight CORS requests, with zero metadata reads and no map growth;
- Console, reserved, invalid, startup, internal namespace, and invalid stored CORS paths;
- least-privilege SSE-C GET, HEAD, and GetObjectAttributes callers with correct, missing, wrong-case, and unauthorized markers;
- marker-only batch-style PUT preserving source ETag/MTime only with
s3:ReplicateObject; - unauthorized
REPLICAPUT and DELETE returning403; - POST policy unable to forge replica status;
- object-lock past-date parsing with and without replica trust;
- marker-only CopyObject with SSE-C source headers copying plaintext rather than ciphertext;
- fake marker on an ordinary SSE-C MPU failing instead of storing raw bytes;
- a real in-process SSE-C multipart replication chain: encrypted source, raw ciphertext part, trusted replica initiation, marker-only PutPart and Complete, and exact plaintext recovery with the original key.
The final local tree passed focused and race tests, the complete cmd suite, object-lock tests, vet, build, and diff checks.
A separate black-box run started two TLS-enabled SILO instances built from the candidate and enabled real site replication. It verified:
- an SSE-C 4 KiB object;
- an SSE-C 12 MiB, three-part multipart object;
- an SSE-C CopyObject result;
- a replicated delete marker.
Source and target ETag, size, version ID, SSE-C key MD5, decrypted SHA-256, and delete-marker version ID matched; targets reported REPLICA.
Two Fable 5 review rounds first corrected the trust model for marker-only batch and multipart calls, then audited the implementation. A final independent Claude Code Opus 5 review reported GO, with no P0/P1 findings, and independently reran build, vet, race, object-lock, and full cmd tests.
Compatibility and operations
- Ordinary clients: no request change. Untrusted internal headers are ignored instead of acquiring internal semantics.
- Unauthorized claimed replica writes: requests carrying
X-Amz-Replication-Status: REPLICAnow return403where some multipart subpaths previously lacked a uniform check. - Batch replication: destination credentials must include
s3:ReplicateObject, as documented in the batch replication requirements. Without it, the receiver processes marker-only writes as ordinary writes and does not preserve source ETag/MTime. - SSE-C: ordinary reads still require the customer key. Authorized replica reads may use the raw ciphertext path needed to preserve encrypted bytes.
- Events: only trusted replication suppresses replica creation/access events; a forged marker no longer silences them.
- Object lock: replica exceptions are permission-derived rather than header-derived.
- Performance: CORS removes pre-authentication backend work. Trusted writes add policy checks already required by the replication contract; no additional object pass is introduced.
- Rolling upgrade: wire and storage formats are unchanged. New receivers enforce the trust boundary; old receivers remain vulnerable to the old header semantics until upgraded. Per-bucket CORS behavior can therefore differ by node during the rolling window.
- Rollback: data written by the repaired version remains readable by the previous version, but rollback reopens both trust defects and restores pre-authentication metadata loads.
Residual risks and follow-ups
-
2026-09-09 replication reliability follow-up: Delete completion, MRF visibility, and resync cancellation records the reproductions, minimal fixes, Fable review, and PR #162 validation for #153, #152, and #137. It addresses reliability after trusted requests enter the replication pipeline, preserving this page’s authorization boundary.
-
Emit a rate-limited diagnostic when a marker-bearing request lacks replication permission; the safe ordinary fallback is otherwise easy to misdiagnose as an ETag/MTime mismatch.
-
Replication validity probes now verify the replication permissions the target credentials need and place the synthetic validation key under the rule prefix (
c9ad74673,5db7be4ee). -
This review covers the named source/replication headers. Other future internal controls must still answer the same question: which authenticated decision allowed this client value to acquire internal meaning?
Conclusion
An internal-looking header is still client input. A bucket-shaped URL segment is still attacker input. The durable repair is to stop either one from becoming authority by accident:
Before authentication, do no backend work. After authentication, derive trust once and pass the decision—not the claim—downstream.
That rule is broader than CORS or replication. It is the boundary future SILO handlers should preserve whenever inexpensive public request syntax meets expensive or privileged internal state.
This repair is SN-2026-008 in the ledger, with delivery in the 20260903 chronicle. The silo-go reference above describes the original temporary dependency. After 0079723d3, Server returned to verified upstream minio-go; the component matrix records selected versions.
19 - One Endpoint, Two Privileges: Separating User and Group Status
Release check (2026-09-16): the original repair described here is included in Server 20260903. Dated review and test accounts below record their original evidence, not a still-pending release or acceptance of a particular production installation. Later source changes and component selections are in the version matrix.
This document records the discussion, repair, and final authorization design for upstream issue minio/minio#21478 and SILO PR #73.
Status on 2026-08-26: SILO PR #73 was merged as
2e2377d1c, preserving the signed-off repair commit58735ee38. All eight reported checks passed. Upstream issue #21478 and PR #21482 remain open, butminio/miniois archived and read-only, so no further issue comment or merge can be made there.
Group follow-up on 2026-08-28: final release review found the same fixed-action defect inset-group-status. Signed-off server commit229fe2b3cnow selectsadmin:EnableGrouporadmin:DisableGroupfrom the requested target state and adds a real four-way IAM authorization test. Local verification and independent review are complete; it was merged intomainon 2026-08-29, and tag and delivery remain pending.
Scope: authorize enabling and disabling a user with their respective existing Admin Actions. Do not change the route, status values, account storage, replication record, or client API.
Security property: possessingadmin:DisableUsermust not grant the ability to enable an account, and possessingadmin:EnableUsermust not grant the ability to disable one.
Release boundary: merge, tag, release package, container image, deployment, and production verification remain separate gates.
Too Long; Didn’t Read (TL;DR)
SILO exposes both admin:EnableUser and admin:DisableUser, but the shared set-user-status handler historically authorized every request with admin:EnableUser. A policy that granted only admin:DisableUser therefore could not disable an account. The workaround was to grant admin:EnableUser as well, which destroyed the least-privilege boundary that the two action names promised.
The selected repair derives exactly one required action from the requested target state before authorization:
| Requested status | Required action |
|---|---|
enabled |
admin:EnableUser |
disabled |
admin:DisableUser |
| invalid or unknown | admin:EnableUser, preserving the previous authorization-before-validation default |
The handler then calls validateAdminReq once. A four-way IAM test proves both positive operations and both denied cross-action operations. This is intentionally stricter than preserving the accidental historical behavior in which an Enable-only policy could also disable users.
The same rule now applies to group status:
| Requested group status | Required action |
|---|---|
enabled |
admin:EnableGroup |
disabled |
admin:DisableGroup |
| invalid or unknown | admin:EnableGroup, preserving the previous authorization-before-validation default |
Before the follow-up, an EnableGroup-only principal could disable a group, while a DisableGroup-only principal received AccessDenied for that exact operation. The group repair uses the same one-selector, one-authorization design rather than treating the two actions as aliases.
The reported defect
The Admin API uses one route for both state transitions:
Before the repair, the handler checked one fixed action before reading the requested status:
The later call to SetUserStatus correctly received either enabled or disabled, but authorization had already treated both as Enable operations. admin:DisableUser existed in the policy vocabulary and documentation while being ineffective for this endpoint on its own.
Issue #21478 supplied the practical counterexample: an operator wanted a policy that could disable accounts during an incident without being able to restore them. A policy containing admin:DisableUser received AccessDenied; adding admin:EnableUser made the request work, but also gave the operator the more powerful recovery transition that the policy intentionally withheld.
This is not a missing convenience permission. It is a mismatch between the policy model and the enforcement point:
Why two actions must mean two capabilities
An account state transition has direction. Disabling is commonly delegated to incident responders, fraud controls, compliance automation, or a break-glass process. Enabling restores access and may require a separate approver.
If either action authorizes both transitions, a policy author cannot express that separation. The server would publish two names while enforcing one combined capability. The design contract is therefore strict:
| Principal policy | Disable target | Enable target |
|---|---|---|
admin:DisableUser only |
allow | deny |
admin:EnableUser only |
deny | allow |
| both actions | allow | allow |
| neither action | deny | deny |
The built-in consoleAdmin policy grants admin:*, so full administrators retain both operations. The compatibility impact is limited to custom restricted policies that relied on the old accidental behavior.
The public PBAC reference now states the same contract for admin:EnableUser and admin:DisableUser.
Design goals and non-goals
Goals
- Make both existing Admin Actions enforceable according to their names.
- Preserve least privilege in both directions.
- Perform one authorization decision and write at most one authorization error.
- Preserve the route, request values, response format, self-mutation guard, IAM storage call, and site-replication hook.
- Encode the contract in tests that fail if the two permissions are broadened or swapped again.
Non-goals
- split the endpoint into separate enable and disable routes;
- add a new combined action or change policy syntax;
- change user status persistence or replication;
- redesign Console permissions;
- infer release, image, deployment, or production delivery from a source merge.
Alternatives considered
Keep checking admin:EnableUser for both states
This preserves behavior but leaves admin:DisableUser unusable and forces over-privileged policies. It is the defect, not a compatibility contract worth retaining.
Require both actions for either transition
This makes the two labels decorative and prevents delegated disable-only operation. It is stricter in quantity but weaker in expressiveness and least privilege.
Try Enable authorization, then retry Disable authorization
Upstream PR #21482 attempted this shape for a disabled request. It first called validateAdminReq with EnableUser, then called it again with DisableUser if the first result was nil.
That helper has an important contract: when it returns a nil object layer, it has already written an error response. A Disable-only request can therefore commit a 403 response before the second authorization succeeds and the handler proceeds to mutate account state. Authorization fallback must never continue after an error response has been committed.
Accept either Enable or Disable for a disabled request
validateAdminReq already accepts multiple actions and succeeds if any one is allowed, so compatibility behavior could be implemented safely with one variadic call. That would let Disable-only policies work while preserving the historical ability of Enable-only policies to disable.
SILO rejected this option because the historical ability was the enforcement bug. It would solve the reporter’s positive case but retain a cross-action privilege that contradicts the two-action model. Operators who want both transitions can grant both actions explicitly.
Validate the status before authenticating
Rejecting unknown status values first would change error precedence: a caller that previously had to pass the Enable authorization gate could now receive a validation result before authorization. The repair does not need that broader behavioral change.
Unknown values therefore retain admin:EnableUser as the authorization default. Valid disabled is the only value that selects admin:DisableUser; the existing IAM layer remains responsible for rejecting invalid status values after authorization.
The selected implementation
The repair adds a pure selector:
The handler reads the route variables, selects the action, and authorizes exactly once:
Everything after the gate remains unchanged:
- a caller still cannot enable or disable its own account;
globalIAMSys.SetUserStatusvalidates and persists the requested status;- site replication records the same status and timestamp;
- response and audit behavior use the existing path.
The selector depends only on the requested target state. It does not load the current user, infer a transition from stored state, or make authorization depend on whether the target exists. This keeps authorization deterministic and avoids a read-before-authentication dependency.
Why the repair is safe
The correctness argument consists of five invariants:
- Every valid status maps to exactly one Admin Action.
validateAdminReqis invoked once, so a failed authorization cannot be followed by mutation.- The mutation call is reachable only after the selected action succeeds.
- Invalid status values preserve the old Enable authorization boundary and are still rejected by the existing status-validation path.
- No storage, replication, wire, or client contract changes; only the permission required to reach the existing mutation changes.
The change is a deliberate authorization tightening for Enable-only custom policies that used the disable operation. That tightening is the mechanism that makes admin:DisableUser a real independent capability.
Test design
Pure action mapping
The unit test fixes three selector cases:
| Input | Expected action |
|---|---|
enabled |
EnableUser |
disabled |
DisableUser |
| invalid | legacy EnableUser default |
Four-way IAM authorization matrix
The integration test creates separate users and policies, then exercises the real Admin API:
- a Disable-only client successfully disables a target;
- the same client receives
AccessDeniedwhen enabling it; - an Enable-only client successfully enables the target;
- the same client receives
AccessDeniedwhen disabling it.
Positive assertions alone would not prove least privilege: both policies could accidentally authorize both states and still pass. The two negative cross-action assertions are the security regression tests.
The test removes every temporary user and policy after execution. It runs inside the existing IAM server suite, so it covers request signing, policy attachment, handler authorization, persistence, and Admin-client error decoding rather than testing only the helper.
Repair and verification record
The server checkout originally contained unrelated dependency, generated-credit, checksum-test, and security-document changes, while local main was behind the remote. The two user-status files were isolated into a clean worktree based on current origin/main; no unrelated file entered the repair commit.
Local verification passed:
The signed-off commit 58735ee38 was pushed in PR #73. Its eight remote checks all passed:
- DCO sign-off;
- format, build, and vet;
- lint and generated files;
cmd/tests;internal/tests;- race detector and S3 Select;
- cross compilation;
- vulnerability analysis.
The PR was merged with the repository’s normal merge strategy as 2e2377d1c. Local main was then fast-forwarded only after the two original working files were byte-for-byte and patch-ID identical to the merged result. The unrelated local changes remained intact, and the temporary worktree and task branch were removed after the code became recoverable from main and PR #73.
Least-privilege policy examples
Disable-only operator
This principal can inspect and disable another user, but cannot enable it.
Enable-only operator
This principal can inspect and enable another user, but cannot disable it. Grant both actions explicitly to roles responsible for the complete account lifecycle.
Group-status follow-up
The group endpoint has the same shape as the user endpoint:
It also publishes two existing actions, admin:EnableGroup and admin:DisableGroup. The inherited handler nevertheless authorized every request with EnableGroup before reading status. This was not merely a dead permission: it reversed least privilege in both directions. The wrong principal could disable a group, and the intended disable-only principal could not.
The follow-up adds setGroupStatusAdminAction, deliberately matching setUserStatusAdminAction:
The integration test creates separate EnableGroup-only and DisableGroup-only administrators and a real target group. It proves:
- DisableGroup-only can disable;
- DisableGroup-only cannot enable;
- EnableGroup-only can enable;
- EnableGroup-only cannot disable.
The suite exercises signed Admin requests, policy attachment, handler authorization, IAM mutation, response decoding, and cleanup. Invalid status still selects the legacy Enable action before the existing validation error, so the change does not expose a new pre-authentication oracle. The successful site-replication hook remains after mutation and is not called for denied requests.
This follow-up changes no user behavior and introduces no new policy action. It makes the two already documented group actions enforce the same state-specific contract as their user counterparts.
Compatibility and migration
No client or API migration is required. The endpoint, query parameters, status strings, success response, and Admin-client method are unchanged.
Policy review is required for restricted administrative roles:
- a role that should only disable users needs
admin:DisableUser; - a role that should only enable users needs
admin:EnableUser; - a role that must do both needs both actions;
consoleAdminand otheradmin:*policies are unaffected;- a legacy custom policy containing only
admin:EnableUsercan no longer use that permission to disable users and must addadmin:DisableUserif both operations are intended.
The equivalent rules now apply to group-management roles:
- a role that should only disable groups needs
admin:DisableGroup; - a role that should only enable groups needs
admin:EnableGroup; - a role that must do both needs both actions;
- a legacy EnableGroup-only role can no longer disable groups.
This is a source-level compatibility change in authorization behavior, not a wire-protocol break.
Upstream disposition
As of this record, upstream issue #21478 and PR #21482 are still displayed as open. The upstream repository is archived and read-only. An attempt to leave the single-authorization analysis on the PR was rejected by GitHub because archived, locked discussions cannot accept comments.
The upstream artifacts remain useful provenance but are no longer an actionable delivery path. SILO owns its implemented semantics, tests, merge, release note, and eventual production verification.
Delivery state
| Gate | User repair | Group follow-up on 2026-08-28 |
|---|---|---|
| Design decision | complete | complete |
| Implementation and local tests | complete | complete |
| Independent adversarial review | complete | complete, GO |
| Signed-off commit | complete | 229fe2b3c on main |
| Push, PR CI, and merge | complete | merged 2026-08-29 |
| Tagged SILO release | Server 20260903 | Server 20260903 |
| Release package or container image | See the 20260903 release record | See the 20260903 release record |
| Deployment | not established | not established |
| Production behavior | not established | not established |
| Upstream merge | unavailable; repository archived | not applicable |
Conclusion
The repairs make the authorization model tell the truth. Enabling and disabling users or groups are opposite state transitions with different operational risk, and SILO already exposes different policy actions for each direction. Each handler must therefore select the action from the requested target state and authorize once before mutation.
The code change is small because the design boundary is clear. The durable result is larger: an explicit permission matrix, rejected compatibility alternatives, an invalid-input rule, a four-way integration test, a clean merge record, migration guidance, and an honest release boundary.
The enable/disable repair is separate from the later admin:ChangeMyPassword split. Before upgrading the September candidate, preserve the paired Deny needed for an existing self-password restriction using the password migration guide.
20 - Config Environment Files Are Not Shell Scripts
Release check (2026-09-16): the original repair described here is included in Server 20260903. Dated review and test accounts below record their original evidence, not a still-pending release or acceptance of a particular production installation. Later source changes and component selections are in the version matrix.
This record defines the startup contract for MINIO_CONFIG_ENV_FILE and explains the compatibility repair committed in SILO as 2aea7fe9c.
Status on 2026-08-28: implementation, focused tests, the complete
cmdandinternalsuites, tagged tests, race tests, vet, lint, generated-file checks, rebrand guards, build, and an independent local Fable Max review are complete. The commit was merged intomainon 2026-08-29 as2aea7fe9c; tag, package, image, deployment, and production verification remain separate gates.
Scope: environment-file parsing and named-target discovery only. No configuration key, subsystem, value, precedence, storage format, or client API changes.
Compatibility rule: the file is a SILO input format. Supporting an optionalexportprefix does not make it a POSIX shell program.
Too Long; Didn’t Read (TL;DR)
SILO can load startup variables from a file:
The parser accepts assignments such as:
The last two names are important. Multi-target configuration appends the target name verbatim after an underscore. The configuration subsystem does not restrict a target to a shell identifier; names containing -, ., :, digits, or printable Unicode can be discovered and resolved exactly.
A hardening change accidentally validated every key as [A-Za-z_][A-Za-z0-9_]*. It made my-hook invalid and stopped the server during restart even though the previous loader and the configuration target model accepted it. The repair validates what SILO actually needs instead:
- the name is non-empty, valid UTF-8, and made of visible non-whitespace characters;
=and NUL are not allowed in a name;- NUL is not allowed in a value;
- invalid input reports file and line without reporting the value;
- the complete file is parsed before any assignment is applied.
Why the regression was real
The environment-file loader calls os.Setenv after parsing. An operating-system environment is a list of strings, not a shell variable namespace. Shell assignment syntax is narrower because the shell must tokenize and expand variable names in its own language.
Named SILO configuration targets are built differently:
For example:
Target discovery lists variables by the fixed parameter prefix and treats the remaining suffix as the target name. Target lookup reconstructs the same name without uppercasing or sanitizing that suffix. Rejecting - in the file parser therefore broke a valid discover-to-resolve path; it did not protect a shell evaluation path because no shell evaluates the file.
The failure is operationally sharp. MINIO_CONFIG_ENV_FILE is loaded only at startup. A server can continue running with an old process environment, then fail on its next restart after the file or binary changes. Startup must fail on malformed input, but it must not invent a narrower target grammar than the configuration system.
The file grammar
Lines and comments
- blank lines are ignored;
- a line whose first non-whitespace character is
#is ignored; - an optional standalone
exportfollowed by whitespace is removed; exportFOO=valueremains the keyexportFOO; it is not mistaken for the prefix;- the first
=separates key and value, so additional=characters remain part of the value.
The file is not a shell. It does not perform variable expansion, command substitution, backslash processing, or inline-comment interpretation.
Keys
Surrounding whitespace around the key is removed. The remaining key must:
- be non-empty valid UTF-8;
- contain only Unicode graphic characters;
- contain no whitespace,
=, NUL, control, or invisible format characters.
This preserves OS-compatible names and multi-target suffixes while rejecting visually empty or structurally ambiguous keys. A key beginning with a digit or punctuation is accepted by the parser; SILO still reads only the exact names used by its configuration and runtime components.
Values and quoting
Unquoted values are trimmed. To retain leading or trailing spaces, quote the complete value with matching single or double quotes:
The parser removes one matching outer quote pair. It does not interpret escapes inside the quoted value. NUL is always rejected because it cannot be represented in an environment entry.
Failure and secrecy contract
Syntax errors stop startup. Diagnostics include the file path, line number, and the invalid key or error class, but never the value. A password on a malformed line must not be copied into logs.
Parsing is all-or-nothing: a syntax error returns no entries, and assignment starts only after the complete file has parsed. If the operating system rejects a validated assignment, SILO also stops startup and identifies the key and file. Since the process exits, it never serves requests with a partially loaded environment.
The file itself remains a privileged secret-bearing input. Operators must protect it with appropriate ownership and mode; parser validation is not a substitute for filesystem permissions.
Regression matrix
The committed tests cover:
- spaces and tabs around
=; - quoted values with significant spaces;
- standalone
export, including Unicode whitespace after it; - keys beginning with
_, a digit, or punctuation; - named targets using
-,.,:, and Unicode; - exact named-target discovery through the configuration subsystem;
- empty keys, whitespace, NUL, and invisible format characters;
- NUL values;
- multiple
=characters in URLs and tokens; - file-and-line diagnostics that redact values;
- all-or-nothing parse results.
The implementation passed the complete local server verification matrix and a read-only adversarial review. Windows-specific os.Setenv behavior has not been exercised on a Windows runner; unsupported platform rejection remains fail-fast rather than silent.
Compatibility and delivery
No configuration migration is required. Existing ordinary environment names behave unchanged. Files using shell-style whitespace become more predictable, and previously accepted named targets work again.
The visible compatibility changes are intentional:
- invalid or invisible names now fail instead of being silently ignored;
- unquoted surrounding value whitespace is trimmed; quote it when significant;
- malformed input stops startup with a redacted location-aware error;
- a valid punctuation-bearing target is no longer rejected merely because a shell could not assign it with
NAME=valuesyntax.
The parser contract is included in Server 20260903. Verify the artifact running in each deployment separately.
Conclusion
Configuration compatibility depends on validating the format SILO actually consumes. MINIO_CONFIG_ENV_FILE borrows a small amount of dotenv-like syntax for operator convenience, but it is not executed by a shell. The repair restores named-target compatibility while retaining strict NUL, invisibility, redaction, and fail-fast guarantees.
The original #65 also exposed a restart trap: the old parser stored KEY = new as a trailing-space key KEY . A systemd cold start could work because systemd parses EnvironmentFile itself; an admin API re-exec inherited KEY=old, and the malformed new key did not replace it. Restart could succeed with stale credentials or KMS/IdP settings. The repair makes whitespace assignments take effect, so review their intended values before upgrading.
21 - Two SSE-C Keys, One CopyObject Response
Release check (2026-09-16): the original repair described here is included in Server 20260903. Dated review and test accounts below record their original evidence, not a still-pending release or acceptance of a particular production installation. Later source changes and component selections are in the version matrix.
This record explains the CopyObject SSE-C checksum response repair committed in SILO as e73436c99.
Status on 2026-08-28: implementation, encryption and key-rotation tests, complete server suites, race tests, static checks, build, and independent Fable Max acceptance review are complete. The commit was merged into
mainon 2026-08-29 ase73436c99; tag, package, image, deployment, and production verification remain separate gates.
Scope: the successful CopyObject XML and HTTP response after the destination object has committed. Stored object bytes, checksum metadata, encryption format, source decryption, federation, replication, and historical objects are unchanged.
Security property: source SSE-C headers may decrypt only source state; destination SSE-C headers may decrypt only committed destination state.
Too Long; Didn’t Read (TL;DR)
An SSE-C copy can use two independent keys:
| Role | Request headers | Purpose |
|---|---|---|
| source | X-Amz-Copy-Source-Server-Side-Encryption-Customer-* |
decrypt the source object |
| destination | X-Amz-Server-Side-Encryption-Customer-* |
encrypt and later interpret the committed destination object |
SILO correctly wrote the destination with its destination key. However, after commit, both the XML generator and the generic PUT-response header helper received the complete CopyObject request. The checksum metadata decrypter intentionally prefers copy-source SSE-C headers when they are present. That priority is correct while reading the source, but wrong when interpreting the committed destination.
With source key A and destination key B:
The object and stored checksum were correct; only the successful response was incomplete. The repair constructs a destination response-header view by removing exactly the three copy-source SSE-C customer headers. It decrypts the destination checksum once, then reuses the resulting map for both XML and HTTP response headers.
Observable failure
The failure requires a checksum-bearing destination and distinct source/destination SSE-C contexts. A representative request supplies:
Before the repair:
- CopyObject returned HTTP 200;
- reading the destination with key B returned the correct body;
- stored destination checksum metadata decrypted with key B and matched the logical bytes;
- the CopyObject XML and HTTP response omitted CRC32 and
ChecksumType.
This is a response-contract defect, not evidence of corrupted object data.
The same ambiguity affects same-object SSE-C key rotation. After metadata has been resealed under key B, the request still carries source key A in the copy-source headers. Response generation must describe the post-rotation object, so it must use B.
Why the global decrypter must not change
The metadata decrypter’s copy-source priority is not itself a bug. Earlier in CopyObject, the server examines source checksum metadata to decide whether to preserve its algorithm, recompute a full-object value, or add the default CRC64NVME checksum. For an SSE-C source, that metadata is protected by the source object key and therefore requires the copy-source headers.
Changing the global priority to prefer destination SSE-C headers would fix the final response while breaking source checksum interpretation. The safe boundary is temporal and object-specific:
The repair applies only at that post-commit boundary.
Selected implementation
Destination response view
The handler clones the request headers and removes exactly:
X-Amz-Copy-Source-Server-Side-Encryption-Customer-Algorithm;X-Amz-Copy-Source-Server-Side-Encryption-Customer-Key;X-Amz-Copy-Source-Server-Side-Encryption-Customer-Key-MD5.
Regular destination SSE-C headers remain. SSE-S3 and SSE-KMS destination metadata needs no customer key and continues through the existing path.
Decrypt once, project twice
Before the repair, CopyObject called decryptChecksums once while building XML and again while writing success headers. For SSE-S3 or SSE-KMS this could repeat KMS unseal work.
The repaired flow is:
The generic setPutObjHeaders wrapper remains available to PutObject, CompleteMultipartUpload, and DeleteObject. CopyObject calls a narrow helper that accepts the already decrypted checksum map. ETag, VersionID, delete-marker, lifecycle prediction, and checksum header behavior remain in one shared implementation.
Regression matrix
The tests cover:
- plaintext source to SSE-C destination;
- compressed and uncompressed SSE-C destinations;
- SSE-C source key A to destination key B;
- checksum value and type in both CopyObject XML and HTTP headers;
- stored checksum decrypted with destination key B;
- destination body readable with B;
- same-object key rotation from A to B;
- checksum response after rotation;
- SSE-S3 source and destination combinations;
- all object-layer backends used by the API test harness.
The final combined tree passed focused encryption tests, the complete cmd and internal suites, the project’s tagged test configuration, full go test -race ./..., vet, lint, generated-file checks, rebrand guards, and a local build. A mirror Fable Max review reported no P0–P2 findings and independently confirmed that source decryption still receives the full request while destination response decryption receives the filtered view.
Compatibility and operational impact
- Successful CopyObject responses: checksum fields that were previously missing now appear when the committed destination has a checksum.
- Stored objects: no rewrite, migration, metadata-format, or encryption-format change.
- Existing objects: unaffected; the defect existed only in the one-time successful response.
- Clients: no request change. Clients already providing both source and destination SSE-C keys receive a more complete S3-compatible result.
- Performance: one metadata checksum decryption instead of two; no additional object read or hash pass.
- Rolling upgrade: old nodes may omit the fields while new nodes return them. Stored objects remain mutually readable.
- Rollback: restores response omission but does not damage objects created while the repair was present.
- Security: no key or digest value is added to logs or error responses. The response carries only the checksum already authorized for the successful write.
This repair does not resolve the separately deferred legacy federation CopyObject branch and does not audit or modify historical compressed-object checksums. Those questions have different data and operational boundaries.
Conclusion
CopyObject is one request with two object identities. Reusing the full request after commit erased that distinction: a source key was allowed to shadow the destination key while describing destination metadata. The durable repair is not a new encryption scheme; it is an explicit context boundary, followed by one decryption and two faithful response projections.
22 - Why CompleteMultipartUpload Must Return ChecksumType: Review of PR #57
PR #57, contributed by Shooks (@Dansyuqri), fixed #47. The repair is included in Server 20260903. This records the response contract; production deployment remains specific to the artifact and installation an operator actually runs.
The defect and its scope
Investigation of @cbornet’s #31
separated two defects. Multipart CRC32 completion itself was repaired by
c8590413f and 3e14733f1. After successful completion, the object retained its
checksum type and HEAD could return it, but completion XML omitted that type.
SDK callers therefore saw a missing ChecksumType beside a valid checksum value.
This omission did not corrupt stored data. It made a full-object checksum and a composite checksum harder to distinguish. Their Base64 encodings do not tell a consumer which calculation to reproduce.
Response contract
| Stored state | Completion response |
|---|---|
| Full-object checksum | ChecksumType=FULL_OBJECT with its algorithm value |
| Composite multipart checksum | ChecksumType=COMPOSITE with its algorithm value |
| No additional checksum | No ChecksumType element; no invented checksum |
The S3 completion API defines the two type values. ETag is a separate field and is not a substitute for the additional S3 checksum.
Implementation and evidence
The production change adds ChecksumType string with
xml:"ChecksumType,omitempty" to CompleteMultipartUploadResponse and assigns
cs[xhttp.AmzChecksumType] after decoding the stored checksum metadata. The
generator reuses the same state as the other checksum APIs; it does not
recalculate content or infer a type from part count.
The merged PR
contains response tests for full-object, composite and absent checksums. The
original change also registered the exported field in the then-current rebrand
inventory. That inventory’s exported-symbol section was subsequently removed by
bc3b35f97; it is not a current public API compatibility guarantee.
Compatibility and adjacent work
Readers that ignore unknown XML elements remain compatible. No new algorithm,
stored metadata format, object migration or checksum bypass was introduced.
The UploadPart repair and
completion validation
are separate changes, also included in Server 20260903. In particular,
CRC64NVME + COMPOSITE is rejected, not silently normalized.
23 - When the Total Is Unknown: Folder Download Progress
Historical scope: this is the Console 2.2.0 progress repair, included in Server 20260903’s embedded Console. Console 2.4.1 subsequently added file streaming and native browser handoff, eliminating the complete JavaScript ZIP Blob. Transport descriptions and follow-ups in the original PRD below refer to the 2.2.0 design point.
Status: Shipped in SILO Console 2.2.0 (
16960f7ab); the server embeds it since its Console pin was updated (edc8be6ed) · Priority: P1 · Owner:pgsty/silo-console· Related issue:pgsty/silo#62· PRD review: Claude Fable 5 (xhigh) — APPROVE · Implementation review: Claude Fable 5 (xhigh), 2026-08-23 — APPROVE, no P0/P1/P2 findings
SILO Console shows NaN% in Downloads / Uploads while downloading a folder. The ZIP normally keeps streaming and the stored objects are intact, but the progress bar has crossed from “unknown” into an invalid determinate state. Users see a full-looking bar, assume the transfer failed or finished, and retry it.
The proposed repair is intentionally narrow:
A download may enter determinate mode only when it has a finite, positive total measured in bytes applicable to that response. Without such a total, it remains indeterminate until completion, failure, or cancellation.
The server keeps streaming ZIPs. Ordinary files keep their percentages. The frontend gains one safe calculation boundary, reuses its existing indeterminate renderer, and closes one missing cancellation transition. This record defines why that is both sufficient and the smallest truthful fix.
The observed failure
The defect was observed in the then-current silo-console v2.1.1, which is embedded by Silo RELEASE.2026-08-06T00-00-00Z.
Reproduction:
- Put several objects below a prefix such as
folder/. - Stay in the parent listing, select
folder/, and click Download. - Open Downloads / Uploads before the transfer finishes.
- The row displays
NaN%; the ZIP request continues.
The runtime check used a prefix containing about 88.7 MiB and throttled Chromium to preserve the observation window. Two independent downloads produced the same NaN% state.
This is a frontend correctness bug. It is not evidence of corrupted objects, an altered disk format, or a failed S3 GET.
What is actually happening
The visible NaN% is the end of a contract mismatch across three layers.
A prefix has no object size
S3 folders are common prefixes, not stored directory objects. In the listing model, a prefix ends in / and carries size=0. The Console already renders that size as -, correctly treating it as not applicable.
The generated API model marks size as omitempty, so logical zeroes are absent from listing JSON. The single-selection thunk nevertheless passes object.size straight into the download helper: a prefix or zero-byte object therefore supplies undefined at runtime (while synthetic prefix records may supply 0). Neither value is a valid denominator.
A streamed ZIP has no known wire length
The server recognizes the trailing /, recursively lists the objects, then connects a zip.Writer to an io.Pipe. Objects are read, deflated, and copied to the HTTP response as the archive is produced.
That behavior is desirable: the server can send the first bytes without holding the complete archive in memory or on disk. Its consequence is equally deliberate: the final compressed byte length does not exist when headers are sent, so the response has Content-Type: application/zip and a filename, but no Content-Length.
The sum of source object sizes is not a substitute. Source sizes are uncompressed bytes; ProgressEvent.loaded counts response bytes after ZIP compression and framing. They are different units.
A progress event does not imply a computable percentage
The client currently computes every event as:
For a prefix, the denominator is zero or absent. Depending on the value and event, JavaScript produces NaN (loaded / undefined or 0 / 0) or Infinity (positive bytes divided by zero).
The progress callback then writes that non-finite value into Redux and sets waitingForFile=false. That second operation is the decisive state error: the task leaves the existing indeterminate branch merely because an event arrived, not because the event contained a usable total. The determinate progress component receives the invalid value and renders an invalid label.
The complete chain is:
Ordinary non-empty files avoid the defect because the server can stat the object, sets Content-Length, and the list size is positive. If the browser emits a progress event for an empty response, a zero-byte file reaches the same arithmetic boundary as a prefix even though it is a real object; it therefore belongs in the regression contract.
Product contract
The UI needs one honest distinction:
- Determinate means both transferred bytes and total bytes are known in the same unit.
- Indeterminate means the request is active but the total is unknown.
This yields four load-bearing invariants:
These invariants are more general than objectPath.endsWith("/"): they cover prefixes, zero-byte files, malformed metadata, and any future unknown-length response without inventing object-type exceptions.
Goals and non-goals
Goals
- A folder download never displays
NaN%,Infinity%, or a fabricated percentage. - Unknown-length transfers use the existing indeterminate animation.
- Known-length ordinary files retain their current percentage behavior.
- Completion, failure, and cancellation always leave indeterminate mode.
- A zero-byte file never produces a non-finite percentage and still reaches success.
- No non-finite or out-of-range download percentage enters Redux.
- The fix can ship in Console first and then be consumed by Silo as a dependency update.
Non-goals
- Do not pre-generate or buffer a complete ZIP on the server.
- Do not use the sum of uncompressed object sizes as network progress.
- Do not redesign the entire Object Manager state model.
- Do not route folders through the current immediately-completing
BrowserDownloadpath. - Do not solve the browser memory cost of
XMLHttpRequest.responseType="blob"here. - Do not change whether a cancelled row remains visible until the user clears it.
- Do not redesign mid-stream ZIP error signaling after HTTP headers have been sent.
- Do not modify the S3 API, Console API, object layout, or archive contents.
Those are legitimate follow-ups, but coupling them to this defect would enlarge risk without being necessary to restore truthful progress.
The decision
The minimum production repair has four parts.
D1. Calculate only from a valid total
Add a small pure function, separate from DOM and Redux side effects:
The source priority preserves compatibility:
- A finite positive
objectSizeretains the current ordinary-file calculation. - If object size is unavailable but the browser declares the response length computable and supplies a finite positive
event.total, use it. - Otherwise return
null: no truthful percentage exists yet.
The helper’s output contract is complete: either null, or a finite number in [0,100].
D2. Keep unknown totals indeterminate
Change the XHR handler to dispatch only a real percentage:
Download rows already start with waitingForFile=true, and ObjectHandled already renders that state with variant="indeterminate". There is no need to widen Redux to number | null, add another boolean, or change MDS.
When the first valid percentage arrives, the existing updateProgress action stores it and sets waitingForFile=false. When no valid percentage ever arrives, the row remains indeterminate until a terminal action.
D3. Make cancellation terminal
Completion and failure already clear waitingForFile. Cancellation does not. Add the missing transition in cancelObjectInList:
Without that line, the repaired prefix download would remain in the indeterminate rendering branch after abort, masking the Cancelled state. The row continues to follow the current product behavior: it remains as a cancelled record and can be removed manually. Automatic removal is not part of this change.
There is one event-order guard at the XHR boundary as well. abort() first produces readystatechange(DONE, status=0) and only then the abort event; without a status-zero return, the generic DONE branch marks the request failed before onabort can mark it cancelled. DONE/status zero is therefore left to the dedicated onerror or onabort handler, and onabort removes the stored request reference.
D4. Normalize an omitted zero-byte size
The single-selection thunk passes object.size || 0, matching the other download entry point. This restores the API model’s omitted logical zero before the helper checks Blob.size === fileSize, so an HTTP 200 zero-byte object completes at 100% instead of being reported as incomplete.
D5. Keep the server stream unchanged
The folder handler continues to generate a deflated ZIP through io.Pipe and omit Content-Length. No API, archive, storage, or resource-management contract changes.
State machine
| State | waitingForFile |
percentage |
Terminal flag | Rendering |
|---|---|---|---|---|
| Queued / no valid progress yet | true |
0 |
none | indeterminate |
| Unknown-total transfer | true |
0 |
none | indeterminate |
| Known-total transfer | false |
0..100 |
none | determinate percentage |
| Completed | false |
100 |
done=true |
success |
| Failed | false |
last value | failed=true, done=true |
error |
| Cancelled | false |
0 |
cancelled=true, done=true |
cancelled |
The state does not move back from determinate to indeterminate. If a later event lacks a valid total after a valid percentage was observed, the handler simply retains the last valid value.
Failed and Cancelled both set done=true in the existing reducers. ObjectHandled uses done to change its close button from “abort request” to “remove record”; this repair preserves that behavior. The cancelled Redux value remains 0, while the existing ProgressBarWrapper renders a full orange terminal bar with a Cancelled label because ready=true. That established presentation is not part of this repair.
waitingForFile is not the ideal long-term name for “no computable progress.” Renaming it or replacing the booleans with a discriminated union would improve the model, but that is a separate refactor. In this repair, the field already expresses and renders the required state, so reusing it minimizes compatibility risk.
Why this is sufficient
The repair closes the bug by cases.
Ordinary non-empty file
objectSize > 0, so the helper uses the same denominator as today. The result is finite and clamped, updateProgress enters determinate mode, and completion still sets 100%.
Current streamed folder
objectSize is normalized to 0, while lengthComputable=false and event.total=0. The helper returns null; no invalid action is dispatched, so the row remains indeterminate. Completion sets waitingForFile=false, percentage=100, and done=true.
Future response with a real length
If a proxy or later server implementation provides a trustworthy response total, lengthComputable=true and event.total>0. The same code automatically produces a real percentage without another product change.
Zero-byte file
The omitted listing size is normalized to zero, and both totals are then zero, so an intermediate percentage is mathematically undefined. The row stays indeterminate for its usually brief lifetime; the zero-byte Blob now equals the normalized expected size, and the successful response transitions directly to 100%. 0/0 is never evaluated.
Failure and cancellation
Failure already exits indeterminate. The added cancellation transition does the same on abort. No terminal row can continue to look active merely because its total was unknown.
Mathematically, division occurs only when total belongs to (0, +infinity). The result is then clamped to [0,100]. Therefore neither NaN nor Infinity can cross the calculation boundary into Redux or the determinate renderer.
Rejected alternatives
Buffer the ZIP to obtain Content-Length
The server could generate the complete archive in memory or a temporary file, measure it, and then send it. That would provide an exact wire total, but at the cost of memory or disk pressure, delayed first byte, cleanup complexity, and worse concurrent-download behavior. An observability defect does not justify discarding streaming.
Sum the objects under the prefix
That sum is uncompressed logical data. event.loaded measures compressed response bytes plus ZIP framing. The units differ, so the bar could stop below 100%, exceed 100%, or move according to compression ratio rather than transfer completion. Reject.
Convert invalid progress to 0%
This hides the string but lies about the state: determinate 0% means the total is known and no portion has transferred. Users would still interpret the transfer as stalled. Unknown must remain unknown.
Special-case paths ending in /
That fixes the reported prefix but misses a real zero-byte object, invalid metadata, and other unknown-length responses. The correct boundary is denominator capability, not object type.
Send folders through BrowserDownload
The current large-file path creates an anchor and immediately calls the completion callback after clicking it. It cannot report true completion, console-managed cancellation, or a subsequent HTTP failure. It may be the basis of a later streaming-download design, but today it would replace one lie with another.
Sanitize inside ProgressBar
A generic component guard could be useful defense in depth, but it would leave invalid data in Redux and hide the broken state transition from every other consumer. The primary repair belongs where progress becomes application state.
Introduce percentage: number | null now
A discriminated progress state would be cleaner than the current booleans if the Object Manager were being redesigned. Adding null while retaining waitingForFile, done, failed, and cancelled would instead create more contradictory combinations. Removing the old fields is larger than this bug requires. Reuse the already-rendered indeterminate state now; redesign it separately.
Requirements and acceptance
Functional requirements
- FR1: An unknown total keeps the task indeterminate.
- FR2: A finite positive object size preserves ordinary-file percentages.
- FR3: A finite positive
event.totalis a fallback only whenlengthComputable=true. - FR4: Every dispatched percentage is finite and within
[0,100]. - FR5: A zero-byte file never displays non-finite progress and reaches success.
- FR6: Completion, failure, and cancellation leave indeterminate mode.
- FR7: Versioned objects, anonymous downloads, previews, and long-filename entry points retain their existing call contract.
Non-functional requirements
- No new server CPU, memory, disk-buffer, or request cost.
- No new frontend dependency or build step.
- No change to the S3 API, Console API, ZIP content, or stored objects.
- The calculation must be testable without a DOM or live store.
- TypeScript typecheck and the production frontend build must pass.
Acceptance criteria
- While a folder ZIP without
Content-Lengthis active, its row shows an indeterminate animation and no percentage text. - On successful completion, the row reports success/100% and the ZIP can be opened.
- A normal non-empty file continues to show finite determinate progress and completes at 100%.
- A zero-byte file never shows
NaN%orInfinity%and completes successfully. - Cancelling an unknown-total download aborts the request and shows Cancelled, not an active animation.
- No download path can place a non-finite or out-of-range percentage in Redux.
Test plan
Pure calculation matrix
Use the existing @playwright/test runner for the pure module rather than adding a test framework. This needs one config-only addition in web-app/playwright.config.ts: a dependency-free unit project, for example with testMatch: /.*\.unit\.ts/. The existing chromium project depends on the auth setup against a live Console at localhost:9090; pure calculation and reducer tests must not be gated by that environment. No new dependency is introduced.
| Case | loaded |
objectSize |
lengthComputable |
event.total |
Expected |
|---|---|---|---|---|---|
| Ordinary file, halfway | 50 | 100 | false | 0 | 50 |
| Common prefix | 1024 | 0 | false | 0 | null |
| Initial zero over zero | 0 | 0 | false | 0 | null |
| Response-total fallback | 50 | 0 | true | 200 | 25 |
| Zero total is unusable | 0 | 0 | true | 0 | null |
| Loaded exceeds total | 150 | 100 | true | 100 | 100 |
| Invalid object size | 10 | NaN |
false | 0 | null |
| Omitted zero size | 10 | undefined |
false | 0 | null |
| Invalid response total | 10 | 0 | true | Infinity |
null |
| Negative loaded | -1 | 100 | true | 100 | null |
State tests
Cover the transition contract directly:
- A new download starts with
waitingForFile=true. - No valid progress action means it remains indeterminate.
- Valid progress produces a finite value and
waitingForFile=false. - Complete produces
done=true,waitingForFile=false,percentage=100. - Failure produces
failed=true,done=true,waitingForFile=false. - Cancel produces
cancelled=true,done=true,waitingForFile=false,percentage=0.
Browser regression
Use the real Console test instance and Chromium:
- Create a temporary bucket with several objects below
folder/. - Select the prefix from its parent and start the download.
- Apply CDP download throttling so the intermediate state is observable.
Throttled runs must raise the default 30-second test timeout with
test.setTimeout. - Open Downloads / Uploads and verify that the row exists, has no percentage label, and contains neither
NaN%norInfinity%. - Cancel it and verify the Cancelled terminal state.
- Restore network conditions in
finally. - Download again without throttling, wait for the browser download, and verify the ZIP.
- Repeat the relevant assertions for one ordinary non-empty file and one zero-byte file.
- Remove the bucket, objects, downloads, and temporary files in teardown.
The current Playwright project is Chromium-only, so CDP is an acceptable test mechanism. If Firefox or WebKit projects are later enabled, keep the pure and state tests cross-browser and gate only the throttled observation behind the Chromium project.
Implementation boundary
Expected Console changes:
- Add
downloadProgress.tscontaining the pure calculation. - Change
Objects/utils.tsto dispatch only a non-null percentage, let status-zero terminal events reach their dedicated handlers, and clean up an aborted request. - Normalize omitted zero sizes in the single-selection thunk.
- Change
cancelObjectInListto clearwaitingForFile. - Add calculation, state, and browser regression coverage using existing dependencies, with a dependency-free
unitproject inplaywright.config.ts.
Expected unchanged code and contracts:
- The Go folder-download handler and its streaming ZIP.
ObjectHandled,ProgressBarWrapper, and MDS.IFileItem.percentage: numberand the existing thunk callback types.- S3 and Console API routes.
- Stored object and archive formats.
Delivery and rollback
The fix belongs in pgsty/silo-console, not the Silo server repository where the issue was reported.
Delivery order:
- Transfer or cross-reference issue #62 to
pgsty/silo-console. - Implement the bounded Console change.
- Pass typecheck, production build, pure/state tests, and real browser regression.
- Publish a new Console release.
- Update Silo’s pinned Console pseudo-version or release dependency.
- Build a Silo candidate and repeat folder, ordinary-file, zero-byte, cancel, and ZIP-integrity checks.
- Publish Silo and record both affected and fixed versions on the issue.
There is no data migration. If the frontend change regresses, Silo can roll back only the Console dependency; server data and API behavior remain compatible.
Definition of done
- The calculation returns only
nullor a finite[0,100]number. - Active unknown-total folder downloads render indeterminate.
- Ordinary files retain determinate progress.
- Zero-byte files never render invalid progress.
- Complete, failed, and cancelled rows all leave indeterminate mode.
- The streamed ZIP and server response contract remain unchanged.
- Typecheck, production build, and automated regressions pass locally.
- Console 2.2.0 is published.
- Server 20260903 includes the updated Console; see its release evidence.
Follow-up work
Five adjacent improvements deserve separate design records:
- Stream large folder downloads directly to the browser or filesystem instead of holding the full Blob in memory.
- Replace the Object Manager’s boolean combination with a discriminated progress/terminal state.
- Improve end-to-end integrity and error signaling for ZIP failures after headers have been sent.
- Add a generic non-finite-value guard to shared progress components as defense in depth.
- Repair the pre-existing Blob JSON error decoder and request-trace cleanup on HTTP failure paths.
None is required to stop the current UI from lying. The next maintenance iteration should first restore the smallest honest contract: known totals get percentages; unknown totals remain unknown.
Later implementation and the original design
The Console 2.2.0 integration also made size an always-present JSON field and repaired ZIP error propagation: unreadable entries are no longer silently skipped; errors before output can return 500 and errors afterward interrupt the stream. The no-API/resource-management-change statements above describe only the original progress-calculation patch, not the whole release.
Current single-folder downloads are handed to the browser. A completed Console row records the handoff, not completion of all bytes; track and cancel the transfer in the browser download manager. Multi-selection file-writer/native handoff paths are described in Console 2.4.1. Size normalization remains defensive support for old responses, not evidence that the current model omits zero values.
24 - A Listing Must Not Drop a Null Version That Still Has Quorum
Publication update, 2026-09-20: The null-version quorum repair shipped in Server 20260916. The rolling-restart limitation tracked in #218 remains open. The dated investigation and validation record below retains its original scope.
This record covers the listing omission investigated on 2026-09-16: an unversioned bucket under concurrent overwrites returned HTTP 200 listings that lacked keys the same endpoint could still GET. It describes the confirmed mechanism, the narrow repair in commit 8d06424b1, what that repair proves, and the counterexample that remains open in #218.
Status: on main after the published Server 20260903. The repair was reviewed in three external review rounds before implementation and validated with deterministic order counterexamples, signed HTTP listings and differential inputs. It does not close the listing problem; see the boundary.
The defect
Listing merges the version streams of the drives that were asked. When the drives disagree, mergeXLV2Versions picks a top version and counts the streams that agree with it. A 2022 upstream change (PR #14125) made that count tolerant of signature differences left by healing, which is necessary, but the way it selects and counts is sensitive to input order. With one ordinary null version per drive, a newer minority that sorts last resets the count for the older generation. The older generation that actually has quorum is then discarded, the merge returns nothing, and the resolver reports the key as absent. The request still succeeds, so the client sees a complete-looking listing with a hole. Reads of the key succeed because the read path resolves quorum on its own.
Nine drive orders on the old code failed this way, with complete signed LIST evidence, while the same content in other orders listed correctly. The defect also exists in the published Server 20260903.
The repair
The recount is added only where the original selection ends without quorum, and only for inputs where it can be exact: every non-empty drive stream holds exactly one ordinary null version, none is a free version, and all share the same erasure parameters, judged on the original inputs before any stream is pruned. In that case the versions are regrouped by header and a group is returned only if it reaches the original effective quorum. Mixed histories, explicit version IDs, delete markers, free versions and mixed erasure layouts keep their previous behavior; the strictness, signature, tie-break and representative-entry rules are unchanged, and shared callers of the merge see the same result for every previously successful input.
Validation: the nine old-order and seven new-order counterexamples pass; 5,620 differential inputs keep the contracted result; full signed HTTP listings, object reads, normal and race runs, and the scanner, healing and migration consumers of the merge were checked. Before integration into main, six top-level and 23 nested cases passed in normal and race mode. A stable-state run of 20,000 successful LISTs under concurrent overwrites on a four-node, sixteen-drive cluster omitted nothing.
What remains open
On the same cluster, a rolling restart of the four nodes under concurrent overwrites returned 27,966 successful LISTs, of which 2,624 omitted one to four keys. The keys were readable: 8,186 of 8,192 same-endpoint GET/HEAD checks returned 200 with the last confirmed content. The strongest sample started 17.6 s after the last node’s health check returned and omitted a key whose last successful PUT had been confirmed 11.5 s earlier with no later write. Raw XML, HTTP correlation and body hashes were re-verified independently, so pagination and parsing do not explain it.
The repaired build still has exits that turn “cannot decide” into “absent” without failing the request: a merge in which no generation reaches quorum, a resolver with fewer valid entries than quorum or no cached metadata, the partial-resolution callback that drops the entry and only fails the listing when more readers have failed than the quorum allows, and a walker that skips an entry whose metadata read fails unexpectedly. The listing quorum is computed from the number of drives asked and does not follow readers that fail mid-listing. Which of these the sample hit cannot be told from HTTP evidence; the per-drive streams, effective quorum, recount eligibility and cache state of the failing request were not captured, and metadata recovered after the run is not the failing input.
The next step is a bounded diagnostic capture of one failing LIST, then a deterministic regression built from that input, and only then a decision on whether the contract changes to fail the request or return a three-state result. Until then, do not run tools that delete destination objects missing from a source listing during rolling restarts, and list again once the cluster is stable.
25 - A ListObjects Shortcut Must Not Turn a Missing Bucket into an Empty One
Release check (2026-09-16): the original repair described here is included in Server 20260903. Dated review and test accounts below record their original evidence, not a still-pending release or acceptance of a particular production installation. Later source changes and component selections are in the version matrix.
This document records the problem analysis, design discussion, and repair decision for SILO #32 and PR #37.
Status on 2026-08-26: PR #37 was updated to the DCO-signed head
e9c5340be, formally approved, and merged as49c8aeac4; #32 closed automatically. DCO, VulnCheck, and all six Go CI jobs passed on the exact PR head; the post-mergemainVulnCheck and all six Go CI jobs also passed. No tagged release, package, container image, deployment, or production endpoint has yet been verified to contain the repair.
Scope: verify bucket existence only for three listing shortcuts that bypass storage; do not restore the genericcheckBucketExist, change the normal listing path, or introduce an existence cache.
Release boundary: local commit, push, remote CI, merge, tag, package, container image, deployment, and production verification are independent gates.
Too Long; Didn’t Read (TL;DR)
The problem is real and worth fixing. A normal ListObjects, ListObjectsV2, or ListObjectVersions request against a missing bucket reaches storage and receives BucketNotFound. Three inputs, however, return early:
- a marker outside the prefix;
max-keys=0;- a prefix beginning with
/, including thePrefix="/"boto3 reproduction from #32.
Those branches return io.EOF directly. The caller treats EOF as a successful end of listing, so the client receives an empty 200 rather than S3’s 404 NoSuchBucket. The identity of the same missing resource changes from an error to success solely because the selection parameters differ. That breaks S3 compatibility and blocks a real user’s upgrade from the pre-regression release.
The repair must not put an expensive bucket check back in every listing. The selected design replaces only the three bare io.EOF returns with a small helper. The helper calls GetBucketInfo once: it returns the real error if the bucket is absent or cannot be confirmed, and preserves io.EOF when the bucket exists. The normal listing hot path is untouched. Only requests that would otherwise exit before storage pay the extra peer-and-disk fan-out.
That decision has now been executed: the strengthened repair passed local review, the exact PR head passed every remote check, and the expected-head-guarded merge entered a green main.
What is the problem?
One API exposes two bucket-existence semantics
#32 reproduces the defect by calling the following against a missing bucket:
AWS S3 raises NoSuchBucket; SILO returns a successful empty listing. The difference is not in authentication, routing, or XML serialization. It comes from the object-layer listPath control flow:
/ is not the only trigger:
| Shortcut condition | Why the result must be empty | Defect before the repair |
|---|---|---|
| Marker does not begin with the prefix | The implementation does not scan this disjoint range | Returns EOF without confirming the bucket |
max-keys=0 |
The caller asks for zero keys | Incorrectly equates “zero results” with “valid resource” |
Prefix begins with / |
SILO’s flat key space produces no entries for this form | The filter short-circuits before bucket identity |
For an existing bucket, returning an empty listing from these branches is a reasonable optimization. For a missing bucket, the same EOF masks the resource error that should take precedence.
The regression has a known origin
The reporter confirmed correct behavior in RELEASE.2024-01-29T03-56-32Z and the regression beginning with RELEASE.2024-01-31T20-20-33Z. The corresponding upstream change is minio/minio#18917 / 80ca12008. It removed GetBucketInfo from generic argument checks and relied on actual Put, List, and Multipart storage operations to expose a missing bucket.
That optimization works on normal paths but leaves a gap: an early-return path never reaches the storage operation that is now responsible for producing the error. #32 does not require a broad rollback of the upstream optimization. It repairs the overlooked control-flow exits.
Why fix it?
The S3 contract explicitly requires NoSuchBucket
Both AWS ListObjects and ListObjectsV2 define NoSuchBucket as HTTP 404 when the specified bucket does not exist. prefix, marker, start-after, and max-keys select listing results; they must not turn a missing bucket identity into a successful request.
ListObjectVersions shares the same object-layer listing engine. Giving V1, V2, and version listings the same existence behavior on the same shortcut inputs prevents the three public APIs from diverging further.
An empty 200 changes client decisions
An empty 200 and a 404 are not interchangeable presentation details:
- 404 tells provisioning or test code to create the bucket, fix configuration, or stop;
- an empty 200 asserts that the bucket exists but has no matching objects;
- SDKs, synchronization tools, and integration tests continue down different branches;
- a test using SILO as an S3 substitute can pass locally and fail against AWS.
#32 also establishes a direct upgrade impact: an application relying on the older correct behavior cannot upgrade past the regression. The repair restores both S3 parity and upgrade compatibility.
The repair surface is narrow and testable
The bug is confined to three adjacent early returns. It does not involve object data, metadata formats, sorting, pagination-token encoding, permissions, or wire schemas. A very small production change can be pinned down with object-layer and HTTP-level contracts, so the benefit clearly exceeds the implementation risk.
Why not restore the global check?
Upstream did not remove generic GetBucketInfo as incidental cleanup. The motivation for #18917 states that checking the bucket before every Put, List, and Multipart operation fans out across servers; even after vectorization, the cost becomes visible beyond 100 nodes.
In current SILO, erasureServerPools.GetBucketInfo calls S3PeerSys.GetBucketInfo. That operation concurrently asks every peer and reduces quorum per pool, while each peer checks its local bucket state. It is not a cheap in-memory map lookup.
Two extremes are therefore unacceptable:
- never check: keep the incorrect empty 200;
- check before every List: restore semantics while undoing a critical large-cluster optimization.
The actual design question is whether the check can be confined to branches that never touch storage and therefore cannot discover the missing bucket naturally. It can.
How is it fixed?
Replace only three bare EOF returns
In cmd/metacache-server-pool.go, each shortcut previously executed:
It now executes:
The helper has only two classes of outcome:
- existing bucket: preserve the previous empty-list behavior;
- missing bucket: pass
BucketNotFoundinto the existing error mapping, producing HTTP 404NoSuchBucket; - state cannot be confirmed: propagate quorum, offline, timeout, or context errors instead of fabricating success.
The normal listMerged, metacache scan, sorting, pagination, and response-generation paths do not change.
Why the helper belongs here
The check must sit next to the shortcut for three reasons:
- only this layer knows that it is about to bypass every storage access;
- moving it into generic argument validation charges every call;
- moving it into the scan layer cannot help because these branches never scan.
The name intentionally states the boundary. This is not a new generic checkBucketExist; it restores missing existence semantics immediately before a shortcut returns EOF.
Do not add a cache
A bucket-existence cache could reduce fan-out but immediately creates invalidation questions for create, delete, site replication, recovery, and expiry. Adding a second source of truth for three low-frequency shortcuts costs more complexity and consistency risk than it saves.
The selected implementation uses the existing GetBucketInfo source of truth. If future telemetry shows that large clusters receive frequent max-keys=0, slash-prefix, or disjoint-marker probes, the project can evaluate a dedicated metadata fast path, rate limiting, or a carefully invalidated cache using real data rather than speculative machinery in this compatibility patch.
Test and review evidence
Object-layer contract
The object-layer test runs against single-drive and multi-drive erasure setups and exercises four inputs:
- slash-prefixed prefix;
- zero limit;
- marker outside prefix;
- a regular prefix as a control that still receives the error naturally from storage.
Each case covers ListObjects, ListObjectsV2, and ListObjectVersions, using the typed isErrBucketNotFound predicate rather than brittle English error-string comparison.
HTTP contract
The handler test sends genuine signed requests for all three public APIs:
| API | Request shape | Assertion |
|---|---|---|
| ListObjects | GET /missing-bucket?prefix=/ |
HTTP 404 and XML code NoSuchBucket |
| ListObjectsV2 | Add list-type=2 |
HTTP 404 and XML code NoSuchBucket |
| ListObjectVersions | Add versions |
HTTP 404 and XML code NoSuchBucket |
The HTTP test uses the real slash-prefix reproduction from #32. The other two shortcuts are enumerated at the object layer. This proves final wire behavior without repeating the full matrix in the slower handler fixture.
Local quality gates
The improved local commit passed:
The full local cmd test completed in 116.215 seconds. An independent local Claude Code review used the Fable model at Max effort to inspect the exact tree, call paths, error mapping, tests, performance boundary, and this decision. Its verdict was GO, with no mandatory pre-merge change.
The DCO-signed PR head e9c5340be then passed eight remote checks: DCO, VulnCheck, and six jobs in Go CI. After merge, the resulting main commit 49c8aeac4 independently passed VulnCheck and all six Go CI jobs. The slowest checks were PR cross-compile at 9 minutes 47 seconds and post-merge cross-compile at 9 minutes 30 seconds.
Can it introduce new problems?
Shortcut requests now fan out across the cluster
This is the most important and deliberately accepted cost. A shortcut on an existing bucket used to be little more than a local branch; it now calls GetBucketInfo. Directional local microbenchmarks observed:
| Path | Observed magnitude |
|---|---|
| Shortcut before the repair | about 0.55 μs, 7 allocations |
| Repaired single-drive shortcut | about 7.8–8.1 μs, 45–47 allocations |
| Repaired 32-drive shortcut | about 70–81 μs, 977 allocations |
| Normal 32-drive listing | about 0.95 ms |
These numbers show local relative cost only; they are not a latency prediction for a 100+ node deployment. Real distributed execution adds peer networks, quorum, and slowest-node tail latency, potentially making the gap much larger. That is precisely why the check must not expand into the normal listing path.
The risk concentrates in malformed or probe-style traffic. A misconfigured client polling max-keys=0, a slash prefix, or disjoint markers at high frequency can amplify what was a cheap request into peer-and-disk work. After merge, the actual frequency of these inputs should be observed through S3 traces or metrics; rate limiting or optimization should follow evidence.
A degraded cluster exposes more real errors
Previously, a shortcut could return an empty 200 while peers were offline or bucket quorum was unavailable because it never consulted cluster state. The repair can return quorum, timeout, or service errors in those conditions.
That is more honest behavior, not an availability regression: if the server cannot establish that the bucket exists, it must not assert a valid empty bucket. Clients depending on unconditional empty success will nevertheless observe a behavior change.
Bucket create/delete races are not linearizable
GetBucketInfo and returning the empty result are two actions. The bucket can be deleted immediately after the check, or created immediately after a missing-bucket result is formed. This patch does not and should not add a transaction spanning bucket lifecycle to a listing shortcut.
This is the same concurrency class as other APIs that validate a resource before acting. The repair guarantees that the request no longer succeeds with no existence evidence at all; it does not promise a cross-node, cross-lifecycle linearizable snapshot of an empty listing.
Clients relying on the bug will receive 404
Some clients may have adopted the missing bucket’s empty 200 as fact. They will now enter an error branch. This is a visible compatibility change, but it restores the documented S3 contract and the pre-regression behavior. Preserving the bug merely transfers upgrade cost to clients that correctly rely on 404.
Two adjacent edges remain out of scope
The adversarial review recorded two non-blocking P3 boundaries:
- When resuming a metacache continuation, the
c.fileNotFoundbranch still returns bareio.EOF. A stale or crafted continuation token used after bucket deletion could theoretically receive an empty 200. AddingGetBucketInfothere would affect normal continuation traffic and needs a separate performance and error-precedence design. - Some V1 and version-list marker/prefix combinations return
NotImplementedduring HTTP handler validation before reaching the object layer; the V2start-afterroute can reach it. This patch fixes storage shortcuts masking a missing bucket; it does not redefine precedence between malformed parameters and resource errors.
Neither blocks merge. The first is outside #32’s ordinary initial-list reproduction; the second is inherited handler behavior. Recording them prevents “all three shortcuts are covered” from being overstated as byte-for-byte AWS parity for every possible parameter combination.
Alternatives considered
Keep upstream behavior
This has zero performance change and minimizes fork divergence. It also keeps a documented S3 incompatibility, a regression with a known release boundary, and a misleading result when SILO is used as an integration-test substitute. For a narrow and well-tested compatibility repair, that tradeoff is no longer justified.
Restore generic checkBucketExist
This covers every path at once but reintroduces peer fan-out into every Put, List, and Multipart operation, directly undoing the large-cluster optimization from #18917. The cost is disproportionate and the option is rejected.
Fix only Prefix="/"
That passes the single issue reproduction but leaves the same root defect in max-keys=0 and marker-outside-prefix. The branches are adjacent and share the same semantics, so one helper is simpler and less likely to regress.
Add a bucket-existence cache
This makes shortcuts cheaper but requires semantics for create, delete, replication, recovery, and stale TTL windows. There is no telemetry showing enough shortcut traffic to justify that complexity, so it is not selected.
Complexity and cost-benefit
| Dimension | Assessment | Rationale |
|---|---|---|
| Production-code complexity | Low | Three call sites and a seven-line helper; no new state, dependency, or format |
| Test complexity | Low to medium | V1, V2, versions, three shortcuts, a control, and HTTP mapping all need coverage |
| Normal-path risk | Very low | No check is added to the listMerged hot path |
| Shortcut runtime cost | Materially higher | A local EOF becomes cluster-wide GetBucketInfo |
| Compatibility value | High | Restores 404 NoSuchBucket, pre-regression behavior, and S3 test fidelity |
| Operational complexity | Low | No migration, configuration, feature flag, cache, or cross-repository dependency |
The overall cost-benefit is favorable. The reason is not that GetBucketInfo is cheap—it is not—but that its cost is strictly limited to three shortcuts that otherwise cannot discover the missing bucket. A narrow performance cost in exchange for explicit protocol correctness is better than either a global rollback or indefinitely preserving the incorrect behavior.
Acceptance decision and remaining gates
The final decision was: accept and merge the strengthened PR #37 revision without expanding the production scope.
The accepted sequence was:
- replace the old fork head with the current-
main, DCO-signed revision while preserving Jason Lin as a co-author; - retain typed error predicates, V1/V2/version-list object-layer coverage, and HTTP-level 404 /
NoSuchBucketassertions; - update the PR description with the shortcut fan-out cost and unchanged normal-path boundary;
- approve the fork workflows and require all eight reported checks to pass on exact head
e9c5340be; - submit a formal approving review against that head;
- merge with an expected-head guard, producing
49c8aeac4, automatically close #32, and require the resultingmainGo CI and VulnCheck to pass independently.
No cache, feature flag, additional abstraction, or continuation-token redesign was required. High-frequency shortcut traffic and large-cluster tail latency remain observability follow-ups, not reasons for speculative code expansion.
Repository integration is complete. A tag, package, docker.io/pgsty/silo image, deployment, and real S3-client verification must still complete before the repair can be described as delivered to users.
Conclusion
The issue is not merely “a slash prefix reports the wrong error.” The listing engine uses io.EOF to mean two different things: an empty result from an existing bucket and an early exit that never established whether the bucket exists. Removing generic existence checks for large-cluster performance was a sound upstream optimization, but the shortcuts violate its premise that a real storage operation will naturally surface a missing bucket.
The selected repair restores that premise by calling the existing GetBucketInfo only at three storage-bypassing exits. It makes those requests more expensive and exposes real errors on degraded clusters; both are explicit costs. In return, SILO restores S3’s 404 semantics, upgrade compatibility, and test fidelity while preserving the upstream optimization on the normal listing hot path.
This worthwhile, controlled compatibility fix is now merged and green on main; the repair is included in Server 20260903; production deployment remains installation-specific.
26 - Read-Only Checksum Audit and Reliable CLI Output
This is the design and implementation record for MCLI’s read-only checksum verification workflow and pgsty/mc#5, the non-TTY output defect found during release review.
Status: shipped in the final mcli 20260903 release. The command merged to
mainthrough pull requests #8 and #13, is exercised against a real SILO server in hosted CI, and pgsty/mc#5 is closed. Server 20260903 images bundle mcli 20260903, including this command.
Owner:pgsty/mc.
Tracking: pgsty/mc#5.
Safety boundary: verification is read-only; repair is not part of this command.
Too Long; Didn’t Read (TL;DR)
Historical CopyObject implementations could calculate a stored additional
checksum over transformed storage bytes instead of the logical bytes returned
by S3. mcli checksum verify inventories objects and independently streams the
logical body through the recorded algorithm. Each candidate becomes MATCH,
MISMATCH, NO_CHECKSUM, WOULD_VERIFY (dry run), one of ten UNKNOWN_*
classifications, or one of three SKIPPED_* results.
The first implementation worked in a terminal but printed nothing when stdout was redirected. MCLI automatically marked non-TTY execution as quiet to disable progress UI, and the new command accidentally treated that internal state as a user request to suppress audit records. The repair separates semantic output from progress suppression without changing global quiet behavior or enabling progress bars in CI.
Command and scope
Version one supports CRC32, CRC32C, CRC64NVME, SHA1, and SHA256 checksums marked
as FULL_OBJECT. It can select one object, an exact VersionID, current objects
under a prefix, all versions, or exact entries from a JSON Lines manifest. It
also supports SSE-C key mappings, time and size filters, dry-run estimation,
bounded workers, download limits, JSON output, and an optional JSON Lines report.
It does not verify COMPOSITE checksums, infer type from an ETag, inspect
xl.meta, identify the historical writer with certainty, or repair metadata.
The endpoint must report the checksum type (x-amz-checksum-type) alongside
the checksum; on one that does not, every checksummed object is classified
UNKNOWN_CHECKSUM_TYPE rather than guessed at.
Read-only data path
For every selected object, MCLI:
- sends
HEADwith checksum mode enabled and retains every supported checksum plusChecksumType; - rejects unsupported or ambiguous states as
UNKNOWN_*instead of guessing; - streams
GETlogical bytes through bounded hashers without writing the body to disk; - uses VersionID pinning, or
If-Matchplus a secondHEADfor mutable unversioned/null objects; - compares independently calculated values with the stored values.
The S3 boundary allows LIST, HEAD, and GET only. Tests fail if a write method reaches the mock endpoint.
Result and exit contract
Every candidate produces one stable result:
| Result | Meaning |
|---|---|
MATCH |
Every supported stored checksum matches the returned logical bytes |
MISMATCH |
At least one stored checksum differs |
NO_CHECKSUM |
No additional checksum exists; the body is not read |
WOULD_VERIFY |
Dry-run found a supported full-object checksum |
UNKNOWN_* |
MCLI cannot make a reliable statement |
SKIPPED_* |
A filter intentionally excluded the object |
The summary carries objects, a verified count, the count of every result
status, and incomplete. verified is MATCH plus MISMATCH: the only results
that actually streamed a body through a hasher. A run that enumerated many
objects and verified none is visible as such.
--fail-on accepts mismatch, unknown, no-checksum, any, or none. The
default any returns exit 1 for mismatches and incomplete verification.
no-checksum returns exit 1 when any object carries no checksum or when
nothing was verified at all, so an empty prefix or a stale manifest cannot
pass as a clean audit. Dry-run does not apply --fail-on. Argument,
authentication, enumeration, and report-write failures remain command failures
rather than object classifications.
In particular, SKIPPED_TOO_LARGE makes the default any return exit 1 because
the size cap leaves the audit incomplete. Time-filter and delete-marker skips do
not fail by themselves.
Output and automation contract
Object records and the final summary are semantic output:
- Unless the caller explicitly sets
--quiet,-q, orMC_QUIET=true, stdout receives every object record and the final summary in both TTY and non-TTY execution. - Non-TTY
--jsonemits exactly one compact JSON value per line. TTY JSON keeps MCLI’s existing pretty presentation. - Global flags work at the app,
checksum, andverifylevels. --reportis independent of stdout. It still writes object records and the final summary as JSON Lines when explicit quiet suppresses stdout.- Output transport does not change
--fail-ondecisions.
The distinction matters because MCLI’s historical globalQuiet has two inputs:
an explicit quiet flag and an automatic non-TTY state used to disable progress
UI. Changing that global would risk re-enabling progress output across copy,
get, put, mirror, and other commands.
The selected repair is command-local. It walks the full CLI context chain for
explicit quiet/JSON flags because the CLI library’s GlobalBool stops at the
nearest ancestor flag set. It also restores JSON Lines mode inside the checksum
action because nested Before hooks can reset it after an app-level --json.
No other command’s progress or output behavior changes.
Report, secrets, and operational cost
Report files are created with mode 0600, must not already exist, and contain
metadata/results rather than object bodies or SSE-C keys. The manifest likewise
contains only bucket, key, and optional VersionID.
Verification downloads every supported object body. Operators should use
--dry-run, --max-size, time filters, --max-workers, and the global download
limit to bound cost and load. NO_CHECKSUM and UNKNOWN_* counts must remain
visible; neither may be presented as successful verification.
What a mismatch proves
A mismatch proves only that the additional checksum returned at verification time does not describe the logical bytes returned at verification time. It does not prove that a particular historical compression defect created the object, and it is not an external source-of-truth comparison.
Do not overwrite checksum metadata in place. Audit and classify first. For a
confirmed, operationally relevant mismatch, prefer a new key or new version,
verify the replacement, then switch consumers deliberately. Leave UNKNOWN_*
objects out of automatic repair.
Verification record and release boundary
The local acceptance matrix covers TTY human/JSON, non-TTY pipes, regular-file
redirects, app/parent/leaf JSON and quiet flags, environment quiet, report under
quiet, report-write failure, and MISMATCH/UNKNOWN exit status. It also includes
real historical MATCH, MISMATCH, and unsupported-composite objects on a local
S3 server.
The command shipped in the final mcli 20260903 release from a
signed tag at the tip of main, with the functional suite - including a
checksum verification run against a real SILO server - green for that commit,
and pgsty/mc#5 is closed. Server 20260903 images bundle that client. A production audit remains a separate execution with its own evidence.
JSON consumers should check schemaVersion: 1 and distinguish per-object records (type: object) from the final summary (type: summary). --max-workers defaults to 4 and accepts 1–64; it bounds concurrent object work, not total memory or server I/O.
27 - Optional Checksums, Mandatory Failure: Repairing UploadPart and UploadPartCopy Compatibility
Release check (2026-09-16): the original repair described here is included in Server 20260903. Dated review and test accounts below record their original evidence, not a still-pending release or acceptance of a particular production installation. Later source changes and component selections are in the version matrix.
This is the complete design and implementation record for SILO #46. The repair was not merely a changed if statement. One apparently optional S3 header reached into multipart completion semantics, copy responses, compression and encryption pipelines, compatibility baselines, and release verification.
Status: merged into
mainas7fea6d5a5on 2026-08-24 (pgsty/silo#46 closed); included in Server 20260903; deployment verification is installation-specific.
Owner:pgsty/silo, the SILO server repository.
Tracking: #46.
Independent follow-ups: #63 CopyObject + compression checksum, #64 federated UploadPartCopy checksum.
Adversarial review: local Claude Code, Fable 5,--effort max; final verdict GO, with no blocking findings.
Too Long; Didn’t Read (TL;DR)
A multipart upload splits a large file into smaller parts. A client may attach a checksum to each part so the server can verify the transfer, but AWS defines that checksum as optional. SILO used to treat it as mandatory: an ordinary UploadPart failed without one, and UploadPartCopy could never work because it has no part-body checksum to provide.
After the repair, SILO still validates a checksum when the client sends one. When the client omits it, SILO computes the checksum while reading the original bytes and saves the result. This happens before compression and encryption, requires no second read, and changes no on-disk format. The result is AWS-compatible behavior without weakening data integrity.
Decision
When a multipart upload declares a checksum algorithm in CreateMultipartUpload, SILO applies this contract:
- If the client supplies a part checksum, the server continues to validate it. A wrong value or algorithm fails and is never hidden by fallback computation.
- If the client omits the part checksum, the server computes it in one pass with the MPU algorithm over the logical plaintext stream, before compression and encryption, and persists the result.
- A normal
UploadPartechoes a checksum response header only when the client supplied the checksum. A server-computed fallback is not echoed. UploadPartCopyhas no client part-body checksum, so the server computes the value and returns it inCopyPartResult.ListPartsreturns the persisted part checksum.FULL_OBJECTcompletion continues to linearize the full checksum from stored part checksums.COMPOSITEcompletion continues to require a checksum for every part; clients can recover those values withListParts.- Computation occurs during the existing read. Completion never re-reads the entire object merely to manufacture missing state.
In one sentence:
The optional input is the client-provided checksum value, not the server’s responsibility to maintain a consistent checksum-enabled MPU.
How we found it
The defect surfaced while investigating a different multipart checksum issue, #31.
#31 concerned CompleteMultipartUpload: for FULL_OBJECT, a client can complete with part numbers, ETags, and an optional full-object checksum without retaining every part checksum in the completion XML. Tracing that path backward exposed a stronger, earlier condition in erasureObjects.PutObjectPart:
Once an MPU declared a checksum algorithm, every UploadPart had to carry the matching x-amz-checksum-* value. Omitting it returned:
API-level probes reproduced the behavior on both the single-drive and erasure backends.
Reviewing CopyObjectPartHandler raised the severity from a client-configuration incompatibility to P0. UploadPartCopy has no request body for the caller to checksum. The handler reads the source object, constructs an internal reader, and eventually enters the same PutObjectPart implementation. There is no client header and no SDK setting that can repair the request. Every checksum-enabled MPU therefore rejected UploadPartCopy by construction.
What AWS requires
This cannot be decided by saying that MinIO has historically behaved a certain way. The S3 protocol is the authority.
The AWS UploadPart API describes each algorithm-specific checksum header as something that “can be used as a data integrity check.” More importantly, its response fields say that the checksum is present only when it was provided in the request.
The AWS UploadPartCopy API is different: when the MPU was created with an algorithm, the copy result contains that part checksum. There is no copy request body, so this is necessarily a server-computed value.
The AWS ListParts API is the standard way to recover checksums for parts in an upload that is still in progress.
The algorithm/type matrix also rules out treating the repair as one Boolean flag:
| Algorithm | FULL_OBJECT |
COMPOSITE |
|---|---|---|
| CRC64NVME | Supported | Unsupported |
| CRC32 / CRC32C | Supported | Supported |
| SHA1 / SHA256 | Unsupported | Supported |
FULL_OBJECT is limited to CRCs that can be linearized, but SHA1 and SHA256 still need correct per-part digests for COMPOSITE completion.
SDK configuration makes the gap practical. Current AWS SDKs usually calculate request checksums when an operation supports them, but users can choose request_checksum_calculation = when_required, and low-level callers can initiate an algorithm without repeating it on every part. S3 accepts those requests; SILO did not.
Why removing the check is not a fix
The most tempting patch is to delete the comparison and allow a checksum-less part to proceed. That only moves the failure to completion.
SILO does not reconstruct and re-read all object bytes during MPU completion. It reads ObjectPartInfo.Checksums from each part.N.meta:
- a missing entry immediately becomes
InvalidPart; FULL_OBJECTcallsChecksum.AddPart, combining digests with their part lengths;COMPOSITEconcatenates the raw digest bytes and hashes them into the object checksum.
The actual invariant is therefore:
Deleting the upload check without filling the metadata would make UploadPart appear successful, leave ListParts incomplete, omit the UploadPartCopy response value, and fail later during completion. A delayed failure is harder to diagnose than the original immediate one.
Alternatives considered
| Option | Benefit | Fatal problem | Decision |
|---|---|---|---|
| Delete the strict check | Smallest diff | Part metadata still lacks the checksum; completion must fail | Rejected |
Relax only FULL_OBJECT |
Unblocks some default CRC clients | Leaves COMPOSITE and SHA incompatible; cannot close #46 |
Rejected |
| Re-read every part at completion | Avoids storing a digest during upload | Adds O(object size) second-pass I/O and still cannot fix ListParts or the copy response |
Rejected |
Always return the server value from normal UploadPart |
Makes federation forwarding easy | Violates the AWS response contract | Rejected |
| Copy the AIStor implementation exactly | Commercial precedent | CRC-only fallback and a transformed-stream placement risk | Rejected |
| Compute and persist in one pass over logical plaintext | Complete protocol behavior, no second I/O, CRC and SHA support | Requires an explicit plaintext checksum reader distinct from the storage reader | Accepted |
What the commercial edition taught us
We downloaded and verified the then-current MinIO AIStor RELEASE.2026-08-07T18-34-35Z. Without a commercial license the server enters offline mode and denies S3 operations, so the evidence came from Go pclntab and ARM64 disassembly, not a black-box compatibility run.
The static analysis showed that AIStor already:
- installs a server hasher when the client checksum is absent;
- persists the result in part metadata;
- exposes checksum fields in
CopyPartResult.
It nevertheless applies fallback only to CanMerge() algorithms—CRC32, CRC32C, and CRC64NVME. SHA1/SHA256 COMPOSITE still follows the old checksum missing path. More importantly, the hasher is attached in the object layer to the current r.Reader; under compression or encryption that reader may already represent transformed storage bytes.
AIStor validated the general direction—compute and store—but not an implementation that SILO could copy mechanically.
How adversarial review overturned the first design
The first plan tried to centralize every decision inside erasureObjects.PutObjectPart: read the MPU metadata in the object layer and install a server hasher when the incoming reader had no client checksum. It looked attractive because all internal callers would share one rule.
The first Fable 5 Max adversarial review found that this design was wrong for compression.
newS2CompressReader is not a lazy wrapper. Construction immediately launches a goroutine:
The S2 writer also reads several blocks concurrently. After constructing the compressor, the handler still performs option parsing, encryption preparation, and the object-layer call. By the time PutObjectPart installed a hasher, the plaintext reader could already have lost several MiB:
- a large part would get a checksum with a missing prefix;
- a small part could reach EOF before installation and produce no result;
- mutating
ServerSideHasherconcurrently withReadwould be a data race.
That finding changed the responsibility split:
The handler installs the hasher before any eager transform starts; the object layer validates the algorithm, requires a result, and persists it atomically.
This was the decisive turn in the design. Putting logic in the lowest layer may look more uniform, but stream correctness depends equally on when bytes begin moving and which representation of those bytes a layer can see.
Final implementation
A dedicated logical checksum reader
PutObjReader originally distinguished two concepts:
Reader, the stream sent to storage, possibly compressed or encrypted;rawReader, used by older ETag and checksum code.
Under compression, even rawReader may not directly see plaintext; it can merely carry an ETag through an etag.Tagger chain. The repair therefore did not overload it. It added an unexported field:
This reader always represents the logical S3 part bytes. WithEncryption can replace the storage Reader, but it must preserve checksumReader.
Unexported accessors on PutObjReader then:
- return the effective client or server checksum type;
- prefer the client value whenever it exists;
- otherwise return the server result finalized at EOF.
Keeping the mechanism unexported minimizes public Go API growth and gives #63 a shared internal path without prematurely changing ordinary CopyObject behavior.
Preparing the hasher before transformations
prepareMultipartChecksumReader loads the algorithm and checksum type saved with the MPU:
- no declared algorithm means no work;
- an existing client checksum is compared by base algorithm;
- a wrong algorithm preserves the
InvalidArgumentrejection; - an omitted client checksum installs the corresponding server hasher on the plaintext reader.
For normal UploadPart:
- the compressed path prepares
actualReaderafter request-checksum parsing but beforenewS2CompressReader; - the uncompressed path prepares the request hash reader before the encryption reader is constructed.
For UploadPartCopy:
- a checksum-enabled MPU first gets an inner hash reader over the logical source range;
- a range copy hashes only the selected bytes;
- compression and destination encryption start only after that reader is ready.
The object layer remains authoritative
Early handler preparation does not replace the storage invariant. erasureObjects.PutObjectPart still:
- re-parses the expected MPU algorithm;
- requires an effective checksum type that matches;
- obtains the checksum map after erasure encoding finishes;
- reports an internal error instead of committing if an enabled algorithm has no result;
- writes the checksum with the ETag, sizes, and index into
part.N.meta, then atomically renames the part.
An internal caller that bypasses the HTTP handler without preparing a valid checksum is therefore rejected just as before. It cannot silently commit a part that violates the MPU invariant.
CopyPart response shape
CopyObjectPartResponse gained the five algorithms supported by this source tree:
All are omitempty, so an MPU without checksums produces the old XML. Normal UploadPart still uses the existing TransferChecksumHeader and echoes only a client request value; fallback computation does not alter that response.
Why it works
After the repair, the data flow is:
This satisfies four requirements that previously appeared to conflict:
- Protocol compatibility: omitting an optional header succeeds.
- No integrity downgrade: a supplied client value is still checked end to end and is never hidden by server fallback.
- Correct object semantics: the checksum covers logical S3 bytes, not compressed data or ciphertext.
- Controlled cost: hashing shares the existing read and adds CPU, not a second disk or network pass.
EOF has a precise role. hash.Reader finalizes ServerSideChecksumResult only when it reaches EOF. Closing the compression pipe synchronizes the compressor goroutine with the storage read; the object layer reads the result only after encoding returns. Targeted -race tests verified that concurrency boundary.
The compatibility-baseline blocker
The five new CopyObjectPartResponse fields are exported Go API. SILO’s buildscripts/rebrand-guard rescans imports, environment variables, headers, routes, storage markers, and exported symbols, then compares them in both directions with buildscripts/rebrand-guard/compat-baseline.json. An unacknowledged symbol makes CI fail.
After recording the five #46 fields, the guard still reported two additions:
They did not come from #46. They belong to the earlier database-notification repair f1ba68358 on the local main branch. The cmd startup path intentionally needs the exported type for errors.As, but that earlier commit had not updated the compatibility baseline. Every later change based on that HEAD would therefore fail the CI guard.
We chose “option A”: acknowledge the two notification symbols as part of their original repair while retaining the five #46 fields. The final baseline diff is exactly seven additions and zero deletions, and the guard reports:
This does not disable the check. Exact set equality means that acknowledging a nonexistent symbol also fails. The change explicitly records two intentional compatibility-surface additions.
golangci-lint has not yet run locally; it remains a remote go.yml gate. Green local go test, go vet, race, and rebrand-guard results do not substitute for green remote CI.
Verification evidence
The new tests execute 76 subtests across:
- CRC32, CRC32C, and CRC64NVME
FULL_OBJECT; - CRC32, SHA1, and SHA256
COMPOSITE; - correct client checksums, wrong algorithms, and wrong values;
- absence of a server-computed checksum in normal
UploadPartresponses; - server values in
UploadPartCopyresponses andListParts; - a real 5 MiB + 1 KiB two-part full-object merge;
- zero-length parts and overwriting the same part number;
- a range copy whose SHA256 covers only the copied interval;
- single-drive and 16-drive erasure backends;
- default, versioned, compressed, encrypted, and compressed-plus-encrypted modes;
- explicit SSE-C and SSE-S3.
Local validation included:
All passed. Two subsequent Claude Code Fable 5 Max implementation reviews and the final acceptance review returned GO with no blocking findings.
Cost, risk, and release boundary
When a client omits its value, the server performs one additional hash over the part. CRC cost is small; SHA costs more CPU. Both share the read that already had to occur, without buffering an entire part in memory or adding a completion-time second pass.
During a rolling upgrade, old and new nodes may answer the same checksum-less request differently: a new node accepts it while an old node returns 400. ObjectPartInfo.Checksums did not change format, so stored data remains downgrade-readable, but client-visible behavior stabilizes only after all serving nodes have upgraded. The release note must call that out.
The original local implementation later merged and shipped in Server 20260903. Its historical local verification record does not establish the state of any production deployment.
Why two follow-ups remain separate
Adversarial review found two related but independent issues.
#63: CopyObject + compression
Ordinary CopyObject can also attach a server-side checksum to a transformed stream. It shares the root cause and the new checksumReader mechanism, but it is a different API with a different test matrix and rollback boundary. We chose a separate repair and require that PR to reuse this plaintext-reader contract instead of inventing a second abstraction.
#64: legacy federation
Legacy etcd federation turns UploadPartCopy into an ordinary remote UploadPart. Under the AWS response semantics preserved here, that remote request does not return a server fallback value, so the proxy may still lack the checksum required for CopyPartResult. A follow-up must independently choose between a remote-returned value and an ETag-verified ListParts fallback. It must not make all external UploadPart responses non-compliant merely to simplify an internal proxy.
Separating them does not abandon consistency. Consistency is maintained through one shared rule:
Every server-computed S3 checksum binds to the logical plaintext stream, is installed before any eager transform, and is validated and persisted by the object layer that owns the storage invariant.
Lessons retained
The repair leaves lessons more durable than its individual lines of code:
- An optional header does not make internal state optional. If the protocol lets the client omit a value, the server must produce the state its own completion path needs.
- Request acceptance and response disclosure are separate contracts. A normal UploadPart may compute internally and still omit the value; UploadPartCopy must return it.
- Stream layers are defined by byte semantics. The lowest layer is not automatically correct if it no longer sees logical bytes, and an eager goroutine turns “install later” into a race.
- A commercial implementation is evidence, not the specification. AIStor showed the direction and the boundary that could not be copied.
- A compatibility guard is a change-acknowledgment mechanism.
compat-baseline.jsonexists to assign every new compatibility surface, not merely to make CI quiet. - Independent defects should ship independently while sharing invariants. #63 and #64 remain separate, but both must cite and obey the checksum-reader contract established here.
The final result is not a broad relaxation. It is a stricter and more accurate boundary: clients may omit optional information; the server may not omit correctness.
28 - BadDigest, InvalidRequest, and the CompleteMultipartUpload Checksum Contract
Publication update, 2026-09-17: The streaming-checksum follow-up (PR #143) shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
The repairs for #48 and #50 are included in Server 20260903. This record supersedes the August proposal to leave CRC64NVME canonicalization unchanged.
Why validation belongs at completion
Initiation selects the checksum algorithm and object type; each part records
its checksum. Completion must validate the final value and any explicit type
assertion against that stored contract. A request cannot change COMPOSITE to
FULL_OBJECT just because both share the same base algorithm, or bypass the
assertion by omitting the final digest.
The repair distinguishes an omitted type from an explicit type and normalizes only internal representation flags. It uses operation-specific error types, leaving UploadPart and the global checksum-mismatch mapper unchanged.
Current error contract
All failures below return HTTP 400 and do not commit a new completed object.
| Request condition | S3 error |
|---|---|
| Wrong full-object or composite object digest | BadDigest |
| Explicit supported type differs from the initiated type, including a type-only assertion | BadDigest |
| Unknown or lowercase type token | InvalidArgument |
| Completion algorithm differs from the initiated algorithm | InvalidArgument |
| Missing required composite part checksum | InvalidRequest, naming the algorithm and part |
| CRC64NVME with COMPOSITE at initiation | InvalidArgument |
| CRC64NVME value with COMPOSITE rejected by checksum parsing at completion | InvalidArgument |
| Bare COMPOSITE type at completion of a CRC64NVME FULL_OBJECT upload | BadDigest |
| Incorrect client checksum during UploadPart | Existing XAmzContentChecksumMismatch |
For FULL_OBJECT, part checksums may be omitted; any supplied values remain
validated. Matching type-only assertions are allowed. An omitted optional type
does not assert COMPOSITE. SHA1/SHA256 with FULL_OBJECT are rejected by the
algorithm/type parser before the stored-type comparison.
CRC64NVME decision and source evidence
The original review deferred #50 pending evidence; that is historical, not the current contract. PR #93 rejects the invalid algorithm/type combination in headers and trailers, and PR #96 removes completion’s canonicalization. Both precede the 20260903 tag. Completion can fail at either parsing or stored-type comparison, which explains the two distinct error codes above.
The initial mapping is in PR #74, with
the explicit type follow-up 7e079ff05 merged through
PR #85. Current
handler tests
exercise full-object and composite mismatches, type-only assertions, invalid
tokens and both CRC64NVME rejection stages. This is committed regression
coverage, not a claim that this documentation update reran all Server tests.
Separate streaming-checksum follow-up
PR #143, following @cbornet’s
#107, handles an aws-chunked request
that advertises x-amz-trailer but supplies its checksum in a header, as used by
the AWS Java SDK v2. It also rejects an invalid header-delivered value instead
of dropping validation. This later repair is on main and not in Server
20260903. It is separate from completion error mapping.
Compatibility and remaining boundaries
Callers that inspect error codes now see checksum/type failures as BadDigest
and missing composite values as InvalidRequest. Successful checksum-free
uploads and ETag semantics are unchanged. No metadata migration is required.
Several narrower differences remain: providing a final digest when initiation
recorded no checksum is rejected as BadDigest; composite part-count and value
mismatches share one description; a -N suffix on a full-object checksum is
not itself validated as a part count. These are documented observations, not
claims of separately filed public issues or complete AWS parity.
29 - ListMultipartUploads: Implementation, Upgrade Contract and Design History
Publication update, 2026-09-17: The September source repairs discussed here shipped in Server 20260916. Coordinated upgrades, opt-in prerequisites and remaining limitations still apply. Dated source-status and validation records below retain their original scope.
This is the problem, design, and decision record for SILO issue #79.
September 16 implementation and upgrade contract
PR #198 retains mr javad seydi’s original metadata-and-scan contribution and adds maintainer fixes for disappearing markers, incomplete discovery, cancellation confirmation and upgrade diagnostics. The pre-release compatibility follow-up restores legacy as the default and makes strict mode an explicit process-environment opt-in. This section describes those source changes. It is not included in Server 20260903, and does not establish production performance or deployment acceptance. The August 30 analysis below remains a historical design record.
Listing and cancellation
New uploads persist their bucket and object key in the existing xl.meta; completion removes these upload-only fields. Strict listing discovers durable uploads across pools and sets, verifies metadata with the existing read quorum, and applies prefix, delimiter, CommonPrefixes and a global page limit of at most 1,000. Restarting a node or choosing another endpoint does not depend on rebuilding its upload cache.
Ordering is (key, initiation time from the native upload ID, encoded upload ID). A returned marker defines this boundary even after the corresponding upload is completed or canceled. This supports clients that echo the server’s markers; it does not promise arbitrary lexical comparison of random upload IDs. There is no snapshot across pages under concurrent mutations. Without key-marker, upload-id-marker is ignored. With a key marker, invalid base64 retains the existing 404 response; a decodable but unsupported native ID returns 400. Unsupported persistent IDs are legacy records, never assigned a fabricated current initiation time.
Directory discovery requires floor(N/2)+1 successful drive scans per set. For example, two available drives out of four are insufficient and return 503, even when metadata could still be read from two copies. A source-drive identity read may exclude another bucket only after validating its bucket/key hash; uncertain identities require a quorum read. Strict mode returns MultipartListingNotReady (503) for legacy records and MultipartListingMetadataInvalid (503) for invalid identities. It never silently switches the whole request to cache-based listing.
In default legacy mode, Abort retains the released read quorum, best-effort cleanup and pool-order return behavior, avoiding new 503 responses for previously successful cancellations. In strict mode it checks every relevant pool and requires floor(N/2)+1 deletion acknowledgements per set. Remnants below read quorum can still be retried. When a majority was already absent but remnants were observed, those observed copies must be cleaned successfully; successful responses from empty drives cannot mask their deletion failures. Unknown pools or insufficient confirmations still return 503.
In both modes, an Abort with the wrong key or bucket cannot evict another valid upload’s cache entry. The S3 HTTP layer retains its existing idempotent response: a nonexistent upload also returns 204; malformed input, authorization and quorum errors remain errors. 204 does not prove that every physical copy has been deleted. Offline part data may still need later cleanup.
Unresolved creation-write boundary: these confirmations do not fence a physical creation write that continues after its caller receives a storage timeout. A fault-injection test reproduces a 16-drive/EC:8 case where seven delayed writes and seven offline old copies restore a writable upload after acknowledged cancellation. The test preserves this known limitation; its passing status is not a repair claim. A durable creation fence needs a separate storage-consistency design.
Coordinated upgrade
The default is legacy, retaining the old exact-key/cache listing limitations. An ordinary upgrade does not require pausing production or draining uploads to enable the new listing. The mode is read only from the server process environment, MINIO_API_MULTIPART_LISTING=legacy|strict; no shared configuration key is added. An unset value selects legacy. An invalid value logs a diagnostic and falls back to legacy while preserving other API settings. Do not set this mode with mcli admin config set.
Only before opting into strict mode must operators upgrade all writers, stop introducing old-format uploads, finish or abort legacy uploads using known keys/IDs, run the read-only preflight below, and verify scan capacity against their workload. Once these checks pass, set MINIO_API_MULTIPART_LISTING=strict in every server’s service environment and restart; check the effective mode at each endpoint. Restart paginated traversals after changing modes. Issue #79 remains open for the default listing limitations; this batch does not claim a complete repair.
If a development build already persisted api multipart_listing, back up configuration and record API values, upgrade all configuration writers to the patched version, prevent concurrent configuration writes, then remove only that historical key:
The patched server ignores the historical key’s mode value without automatically rewriting shared configuration or deleting history. The config get/export views omit retired keys, so their absence from those views alone does not prove deletion. Require a successful targeted reset, then verify preserved values after restart and during a controlled rollback check; do not reset the whole api subsystem. A development build can reintroduce the key on its next configuration write, and history containing that key should not be replayed directly. Existing API values can temporarily be pinned through their corresponding environment variables where needed, but that does not replace persistent cleanup. This procedure covers this API configuration change only; other rollback constraints require separate checks.
The read-only, SigV4-authenticated endpoint GET /minio/admin/v3/multipart-preflight requires admin:StorageInfo. For example, with credentials and an endpoint supplied by the operator:
The report contains mode, ready, complete, scannedEntries, legacyUploads, and per-pool/set drive coverage, uncovered drive indexes and oldest legacy initiation time. It bypasses upload caches, inspects even suspended pools and detects minority legacy copies. ready=true requires every drive to be inspected, no unreadable candidate metadata and no observed legacy copies. It cannot attest that every writer has been upgraded or that no concurrent writer will introduce a legacy record. Offline drives, scan errors, timeouts and budget exhaustion prevent readiness; an incomplete count is not a zero count. Rerun after drives return and before enabling strict mode.
For uploads whose original key/ID has been lost, the existing stale-upload cleanup scans each server’s local drives. Age is measured from creation, not recent part activity. Retain the existing cleanup policy, wait and verify actual drain; the default 24-hour expiry and 6-hour interval do not guarantee drain completion. Do not shorten expiry to accelerate an ordinary upgrade, since this can also remove active long-running uploads. Continue using default legacy mode if drain cannot be established. This batch changes neither cleanup policy nor the available deletion APIs.
The shared maxUploadsList cap changes from 10,000 to 1,000 in both modes; clients that assumed one response contained everything must paginate.
A legacy upload lacks bucket/key identity and cannot safely be excluded as belonging to another bucket, so it can block strict listing for any bucket. A valid new-format identity belonging to another bucket can be filtered earlier. This does not impose deployment-wide 503 responses on default legacy listing.
Scan capacity and evidence limits
Each process admits two scans, with 16 identity workers and four full-metadata workers per scan. Directory reads pass a finite count with overflow detection. The aggregate budget is 100,000 returned directory entries, including repeated entries on different drives and hash directories; it is not a promise to list 100,000 unique uploads. Concurrent directory calls may already be in flight when the aggregate budget is exceeded. Overflow returns SlowDown (503), never a successful partial page. A 30-second context budget stops further scheduling; admission stays held until scan workers exit. This is neither a precise memory ceiling nor a guarantee that a canceled physical system call stops immediately.
As a budget illustration only, one upload per unique key on every drive of an N-drive set costs about 2 × N entries across the two directory levels: 100000 / (2 × N) uploads, or about 3,125 at N=16. Multiple uploads under one key share the hash-directory cost; other sets/pools, stale directories and timeouts also matter. This is not a universal upload-count limit.
Every page still rescans durable state: total enumeration cost grows with both stored candidates and page count. The tests cover missing markers, multi-pool coverage, partial-deletion retries, identity fallback, RPC directory bounds, cancellation admission and the known late-write counterexample. Temporary multi-node and maintained-client checks establish functional behavior for their recorded environment. Production-scale latency and foreground-load impact remain deployment-specific acceptance work; merging the source does not certify them.
A September 16 temporary Docker Desktop arm64 run used two nodes, four APFS-backed bind volumes and 11,000 uploads across two buckets, while other local validation was running. One 1,000-entry page for the 10,000-upload target bucket took 18.7 seconds; two concurrent requests returned retryable SlowDownRead responses after about 25–27 seconds. These observations do not meet the provisional five-second page target and are not an isolated SSD benchmark. Capacity and foreground-load acceptance remain open; the bounded scanner must not be advertised as a large-scale performance fix.
Historical analysis (2026-08-30): “current”, “recommended” and gate language below describes that review and proposal. The September 16 section above defines the implementation and remaining limits. Capability advertisement and the old one-day drain estimate are not current guarantees.
The problem in plain language
Imagine that four large files are still being uploaded:
An S3 client asks, “show me every unfinished upload below tables/.” AWS S3 returns the first three. SILO currently treats tables/ as if it were the complete name of one object, looks for exactly that object, and returns an empty list.
If the client removes the prefix and asks for every unfinished upload in the bucket, SILO takes a different shortcut: it reads a process-local memory cache. That cache may contain all four uploads on the node that created them, but it does not survive a restart and is not authoritative across nodes. The upload data is still on disk; the list is wrong.
This is why the defect is more serious than one ignored query parameter. Cleanup tools can receive 200 OK, conclude that no unfinished uploads exist, and report success while uploads remain on disk. The server is not losing committed objects, but it is giving callers a false view of unfinished work.
Executive decision
SILO should fix this behavior if it intends to keep advertising practical S3 compatibility.
The repair is justified because the current endpoint silently claims success, behaves differently after restart or node switching, and breaks standard prefix-based cleanup and pagination. The default 24-hour stale-upload collector limits storage accumulation on default configurations, but it does not make the API result truthful.
The repair is not a small change. Existing upload directories contain only a one-way hash of the bucket and object key, and the original key is not stored in their xl.meta. A correct implementation must begin persisting that identity for new uploads, discover candidates with erasure-aware quorum rules, apply S3 semantics globally across pools and sets, and handle legacy uploads during a rolling upgrade.
The recommended direction is therefore:
- record the bucket and object key in the upload’s existing quorum-written metadata;
- build a bounded on-demand scan as the durable correctness path;
- keep any cache only as a rebuildable optimization;
- enable strict S3 behavior only after every writer has upgraded and all keyless legacy uploads have drained;
- consider a durable secondary index only if measurements prove that scanning cannot meet a product-approved service-level objective.
Sources and provenance
The problem statement and proposed design are grounded in five kinds of evidence.
The S3 contract
The AWS ListMultipartUploads API defines the public contract for general-purpose buckets:
prefixselects every upload whose key starts with that string;delimitergroups matching keys intoCommonPrefixes;max-uploadslimits a page, with 1,000 as the documented maximum;key-markerandupload-id-markercontinue a truncated listing;upload-id-markeris ignored whenkey-markeris absent;- results are ordered by object key, then by initiation time for uploads with the same key.
AWS documentation does not settle every implementation edge unambiguously. Equal timestamps, invalid or out-of-range max-uploads, URL encoding, marker boundaries, and the way CommonPrefixes consume a page should be captured once against AWS and stored as fixtures before implementation.
The reported defect
Issue #79 supplied a self-contained signed reproducer against pgsty/silo:latest and compared SILO with AWS, RustFS, SeaweedFS, and Garage. Its four central observations reproduce:
| Request | Required behavior | Observed SILO behavior |
|---|---|---|
prefix=t/ |
return the three keys beginning with t/ |
returns no uploads |
max-uploads=1 |
return one item and continuation markers | returns every cached upload |
key-marker=t/a_b/p2 |
continue after that key | returns every cached upload |
prefix=t/&delimiter=/ |
return grouped CommonPrefixes |
returns neither uploads nor prefixes |
The issue correctly identifies a compatibility failure, but its statement that max-uploads is always ignored and IsTruncated is always false is broader than the implementation. Those claims hold on the empty-prefix cache path used by the reproducer; the exact-object path can honor max-uploads and upload-id-marker and can set IsTruncated.
The upstream design history
The behavior was inherited rather than invented by SILO:
- MinIO PR #5248 deliberately removed prefix-based listing from the erasure backend in 2017 “to simplify” multipart support.
- MinIO PR #20407 added the empty-prefix multipart cache in 2024, mainly for Alluxio tests.
- A 2025 report of the same exact-key behavior, MinIO issue #20989, was closed as working as intended.
- SILO’s current S3 compatibility reference already records the exact-object-name divergence, although it did not explain the cache, pagination, marker, delimiter, or restart limitations before this design record.
This history explains why the code looks deliberate. It does not make the endpoint compatible with the AWS contract.
Source review
The current source has two mutually exclusive listing paths:
The important locations are:
cmd/erasure-server-pool.go: empty-prefixmpCache, per-pool concatenation, and the internal exact-object lookup used byNewMultipartUpload;cmd/erasure-multipart.go: exact-object listing, upload directory construction, stale-upload cleanup, and the quorum write for a new upload’sxl.meta;cmd/erasure-sets.go: hashing a supplied object name to one erasure set;cmd/bucket-handlers.go: public request validation, including a501 NotImplementedguard whenkey-markerdoes not share the request prefix;cmd/object-api-multipart_test.go: a large expected-results table whose final assertion block checks only echoed scalar fields, not the returned uploads, prefixes, markers, or truncation state.
Independent reproduction and adversarial review
The issue scenario was independently reproduced against the reviewed SILO source with a single-node server and SigV4 requests. Additional probes established that:
- an exact object key can paginate its own uploads;
upload-id-markercurrently affects that exact-key path even withoutkey-marker, contrary to AWS;NextKeyMarkerremains empty on an exact-key truncated page;max-uploads=0behaves as unlimited in the current path;- a plain server restart empties the bucket-wide view while exact-key lookup still finds the on-disk uploads.
A second adversarial architecture review challenged the storage, quorum, migration, suspended-pool, mixed-version, and performance assumptions. The corrections from that review are incorporated below; this record does not treat an AI review as a substitute for code tests or an AWS conformance capture.
What the code actually does
Empty prefix: a volatile node-local view
With no prefix, the pool layer returns every MultipartInfo for the bucket from mpCache, sorted only by initiation time. It does not apply max-uploads, key-marker, upload-id-marker, or delimiter, and it does not compute continuation markers or IsTruncated.
The cache is initialized empty at process startup. Creation populates only the node that handles the request. Completion and abort delete cache entries, including peer notifications in some paths, but creation has no equivalent durable cluster-wide population or startup rebuild. Consequently:
- a restart can change a non-empty listing into an empty one;
- two nodes can return different answers for the same bucket;
- a successful response is not evidence that the server has enumerated durable upload state.
Non-empty prefix: an exact object lookup
With a non-empty prefix, the string is passed through object hashing as though it were a complete object name. One erasure set is selected, and listing reads the directory derived from sha256(bucket/object).
This path can enumerate multiple upload IDs for that exact object. It sorts them by initiation time, applies its upload-id-marker, stops at max-uploads, and sets IsTruncated. It still does not implement lexical prefix matching, CommonPrefixes, general key-marker semantics, or NextKeyMarker.
Multiple pools make pagination less correct
For a non-empty request on a multi-pool deployment, the pool layer invokes each active pool with the same maximum and concatenates the results. It does not perform a global ordered merge or recompute page boundaries and next markers. A request for N items can therefore collect up to N from each pool.
Suspended pools are skipped by listing and by the other public multipart verbs. In-progress uploads left on a suspended or decommissioning pool are therefore inaccessible, not merely unlisted. That is a related lifecycle defect, but listing alone must not advertise handles that PutObjectPart, ListParts, CompleteMultipartUpload, and AbortMultipartUpload cannot use. Pool drain or forced abort should be designed as a separate cross-verb change.
Why current uploads cannot be backfilled
The multipart namespace is flat:
The hash is one-way. The original bucket and key are not encoded in the path. They are also not stored as a name field in the current multipart xl.meta; the supplied object name only influences the erasure distribution during newFileInfo construction.
Therefore an all-directory scan can discover that an upload exists, but it cannot determine which bucket or key it belongs to. The current node-local cache cannot repair this reliably because it is incomplete across nodes and disappears on restart.
This rules out a tempting “small” fix: scanning every existing xl.meta and applying prefix filters. New identity metadata or a durable index is required, and old keyless uploads need an explicit migration policy.
Complexity assessment
The semantic algorithm is not the hardest part. The hard part is obtaining a complete, quorum-valid, globally ordered input set without turning a listing call into an uncontrolled cluster-wide metadata storm.
| Area | Complexity | Why |
|---|---|---|
| Pure S3 filtering and pagination | Medium | Rules are finite, but marker and delimiter edge cases need captured AWS evidence. |
| Persisting bucket/key in new upload metadata | Medium | It reuses an existing quorum write, but completion, rollback, healing, and replication compatibility must be tested. |
| Candidate discovery | High | The namespace mixes every bucket and duplicates each upload across erasure drives. One-disk discovery can miss quorum-valid uploads. |
| Quorum and concurrent deletion | High | A scan must reject minority ghosts while tolerating abort, completion, GC rename-to-trash, and transient ENOENT. |
| Multi-pool global pagination | High | Results must be merged, sorted, truncated, and marked once across all accessible pools and sets. |
| Rolling migration | High | Old writers keep creating keyless uploads; old completers may preserve unknown internal metadata. |
| Performance and resource control | High | A bucket request may require inspecting every active upload in the cluster, not only that bucket. |
Overall, this is a high-complexity compatibility project with medium wire-compatibility risk and high implementation-correctness risk. It is not a destructive object-format migration: the recommended design adds internal metadata for new incomplete uploads and leaves the existing directory scheme in place.
Compatibility and operational impact
Wire behavior changes
A correct implementation deliberately changes observable results:
prefix=foowill matchfoo,foobar, andfoo/..., not only the exact keyfoo;- bucket-wide results will be ordered by key and initiation time rather than only initiation time;
max-uploadswill actually limit a page;- the default and maximum will move from SILO’s current 10,000 constant toward the AWS limit of 1,000, subject to the captured edge-case contract;
- clients must follow
NextKeyMarkerandNextUploadIdMarkerinstead of assuming one response contains everything; - delimiter requests will return
CommonPrefixes; - the current handler-side
501for a marker outside the prefix will be replaced by the captured AWS semantics.
These are compatibility fixes, but they can break software that accidentally depends on SILO’s old non-S3 behavior. In particular, a client that ignores pagination may see fewer entries after the repair. Strict behavior should therefore be introduced through an explicit release and rollout contract, not silently slipped into an unrelated patch.
Storage-format compatibility
The recommended write path adds the bucket and object key as reserved internal metadata inside the new upload’s existing quorum-written xl.meta. It does not rename multipart directories or create a second transactional write.
Before CompleteMultipartUpload renames upload metadata into the completed object, the new upload-only fields must be removed alongside the multipart checksum fields that are already stripped there.
An old binary completing an upload created by a new binary will not know to remove the new internal keys. They would remain inert and hidden from S3 user metadata, but persist in the completed object’s internal metadata. Rolling-upgrade tests must prove that unknown reserved keys do not disturb healing, replication, metadata comparison, or downgrade reads. The product must then choose between tolerating that residue and adding a scrubber; it must not assume the keys disappear.
Operational cost
Because all buckets share one flat hash namespace, an on-demand scan is O(all active multipart uploads in the cluster), not O(uploads in the requested bucket). Bounded parallelism, cancellation, memory limits, and failure behavior are part of correctness, not optional tuning.
Default SILO configuration expires stale multipart uploads after 24 hours and runs cleanup every 6 hours. Once the last old writer has been upgraded, the keyless population should normally drain within roughly 30 hours. Operators with a larger custom expiry have a longer migration window. A zero value is mapped back to the 24-hour default in the current code; no supported “disabled” expiry value was identified in this review.
The collector bounds default storage accumulation, but does not repair a false listing response. It also does not remove the need to test sustained legitimate multipart activity, failure modes, and custom expiry settings.
Severity
The recommended classification is P1 / high compatibility, not P0:
- no committed object data loss was demonstrated;
- no security boundary is bypassed;
- unfinished uploads remain on disk until completed, aborted, or collected;
- default stale-upload cleanup bounds accumulation in the ordinary configuration.
It remains high rather than medium because the server returns fabricated success, the answer changes after restart or node switching, and cleanup or quiescence tooling can be misled into false confidence.
Options considered
Option 0: leave the current behavior unchanged
This has no engineering cost and preserves every accidental behavior. It also preserves false 200 OK responses, node-local inconsistency, restart volatility, broken prefix cleanup, and an inaccurate impression of S3 support.
This option is acceptable only if SILO deliberately downgrades the public compatibility claim and treats the endpoint as unsupported. Even then, silently returning an incomplete success is inferior to explicit rejection.
Decision: reject as a long-term position.
Option 1: explicit documented divergence
Reject combinations that SILO cannot honor with a stable NotImplemented-class error and document the exact supported subset. This is operationally honest and much smaller than full compatibility.
It is still a breaking change: tools that currently receive an empty or unbounded 200 OK may begin failing jobs. It also does not produce an S3-compatible endpoint. The error behavior and default release policy must be deliberate.
Decision: acceptable short-term containment if full compatibility is declined or deferred; not a compatibility fix.
Option 2: persist identity, scan durable state, optionally cache
For each new upload, store the bucket and key in reserved internal metadata in the upload’s existing xl.meta. For listing, discover upload directories across accessible pools and sets, validate candidates with erasure read quorum, then run one global S3 semantic layer. A cache may accelerate this path only if it can be rebuilt and reconciled from durable state.
This avoids a second write transaction and keeps the directory layout stable. Its principal cost is the cluster-wide scan.
Decision: recommended, subject to a performance and failure-mode spike.
Option 3: durable bucket-scoped ordered index
Maintain a secondary index ordered by bucket, key, and upload identity. Listing becomes scalable and naturally paginable, but create, complete, abort, healing, rollback, and reconciliation must keep two locations consistent across failures. The design resembles multipart index structures that upstream MinIO deliberately removed while simplifying this subsystem.
Decision: no-go unless measurements show that Option 2 cannot meet the product-approved service-level objective.
Rejected variant: repair only mpCache
Filtering, sorting, paginating, broadcasting creates, or rebuilding the current cache would improve symptoms but would not by itself establish a durable quorum-valid source of truth. A cache-only patch risks producing a more convincing but still incorrect answer.
Decision: reject. A cache can optimize a correct read path, never define it.
Recommended design
1. Freeze the public contract first
Create a recorded AWS fixture suite for general-purpose buckets covering:
- ordering across keys and multiple uploads of one key;
- equal initiation times and a deterministic total-order tie-break;
- prefix and exact-key overlap;
- key-marker with and without upload-id-marker;
- upload-id-marker without key-marker;
- delimiter,
CommonPrefixes, and page accounting; max-uploadsomitted, 0, 1, 1,000, and greater than 1,000;encoding-type=url;- empty pages, final pages, and next-marker values.
The captured responses should become repository fixtures. CI should not depend on live AWS access.
2. Persist recoverable identity in the existing write
At NewMultipartUpload, add reserved internal metadata for the canonical bucket and object key before the existing writeAllMetadata quorum write. The exact key names are an implementation detail, but they must be versioned, unambiguous, size-bounded by the existing object-key limits, and excluded from client-visible metadata.
At successful completion, delete those upload-only keys before copying fi.Metadata into the final object metadata and before renameData. Abort and stale cleanup already delete the entire upload directory and need no separate index operation.
3. Separate discovery from validation
Candidate discovery and candidate validity are different questions.
For every accessible, non-suspended pool and set:
- list candidate hash and upload directories from all online drives required by the configured list-quorum policy;
- union and deduplicate those names;
- read the candidate
xl.metathrough the normal erasure metadata machinery; - include the upload only when its metadata is quorum-valid and contains a valid bucket/key identity;
- tolerate a candidate disappearing during abort, completion, or stale cleanup;
- under strict list quorum, fail the request rather than return a partial
200 OKwhen a required set cannot be evaluated.
Using the first healthy disk for discovery is insufficient: that disk may have been offline when a still-quorum-valid upload was created.
4. Apply semantics once, globally
Feed the validated candidates from all pools and sets into a pure semantic layer. The layer owns bucket filtering, prefix, delimiter grouping, ordering, markers, maximum-page accounting, URL encoding, IsTruncated, and next markers.
Pool-local limits and markers must not be applied before the global merge. The result should be deterministic under duplicate discovery and independent of which node handles the request.
5. Preserve the internal exact-object operation
erasureServerPools.NewMultipartUpload currently calls ListMultipartUploads(bucket, object, ...) to keep another upload for the same object in the same pool. If the public function starts treating that argument as a lexical prefix, foo could match foobar and select the wrong pool.
Introduce a narrowly named internal helper such as FindMultipartUploadPool or ListMultipartUploadsExact. It should use the existing object hash path and must not share the public prefix semantics.
6. Treat cache as an optimization
The existing mpCache may be removed. If retained, it must satisfy all of the following:
- durable state remains authoritative;
- startup can rebuild it;
- create, complete, and abort updates are propagated consistently;
- reconciliation detects missed events and stale entries;
- a cold or divergent cache falls back to the quorum-valid scan;
- correctness tests pass with the cache disabled.
7. Gate strict behavior through rolling migration
Legacy upload records lack bucket/key identity and cannot be reconstructed reliably. Use two externally meaningful modes:
- legacy mode, the initial upgrade default: new writers persist identity; keyless uploads are counted and drained; the documented response policy for a mixed keyed/keyless population must be selected explicitly;
- strict mode: activation requires every writer node to advertise the new metadata capability and the observed keyless count to be zero. Discovering a keyless upload afterward is an error with anomaly telemetry, not a silent omission.
A short shadow comparison can help validate the new scanner, but a permanent third operating mode is unnecessary unless the spike finds a need. With default expiry, the expected legacy drain is about one day plus one cleanup interval after the last old writer stops.
There is one unresolved product choice in legacy mode:
| Policy | Advantage | Cost |
|---|---|---|
| return the complete keyed subset with documented telemetry | keeps tools operating during the bounded drain | still returns an incomplete 200 OK that ordinary clients cannot see is incomplete |
| fail listing while any keyless upload exists | never fabricates completeness | can block cleanup and existing jobs throughout the drain window |
This choice belongs in the ADR. Strict mode has no such ambiguity: it must fail loud if its precondition is violated.
8. Keep suspended-pool lifecycle separate
Listing should initially mirror the accessibility contract of the other multipart verbs and scan non-suspended pools. Adding suspended-pool entries to listing alone would expose uploads that cannot be extended, completed, or aborted.
Open a separate lifecycle design for in-progress uploads when a pool drains: either keep all multipart verbs available until the uploads finish, migrate them, or force-abort them under a documented policy. Do not hide that problem inside #79.
Performance spike and decision rule
Option 2 is preferred because it has one durable write location, but its scan cost must be measured rather than assumed.
Generate 1,000, 10,000, and 100,000 active uploads across a matrix of pools, sets, and drive counts. Measure:
- cold and warm p50/p95/p99 latency;
- total and per-drive
ListDiroperations; - metadata-read and internode RPC counts;
- peak memory and allocation volume;
- cancellation latency;
- behavior with slow, offline, healing, and intermittently disappearing drives;
- simultaneous create, complete, abort, and stale cleanup;
- first-page and deep-page cost with selective and empty prefixes.
The acceptance threshold is a product decision and must be recorded before interpreting the result. A guessed one- or two-second target is not evidence. If the scan meets the approved target with bounded resource use, reject Option 3. If it does not, use the measurements to design the smallest durable index that solves the demonstrated bottleneck.
Test and release gates
Semantic and unit tests
- pure table tests generated from recorded AWS fixtures;
- ordering, marker, delimiter, encoding, truncation, and maximum-edge coverage;
- property tests ensuring pagination returns each logical upload exactly once;
- deterministic behavior with duplicate candidates and equal timestamps.
Object and handler tests
- strengthen the existing object-layer table to assert uploads, common prefixes, markers, and truncation;
- parse and validate handler XML bodies instead of checking only status codes;
- verify default and invalid
max-uploadshandling; - test exact-helper pool selection independently of public prefix semantics.
Distributed and failure tests
- restart equivalence and node-switch equivalence;
- multiple sets and pools with a single global page boundary;
- candidate missing from one drive but present at quorum;
- minority ghost after partial abort;
- concurrent completion and GC rename-to-trash;
- unavailable set under every supported
list_quorumpolicy; - rolling upgrade, old-writer reintroduction, downgrade completion, and strict-mode gating;
- unknown internal metadata under healing and replication.
Delivery gates
- approve the ADR, including product mode and performance SLO;
- commit the captured conformance fixtures;
- complete and review the storage spike;
- implement and pass focused, full, race, and failure QA;
- update the S3 compatibility reference and operational guidance;
- commit and merge the source change;
- build and identify the release artifact or container image;
- canary a rolling upgrade and observe keyless-drain telemetry;
- enable strict mode only after its gates hold;
- verify the live endpoint before closing #79.
Passing an earlier gate is not evidence that a later gate happened.
Final recommendation: fix it, but do not rush it
Leaving the current endpoint indefinitely is the wrong trade-off. This is not an obscure response-field mismatch: it affects discovery and cleanup of unfinished data, returns successful but false answers, and changes behavior across nodes and restarts. Those properties undermine the practical meaning of S3 compatibility.
At the same time, a direct implementation patch is also the wrong trade-off. The current disk layout cannot identify legacy uploads, a correct scan needs erasure-aware discovery and quorum, and wire-correct pagination changes observable client behavior.
The balanced decision is:
- GO for the ADR, AWS fixture capture, metadata-plus-scan prototype, and performance/failure spike;
- GO conditionally for Option 2 after the product SLO and legacy response policy are approved;
- NO-GO for a cache-only repair, an immediate durable secondary index, strict-by-default behavior in a patch release, or closing the issue before rolling-upgrade reachability is demonstrated;
- if implementation capacity is unavailable, GO for an explicit documented divergence and stable error behavior rather than continuing to fabricate successful listings.
This preserves compatibility discipline without pretending that a high-risk distributed listing change is a two-line bug fix.