Skip to content

fix(store): rejoin a Store rebuilt empty at its old raft address - #3234

Merged
imbajin merged 3 commits into
apache:masterfrom
hugegraph:fix/store-rebuild-rejoin-3227
Sep 24, 2026
Merged

imbajin merged 3 commits into
apache:masterfrom
hugegraph:fix/store-rebuild-rejoin-3227

Conversation

@bitflicker64

@bitflicker64 bitflicker64 commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the PR

A Store whose data volume is lost comes back at the same raft address (on Kubernetes the StatefulSet always rebuilds the Pod under the same DNS name; on bare metal, any rebuild that reuses host:port) but registers with PD under a new store id. Retiring the old id per the documented procedure (Tombstone, then patrolPartitions) never repairs the groups: every group ends up naming a store id that no longer exists, the rebuilt Store holds no partitions, every group runs on two live replicas, and /v1/stores, the cluster state, Hubble and Pod readiness all read healthy. #3227 has the full measurement and a diagram of the loop.

There were two faults, and both had to go:

  1. No data ever reaches the rebuilt Store. reallocShards puts the new id Y into each group and fires ChangeShard. The group leader maps the ids to raft endpoints (shards2Peers), gets the same endpoint set it already has, so changePeers adds nothing. The rebuilt Store never gets a raft node for any group, and jraft only installs a snapshot into a node that exists. On the [Bug] A Store rebuilt with an empty volume never rejoins its raft groups; retirement (Tombstone + patrol) cannot repair it and every health surface reads healthy #3227 run the leaders logged Replicator stateChanged hugegraph-store-2...:8510 OFFLINE for that address 11,814 times on store-0 alone.
  2. The old id keeps coming back. The leader's local ShardGroup still lists the old id X. Its partition heartbeat (HeartbeatService.partitionHeartbeat, every 5 s) sends that list, and PartitionService.partitionHeartbeat overwrites PD's corrected group with it. On the [Bug] A Store rebuilt with an empty volume never rejoins its raft groups; retirement (Tombstone + patrol) cannot repair it and every health surface reads healthy #3227 run PD rewrote group 1 with X three seconds after reallocShards wrote Y (Raft 1 updateShardGroup ... store_id: 1190848017232829764 at 17:59:47, PD leader log). onConfigurationCommitted resolves endpoints through the same stale list, so any later leader change would do the same.

Before: the leader's ChangeShard adds nothing and its heartbeat writes the old store id back to PD, so 12 of 12 groups keep the retired id. After: the leader creates the raft node on the rebuilt Store and takes PD's id, and all 12 groups name the rebuilt Store within 17 s

Main Changes

Store side only: hg-store-core, its tests, and the Store operations guide. No PD, proto or configuration change.

  • PartitionEngine.doChangeShard: after changePeers, call the new syncShardIdentities. When PD's shard list and the raft configuration have the same endpoints but PD names another store id at one of them (never the leader itself), the leader:
    1. sends createRaftNode to that endpoint with the current configuration, so the rebuilt Store creates the raft node and jraft installs the leader's snapshot into it (the same RPC changePeers already uses for new learners), then takes a snapshot. If the RPC fails, it returns that status as changePeers does, so nothing is re-queued; the stored task runs again when the endpoint's replicator comes online;
    2. takes PD's ids into its local shard group (roles from the live configuration), persists it and reports it to PD, which stops the heartbeat overwrite;
    3. waits until the peer's replicator reaches Replicate (the existing wait loop, moved into waitForReplicate so both paths share it). A timeout logs a warning and returns TASK_ERROR; the replicator keeps installing the snapshot. The ids are taken before the wait because a rebuilt Store that restarts before PD names it exits in PartitionManager.loadPartitions.
      No raft membership change happens: the endpoint is already a voter, and the rebuilt node joins that slot with an empty log.
  • PartitionManager.getStoreByRaftEndpoint(ShardGroup, Metapb.ShardGroup, String): when the id the local group holds for an endpoint is no longer in PD's current group, use the id PD's group has at that endpoint. onConfigurationCommitted fetches PD's group once and resolves through it, so a leader change after the repair does not put the old id back. The two-argument overload is unchanged.
  • hugegraph-store/docs/operations-guide.md, Single Store Node Failure: the steps for a replacement that reuses the failed node's raft address (retire the old id, patrol, verify, delete), and that Store versions without this fix never finish them.
  • StoreIdChangeTest (in CoreSuiteTest): endpoint resolution when the local id is still valid, retired, or deleted from PD, and the endpoint-diff rule (finds a new id at an unchanged address, ignores real membership changes, never includes the leader itself).

Not in this PR: PD retiring the old record by itself when a new id registers at its raft address. The documented Tombstone step is still needed to start the repair; with this change it now works. Readiness while a rebuilt Store catches up stays with #3229.

Verifying these changes

Control (83ef9f3) This PR (9b6bf414)
Groups naming the old id after Tombstone + patrol 12 of 12, after 1501 s 0 of 12, at the first sample: 1 s (run 1), 17 s (run 2)
Groups naming the rebuilt Store 0 12
Rebuilt Store :8520/v1/partition/{id} 500 for all 12 200 for all 12, same key counts as the leader per table
Rebuilt Store data dir 38 MB (survivors 246 MB) 242 MB (survivors 243, 244 MB)
After DELETE /v1/store/{old} 12 groups name the deleted id 0 groups name it
/v1/stores and cluster state 3 Up; rebuilt Store partitionCount 0; Cluster_OK 3 Up; rebuilt Store partitionCount 12; Cluster_OK (run 2, from 48 s on)
Acknowledged writes lost 0 of 3182 0 of 429 (run 1), 0 of 1492 (run 2)

Run 2 (same images, fresh cluster) adds a 5-minute health watch after convergence, a second patrolPartitions, and a restart of the rebuilt Store with its PVC kept. It converged 17 s after the Tombstone. From the 48 s sample on, every 30 s sample read Cluster_OK, the rebuilt Store Up with partitionCount 12, and all 12 groups PState_Normal. The second patrol changed nothing. The restarted Store was Ready in 151 s under the same id, with 0 container restarts, no is illegal exit in loadPartitions, and all 12 groups served. The Store-level partitionCount follows the Store heartbeat (30 s), so it lags the repair: run 1 sampled it 1 s after the Tombstone and saw partitionCount 0 and Cluster_Warn while the snapshots were still installing.

Rerun after the review follow-ups (97f5b013, 0c2a1801), images built from 0c2a1801: converged 17 s after the Tombstone; Cluster_OK, rebuilt Store partitionCount 12 and all 12 groups PState_Normal through the 5-minute watch; restart with the PVC Ready in 151 s, same id, 12 of 12; 0 of 1494 acknowledged writes lost. All 12 leaders logged snapshot after create raft node then doChangeShard result is Status[OK]. Unit suites unchanged (30, 9, 104). Evidence: bitflicker64/hugegraph logs/fix-3227-review.

Store leader log on the fixed run, per group (store-0 and store-1 between them cover all 12):

Raft 3 store id changed at raft address [hugegraph-store-2.hugegraph-store.hg.svc:8510], local {...=X}, pd {...=Y}
Send to hugegraph-store-2.hugegraph-store.hg.svc:8510 CreateRaftNode rpc call ...
Raft 3 shard group after store id change [... store_id: 239294855241425130 ...]
Raft 3 doChangeShard result is Status[OK]

Evidence (scripts, run logs, per-Store partition views, PD and Store logs, write ledgers): bitflicker64/hugegraph logs/fix-3227 (s2b-run.sh is the harness, README.md lists the layout, one folder per run)

Does this PR potentially affect the following parts?

Documentation Status

  • Doc - TODO: required documentation is pending; complete it before merging.
  • Doc - Done: documentation is included here or linked below.
  • Doc - No Need: no user-visible documentation is affected.

Documentation files in this PR or paired hugegraph-doc PR:

  • hugegraph-store/docs/operations-guide.md (this PR). The website (hugegraph-doc) only lists these PD endpoints, with no replacement procedure, so no paired doc PR.
  • The Helm chart README is not in master yet; it lives in feat(helm): add HStore deployment chart #3218 and currently tells operators never to delete a Store volume. That warning changes there, with a version caveat, once this merges; see step 2 below.

After this merges

For whoever picks this up next. Nothing here blocks the merge.

  1. Rerun the [Bug] A Store rebuilt with an empty volume never rejoins its raft groups; retirement (Tombstone + patrol) cannot repair it and every health surface reads healthy #3227 scenario on images built from merged master. On a host with Docker, kind, helm and about 12 GB free memory (the runs above used 12 cores and 15 GB):
    • Check out apache/hugegraph master at or after the merge commit of this PR, then build the images with the revision label: SHA=$(git rev-parse HEAD); IMAGE_TAG=m3227 SOURCE_REVISION=$SHA docker buildx bake -f docker/bake.hcl pd store server-hstore --set '*.platform=linux/amd64' --set "*.labels.org.opencontainers.image.revision=$SHA" --load, and tag each as hg3227/{pd,store,server}:master. (Or use the Docker Hub :latest published after the merge, and record its digests and revision instead of building.)
    • Run OBSERVE_S=300 WAIT_S=1500 bin/s2b-run.sh master with the harness linked under Evidence. It expects the chart under ~/hg-3227/chart/helm/hugegraph, and harness/kind.yaml, writer.sh and verify-ledger.sh from the evidence folder under ~/hg-lifecycle (kind.yaml at the top, the two scripts in scripts/). It also needs the images hg3227/hubble:control and hg3227/probe:1, the latter built from harness/probe.Dockerfile.
    • Pass: converged=1; 0 groups naming the old id and 12 naming the new one; the rebuilt Store answers 200 on all 12 :8520/v1/partition/{id}; after the observation window /v1/stores shows the rebuilt Store Up with partitionCount 12 and the cluster Cluster_OK; the store-2 restart with its PVC comes back with 12 of 12 and no is illegal line; missing=0 in the ledger. Attach the s2b.log and health.log to this PR as a comment.
  2. Update the Helm chart README (hugegraph/hugegraph branch feat/hstore-helm-chart-3132, the head of apache PR feat(helm): add HStore deployment chart #3218 and its mirror feat(helm): add HStore deployment chart hugegraph/hugegraph#221, file helm/hugegraph/README.md). Two places say a Store volume must never be deleted: the paragraph starting "Do not delete a Store's PersistentVolumeClaim on current images." under Disaster Recovery, and the first bullet of "## Limitations" ("A Store cannot be recovered in place after its volume is lost"). Replace them with the retirement procedure that now converges (find the old id at the replaced Pod's address, POST /v1/store/{oldId} with {"storeState":"Tombstone"} on the PD leader, patrolPartitions, verify every group lists only Up Stores and the replaced Store's :8520/v1/partition/{id} answers 200, then DELETE /v1/store/{oldId}). Keep a version caveat: images built before this PR's merge commit, including every published 1.7.x image and :latest before that merge, still cannot recover in place, and the old warning applies to them word for word. Only make this change after step 1 passes, and cite the step 1 run in the commit message.
  3. Keep [Feature] Store-side readiness for rolling replacement: registration marks Up before partition restore, and a stopped Store stays Up in every shard group for keepAlive-timeout #3229 open. This PR does not change when a Store reports Up or add a readiness signal. It does change what [Feature] Store-side readiness for rolling replacement: registration marks Up before partition restore, and a stopped Store stays Up in every shard group for keepAlive-timeout #3229 is about: before, a Store rebuilt with an empty volume never received data, so "Up but empty" lasted forever; now it receives a snapshot for every group after the old id is retired, and the remaining gap is the window between Up and caught-up (1 s on the 243 MB test data, longer on real data), plus the stale Up of a stopped Store. Post one short comment on [Feature] Store-side readiness for rolling replacement: registration marks Up before partition restore, and a stopped Store stays Up in every shard group for keepAlive-timeout #3229 saying this, linking this PR.
  4. Add the scenario to the testing skill as a permanent regression case (repo bitflicker64/hugegraph-helm-testing): copy s2b-run.sh from the evidence into scripts/lifecycle/, add a row for it to scripts/lifecycle/README.md, and add a section to references/test-suite.md under Lifecycle with the procedure, the pass criteria from step 1, and the expected control result on images older than this merge (12 of 12 groups keep the old id, rebuilt Store 500 on every group, Cluster_OK throughout). Mark the old S2b finding in that file as fixed by this PR.
  5. Tell maintainer it landed. Once step 1 passes, one comment on this PR mentioning him, with the merge commit, the step 1 result in one line, and a link to the feat(helm): add HStore deployment chart #3218 README change from step 2.

A Store whose data volume is lost comes back at the same raft address
but registers with PD under a new store id. PD's reallocShards puts the
new id into every shard group and fires ChangeShard, but the raft
configuration already holds that address, so changePeers has nothing to
add: the rebuilt Store never gets a raft node or a snapshot. The group
leader still has the old id in its local shard group, and its partition
heartbeat writes that id back to PD a few seconds later. Every group then
names a store id that no longer exists and runs on two live replicas,
while /v1/stores and the cluster state read healthy.

When PD's shard list and the raft configuration have the same endpoints
but PD names another store id at one of them, the leader now creates the
raft node on that endpoint (jraft then installs a snapshot into it),
takes PD's store ids into its local shard group, reports the group to PD
and waits for the peer to catch up. Resolving a raft endpoint to a store
id on a configuration commit also checks PD's current group, so a later
leader change does not bring the old id back.
@codecov

codecov Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 29.52381% with 74 lines in your changes missing coverage. Please review.
✅ Project coverage is 41.37%. Comparing base (83ef9f3) to head (0c2a180).

Files with missing lines Patch % Lines
...va/org/apache/hugegraph/store/PartitionEngine.java 17.04% 73 Missing ⚠️
.../apache/hugegraph/store/meta/PartitionManager.java 94.11% 0 Missing and 1 partial ⚠️
Additional details and impacted files
@@             Coverage Diff              @@
##             master    #3234      +/-   ##
============================================
+ Coverage     41.35%   41.37%   +0.02%     
- Complexity     7299     7322      +23     
============================================
  Files           802      802              
  Lines         69688    69770      +82     
  Branches       9291     9309      +18     
============================================
+ Hits          28816    28867      +51     
- Misses        37576    37613      +37     
+ Partials       3296     3290       -6     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@bitflicker64 bitflicker64 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: no. Summary: The repair path looks correct for the #3227 case. Endpoint equality is checked before any id sync, the leader is never included, createRaftNode is a no-op on a Store that already runs the group, and onLeaderStart and onStartFollowing both resolve through PD's group before the leader's next partition heartbeat, so the old id does not come back after a leader change. One minor retry issue: a failed createRaftNode now returns TASK_CONTINUE, and each retry takes a fresh raft snapshot. Evidence: read the full diff at 9b6bf41 with PartitionEngine.changePeers, doChangeShard, the SYNC_PARTITION_TASK handler, HgCmdClient.tryWithThrowable, HgCmdProcessor.handleCreateRaft, HgStoreEngine.createPartitionEngine and HeartbeatService.partitionHeartbeat. CI is green apart from codecov patch/project (patch coverage 30%) and build-server (hbase), which was still running when this was reviewed.

if (!status.isOk()) {
log.info("Raft {} createRaftNode, peer:{}, reason:{}", getGroupId(), peer,
status.getErrorMsg());
return HgRaftError.TASK_CONTINUE.toStatus();

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor: When createRaftNode fails here, the method returns TASK_CONTINUE, and doChangeShard re-queues the task straight away through addRaftTask(SYNC_PARTITION_TASK). Before that, local ids are not updated yet, so the retry finds the same changed endpoint and calls doSnapshot again on line 474. A single failure takes about 1.5 s: HgCmdClient.tryWithThrowable retries 5 times with a 100/200/300/400/500 ms backoff. Each retry also appends a raft entry, and setSnapshotLogIndexMargin is commented out in PartitionEngine.init, so jraft does not skip these snapshots. If the rebuilt Store goes down again after PD has moved the group to its new id (for example a crash-looping Pod), the leader of every affected group takes a snapshot and appends a log entry about every 1.5 to 2 s until the Store is back. changePeers handles the same createRaftNode failure at line 329 by returning the RPC status (code -1), which is not TASK_CONTINUE and does not re-queue. Please do the same here, or at least call doSnapshot only once per repair (for example after createRaftNode succeeds, or only when getReplicatorState shows the peer is not already installing a snapshot), so a peer that stays down does not trigger a retry loop that snapshots on every pass.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 97f5b01. A failed createRaftNode now returns the RPC status, as changePeers does at line 329, so doChangeShard no longer re-queues it. The task is already stored by the SYNC_PARTITION_TASK handler, and ReplicatorStateListener runs it again when the endpoint comes back online. The snapshot is taken once, after the raft node exists. Rerun on images from 0c2a180: converged 17 s after the Tombstone, all 12 leaders logged snapshot after create raft node, 0 of 1494 writes lost (evidence).

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Timeout retries lose pending repair state, and the required operator documentation is not included.

Get a fresh assessment by requesting another Copilot review.

Review effort: Balanced
Findings: 1 High severity · 1 Low severity

Open (2)
What changed in this PR

Fixes Store rejoining after data loss when its Raft address is reused under a new Store ID.

Changes:

  • Detects Store identity changes at unchanged Raft endpoints.
  • Creates replacement Raft nodes, synchronizes snapshots, and resolves IDs using PD state.
  • Adds Store-ID resolution and endpoint-difference tests.
File Description
StoreIdChangeTest.java Tests Store-ID and endpoint resolution.
CoreSuiteTest.java Registers the new test suite.
PartitionEngine.java Implements identity synchronization and replication waiting.
PartitionManager.java Resolves current Store IDs against PD metadata.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +513 to +514
return waitForReplicate(changed) ? HgRaftError.OK.toStatus() :
HgRaftError.TASK_CONTINUE.toStatus();

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

97f5b01: a timeout now logs a warning and returns TASK_ERROR, so it no longer re-queues a pass that finds nothing to do and reports OK. The ids stay taken before the wait on purpose: a rebuilt Store that restarts before PD names it exits in PartitionManager.loadPartitions (System.exit(0) at line 302), so moving them after the wait would crash-loop it during catch-up. Nothing is lost on a timeout: the node exists and its replicator keeps installing the snapshot on its own; the task result only goes to the log. The timeout path needs a live raft group, so it has no unit test; the kind rerun in the PR description covers the normal path.

* @param shards shard list from PD
* @return OK when nothing changed or the peers caught up, TASK_CONTINUE to retry
*/
private Status syncShardIdentities(List<Metapb.Shard> shards) {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

0c2a180 adds the procedure to hugegraph-store/docs/operations-guide.md (Single Store Node Failure, step 3): retire the old id with Tombstone, run patrolPartitions, verify, then delete it, with the note that Store versions without this fix never finish it. The website only lists these PD endpoints (quickstart hugegraph-pd), with no replacement procedure to correct. The Helm chart README is not in master yet; it lives in #3218, and its "do not delete a Store volume" warning changes there with a version caveat once this is merged and rerun on master images.

When the rebuilt Store cannot be reached, syncShardIdentities returned
TASK_CONTINUE. doChangeShard then re-queued the task as a raft entry at
once, and every pass took a new raft snapshot before the RPC failed
again, about every 1.5 to 2 seconds for as long as the Store stayed down.

Return the createRaftNode status instead, as changePeers does for the
same failure. The task stays stored, and the replicator listener runs it
again once the endpoint comes online. Take the snapshot only after the
raft node exists.

A peer that has not caught up after the wait now returns TASK_ERROR
instead of TASK_CONTINUE. The re-queued pass found the ids already
taken and returned OK without doing anything; the replicator keeps
installing the snapshot either way. The ids stay taken before the wait,
because a rebuilt Store that restarts before PD names it exits in
loadPartitions.
The single-node failure procedure said PD assigns partitions to a new
Store by itself. A replacement that reuses the failed node's raft address
with an empty data directory registers under a new store id, and the old
id keeps its partitions until it is retired. Add the Tombstone, patrol,
verify and delete steps, and say that Store versions without the apache#3227
fix never finish them.
@bitflicker64
bitflicker64 requested a balanced review from Copilot September 24, 2026 13:29

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@imbajin imbajin left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking: no actionable issue found after five independent review lanes and synthesis at 0c2a180. Score: 8.8/10. The earlier retry, timeout-result and recovery-documentation findings are addressed. Exact-head Store CI passed 30 core tests and 9 raft-core tests; live recovery and the >600-second timeout were not independently rerun. This score is not a merge approval; remaining CI and unresolved review threads still need to be cleared.

@imbajin
imbajin merged commit dbb6663 into apache:master Sep 24, 2026
21 of 22 checks passed
@imbajin
imbajin deleted the fix/store-rebuild-rejoin-3227 branch September 24, 2026 14:41
bitflicker64 added a commit to hugegraph/hugegraph that referenced this pull request Sep 24, 2026
Reserve the chart-owned checksum/ podAnnotations prefix on every
component: user annotations render after the chart's own and the last
duplicate key wins, so a fixed value would pin the rotation checksum.
Refuse server.securityContext.readOnlyRootFilesystem=true, mirroring the
Hubble guard, because the Server wrapper rewrites rest-server.properties
inside the image. Reject updateStrategy.rollingUpdate options together
with type OnDelete in the schema for PD and Store; Kubernetes refuses
that combination at apply time, which would otherwise surface
mid-upgrade. Unit tests cover the template guards and the CI
reject-invalid-values step covers the schema constraint.

Use the get-with-default pattern for the $exposed NetworkPolicy check,
matching the rest of validateValues. Remove the constant $pdMeta and
$wrapper indirection from the Server Deployment, and remove the unused
server.restServer.minFreeMemory / batchMaxWriteThreads knobs end to end
(template, values, schema, README); the chart is unreleased, so nothing
depends on them. helm template output on the default, single and
cluster presets is byte-identical before and after.

Refresh the docs: the server.testResources default in the README matches
values.yaml again, and the runbook follows apache#3232, apache#3233
and apache#3234 (merged 2026-09-24). The empty-PVC Store retirement was
re-proven on images built from master at dbb6663: all 12 groups
converged onto the replacement 1 s after Tombstone and patrol, with 0
acknowledged writes lost, so the Disaster Recovery section documents the
working procedure with a version caveat instead of a prohibition.
bitflicker64 added a commit to hugegraph/hugegraph-doc that referenced this pull request Sep 24, 2026
Bring the deployment page level with the chart at 8603cdbb3: the
values-cluster preset now ships NetworkPolicy and 5Gi/8Gi Store memory,
PD startup and liveness derive from the replica count (pd.livenessPath,
single-PD /v1/ready per apache/hugegraph#3222), the Store roll procedure
no longer treats Up in PD as the between-pods check, and the Limitations
follow the merged fixes, including in-place empty-PVC Store recovery on
images carrying apache/hugegraph#3234.

Add an operations page (EN and CN, registered in the version route map)
rewritten for operators from the chart README: ports and health,
scheduling, partition sharding, NetworkPolicy with the per-CNI NodePort
client behavior, safe Store rolls, the disaster recovery runbook with
the post-#3234 procedure, scaling including the Store drain steps,
running Hubble outside the cluster, and the two causes of "Could not
rebind".

Verified: scripts/hugo.sh build passes and both new pages render with
their cross-links in EN and CN.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

3 participants