Stop waits for all background goroutines started via goAsync to finish AND drains the rate-limiter cleanup goroutines spawned at construction time (BUG-851). Safe to call multiple times. Should be called before Store.Close() so in-flight DB writes don't race a closed connection (or worse, the SQLite
()
| 424 | // Store.Close() so in-flight DB writes don't race a closed connection |
| 425 | // (or worse, the SQLite -wal/-shm file removal in t.TempDir cleanup). |
| 426 | func (s *Server) Stop() { |
| 427 | // Signal long-running background loops (orphan GC, etc.) to exit. |
| 428 | // Each loop registers itself on s.bg, so the Wait() below blocks |
| 429 | // until they actually finish and any in-flight goroutines drain. |
| 430 | s.stopOrphanGC() |
| 431 | // Yjs op-log prune sweeper (TASK-1309). Same lifecycle pattern; |
| 432 | // signals BEFORE Wait() so the goroutine sees the close and exits. |
| 433 | s.stopOpLogGC() |
| 434 | // Short-lived-credential reaper (PLAN-1933 DR-5 / TASK-1936). Same |
| 435 | // lifecycle pattern; signal BEFORE Wait() so the goroutine exits. |
| 436 | s.stopTokenReaper() |
| 437 | // Soft-deleted-workspace hard-purge sweeper (TASK-1966). Same |
| 438 | // lifecycle pattern; signal BEFORE Wait() so the goroutine exits. |
| 439 | s.stopWorkspacePurgeSweeper() |
| 440 | // SPEC-3 event outbox drain (TASK-2714). Same lifecycle pattern; an |
| 441 | // in-flight delivery is tracked on s.bg and awaited below. |
| 442 | s.stopOutboxDrain() |
| 443 | // MCP audit writer / sweeper run on s.bg too. Signal first so |
| 444 | // the workers see the close BEFORE Wait() blocks; without the |
| 445 | // signal Wait would hang forever on the writer's blocking |
| 446 | // queue receive. |
| 447 | s.stopMCPAuditWriter() |
| 448 | // MCP session tracker (TASK-1120) runs its sweeper on s.bg too. |
| 449 | // Order with the audit writer doesn't matter — both are |
| 450 | // independent goroutines; we just need the close BEFORE Wait(). |
| 451 | s.stopMCPSessionTracker() |
| 452 | // Close the collab room manager BEFORE bg.Wait() so any in-flight |
| 453 | // op-log GC sweep (TASK-1309) blocked on a per-item lock behind |
| 454 | // an active Join can drain. collab.Close() tears down the Joins |
| 455 | // (their WS readLoops return, runConn unwinds, itemLocks |
| 456 | // release), which unblocks the GC's per-item PruneItemOpLogIfDormantBefore |
| 457 | // call. Without this ordering, Stop() can deadlock: GC waits on |
| 458 | // itemLock; Join holds itemLock until WS closes; WS only closes |
| 459 | // when collab.Close() runs; collab.Close() only runs after |
| 460 | // bg.Wait(); bg.Wait() never returns because GC is stuck. |
| 461 | // Per Codex review of TASK-1309 [P2]. nil-safe: collab is optional. |
| 462 | if s.collab != nil { |
| 463 | s.collab.Close() |
| 464 | } |
| 465 | s.bg.Wait() |
| 466 | // Watch/nudge bus (BUG-2651). Closed AFTER bg.Wait() so a background |
| 467 | // producer cannot publish into a bus that is already tearing down; both |
| 468 | // implementations are safe if one does anyway (MemoryBus finds no |
| 469 | // subscribers, RedisBus finds a cancelled context and fails closed). |
| 470 | // |
| 471 | // This matters more than it did for MemoryBus, whose Close only dropped |
| 472 | // channels: RedisBus holds a receive goroutine and a Redis subscription |
| 473 | // from construction, so skipping it leaks both for the process's life |
| 474 | // (Codex round 1 P2). Closing also closes every subscriber channel, which |
| 475 | // is how a still-open SSE stream learns to unwind. |
| 476 | if s.watchEvents != nil { |
| 477 | s.watchEvents.Close() |
| 478 | } |
| 479 | s.rateLimiters.Stop() // nil-safe via the RateLimiters receiver guard |
| 480 | } |
| 481 | |
| 482 | func New(s *store.Store) *Server { |
| 483 | rl := NewRateLimiters() |