Ask AI
Ask AI Start conversation ↵

I'm an AI assistant with Mewbo's codebase and documentation in context.

Ask me anything about Mewbo.

EXAMPLE QUESTIONS

Python modules and classes

This page is generated from inline docstrings via mkdocstrings. The sections below are grouped by package or client.

packages/mewbo_core (core runtime)

mewbo_core.loop.orchestrator

Session orchestration entrypoint.

Orchestrator

Unified tool-use orchestration loop.

Source code in packages/mewbo_core/src/mewbo_core/loop/orchestrator.py
 120
 121
 122
 123
 124
 125
 126
 127
 128
 129
 130
 131
 132
 133
 134
 135
 136
 137
 138
 139
 140
 141
 142
 143
 144
 145
 146
 147
 148
 149
 150
 151
 152
 153
 154
 155
 156
 157
 158
 159
 160
 161
 162
 163
 164
 165
 166
 167
 168
 169
 170
 171
 172
 173
 174
 175
 176
 177
 178
 179
 180
 181
 182
 183
 184
 185
 186
 187
 188
 189
 190
 191
 192
 193
 194
 195
 196
 197
 198
 199
 200
 201
 202
 203
 204
 205
 206
 207
 208
 209
 210
 211
 212
 213
 214
 215
 216
 217
 218
 219
 220
 221
 222
 223
 224
 225
 226
 227
 228
 229
 230
 231
 232
 233
 234
 235
 236
 237
 238
 239
 240
 241
 242
 243
 244
 245
 246
 247
 248
 249
 250
 251
 252
 253
 254
 255
 256
 257
 258
 259
 260
 261
 262
 263
 264
 265
 266
 267
 268
 269
 270
 271
 272
 273
 274
 275
 276
 277
 278
 279
 280
 281
 282
 283
 284
 285
 286
 287
 288
 289
 290
 291
 292
 293
 294
 295
 296
 297
 298
 299
 300
 301
 302
 303
 304
 305
 306
 307
 308
 309
 310
 311
 312
 313
 314
 315
 316
 317
 318
 319
 320
 321
 322
 323
 324
 325
 326
 327
 328
 329
 330
 331
 332
 333
 334
 335
 336
 337
 338
 339
 340
 341
 342
 343
 344
 345
 346
 347
 348
 349
 350
 351
 352
 353
 354
 355
 356
 357
 358
 359
 360
 361
 362
 363
 364
 365
 366
 367
 368
 369
 370
 371
 372
 373
 374
 375
 376
 377
 378
 379
 380
 381
 382
 383
 384
 385
 386
 387
 388
 389
 390
 391
 392
 393
 394
 395
 396
 397
 398
 399
 400
 401
 402
 403
 404
 405
 406
 407
 408
 409
 410
 411
 412
 413
 414
 415
 416
 417
 418
 419
 420
 421
 422
 423
 424
 425
 426
 427
 428
 429
 430
 431
 432
 433
 434
 435
 436
 437
 438
 439
 440
 441
 442
 443
 444
 445
 446
 447
 448
 449
 450
 451
 452
 453
 454
 455
 456
 457
 458
 459
 460
 461
 462
 463
 464
 465
 466
 467
 468
 469
 470
 471
 472
 473
 474
 475
 476
 477
 478
 479
 480
 481
 482
 483
 484
 485
 486
 487
 488
 489
 490
 491
 492
 493
 494
 495
 496
 497
 498
 499
 500
 501
 502
 503
 504
 505
 506
 507
 508
 509
 510
 511
 512
 513
 514
 515
 516
 517
 518
 519
 520
 521
 522
 523
 524
 525
 526
 527
 528
 529
 530
 531
 532
 533
 534
 535
 536
 537
 538
 539
 540
 541
 542
 543
 544
 545
 546
 547
 548
 549
 550
 551
 552
 553
 554
 555
 556
 557
 558
 559
 560
 561
 562
 563
 564
 565
 566
 567
 568
 569
 570
 571
 572
 573
 574
 575
 576
 577
 578
 579
 580
 581
 582
 583
 584
 585
 586
 587
 588
 589
 590
 591
 592
 593
 594
 595
 596
 597
 598
 599
 600
 601
 602
 603
 604
 605
 606
 607
 608
 609
 610
 611
 612
 613
 614
 615
 616
 617
 618
 619
 620
 621
 622
 623
 624
 625
 626
 627
 628
 629
 630
 631
 632
 633
 634
 635
 636
 637
 638
 639
 640
 641
 642
 643
 644
 645
 646
 647
 648
 649
 650
 651
 652
 653
 654
 655
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
class Orchestrator:
    """Unified tool-use orchestration loop."""

    #: The model the loop was last generating with. ``None`` until a loop has
    #: run, which is the honest answer for a failure raised before one did.
    #: Kept because ``_model_name`` is frozen at construction and a run that
    #: escalates down the ladder leaves it naming a model that served nothing —
    #: the failure record is what an operator benches a model from. Declared on
    #: the CLASS so the emission rule stays drivable from a bare instance, which
    #: is how its tests reach it without paying for a whole orchestrator.
    _served_model: str | None = None

    def __init__(
        self,
        *,
        model_name: str | None = None,
        fallback_models: tuple[str, ...] | None = None,
        session_store: SessionStoreBase | None = None,
        tool_registry: ToolRegistry | None = None,
        permission_policy: PermissionPolicy | None = None,
        approval_callback: Callable[[ActionStep], bool] | None = None,
        hook_manager: HookManager | None = None,
        cwd: str | None = None,
        session_step_budget: int = 0,
        system_instructions_store: SystemInstructionsStoreBase | None = None,
        session_mcp_servers: dict[str, dict] | None = None,
    ) -> None:
        """Initialize orchestration dependencies.

        *session_mcp_servers* are MCP servers a caller attaches to THIS run, in
        the standard ``{name: {command|url, …}}`` config shape. They are merged
        into the registry build below alongside the plugin-contributed ones, and
        their discovered tool ids are admitted through this run's
        ``allowed_tools`` (:meth:`_session_mcp_tool_ids`).

        **It is deliberately NOT named ``extra_mcp_servers``, even though that is
        the parameter it feeds.** The two carry different grants, and reusing one
        name for both senses is the trap the house rules name outright: a
        plugin-contributed server's tools are subject to ``allowed_tools`` like
        any other registry tool, while a server named HERE is admitted by the act
        of naming it — the caller attaching a server for one run IS the grant, and
        it has no other way to express one, since a server's tool ids are not
        knowable until discovery has run. Widening the existing parameter's
        meaning instead would silently admit every plugin server's tools into
        every scoped session in the deployment.

        Discovery is a network/subprocess cost and it belongs HERE rather than at
        whichever caller assembled the config: this constructor already runs on
        the background run thread, which is the offline side of the
        acceptance/execution boundary. A caller resolving the ids itself would be
        paying for discovery on its own request path.
        """
        self._cwd = cwd
        self._session_mcp_servers = dict(session_mcp_servers or {})
        self._session_step_budget = session_step_budget
        # Resolved LAZILY on first use (``_instructions_store``): the factory
        # raises when the configured driver is mongodb and Mongo is unreachable,
        # and a custom-instructions store being down must never stop a session
        # from starting. ``_instructions_store_tried`` makes the failure sticky
        # so we don't retry a dead connection once per run.
        self._instructions_store: SystemInstructionsStoreBase | None = system_instructions_store
        self._instructions_store_tried = system_instructions_store is not None
        # Strong references for fire-and-forget background tasks scheduled on
        # an already-running loop (the emscripten branch of
        # ``_maybe_generate_title``) so asyncio can't GC them mid-run — the
        # done-callback discards its own entry once finished. Mirrors
        # ``AgentHandle.asyncio_task`` (hypervisor.py) / ``_lifecycle_tasks``
        # (spawn_agent.py).
        self._background_tasks: set[asyncio.Task] = set()
        self._model_name = (
            model_name
            or get_config_value("llm", "action_plan_model")
            or get_config_value("llm", "default_model", default="gpt-5.2")
        )
        self._fallback_models = (
            fallback_models
            if fallback_models is not None
            else tuple(effective_fallback_models())
        )
        self._session_store = session_store or create_session_store()
        self._permission_policy = permission_policy or load_permission_policy()
        self._approval_callback = approval_callback or approval_callback_from_config()
        self._hook_manager = hook_manager or default_hook_manager()

        self._project_instructions = discover_project_instructions(cwd)
        # Loaded exactly once, here, from operator-owned config — never from a
        # tool a model can call. ``None`` when the deployment-wide switch is off
        # (the default): every call site below tests ``is not None`` and does
        # nothing else, so an unconfigured install never stats ``.mewbo/``,
        # never parses a document and never builds a rule.
        self._safety_plane = SafetyPlane.load(
            cwd, enabled=bool(get_config_value("safety", "enabled", default=False))
        )
        self._skill_registry = SkillRegistry()
        self._skill_registry.load(cwd)

        # Plugin loading (before load_registry so MCP servers are collected first).
        # Uses the shared load_all_plugin_components() so the same logic is
        # reused by the API /skills and /tools endpoints (DRY).
        self._agent_registry = None
        self._session_tool_registry = SessionToolRegistry()
        # schedule_trigger rides the ordinary SessionToolRegistry so
        # a spawned sub-agent whose allowlist admits it can bind it — the old
        # root-only extra_session_tools seam structurally could not. The factory
        # exists only once the app pushed its store+policy; None (CLI/tests) ⇒
        # the tool is simply absent, exactly as before.
        from mewbo_core.triggers.session_tool import schedule_trigger_factory

        _trigger_factory = schedule_trigger_factory()
        if _trigger_factory is not None:
            self._session_tool_registry.register(_trigger_factory)
        plugins_cfg = get_config().plugins

        if plugins_cfg.enabled:
            # Reconcile missing plugins on fresh containers / volume wipes.
            if plugins_cfg.enabled_plugins:
                self._reconcile_missing_plugins(plugins_cfg)

            from pathlib import Path as _Path

            from mewbo_core.agents.agent_registry import AgentRegistry, parse_agent_file
            from mewbo_core.hooks import merge_plugin_hooks
            from mewbo_core.tooling.plugins import load_all_plugin_components

            fan_out = load_all_plugin_components()
            self._agent_registry = AgentRegistry()

            self._skill_registry.load_plugin_components(fan_out)
            for pc in fan_out.components:
                if pc.manifest is None:
                    continue
                plugin_source = f"plugin:{pc.manifest.name}"
                plugin_caps = pc.manifest.requires_capabilities
                plugin_root = pc.manifest.install_path
                for af in pc.agent_files:
                    agent_def = parse_agent_file(_Path(af), source=plugin_source)
                    if agent_def is None:
                        continue
                    self._agent_registry.register(
                        agent_def,
                        capabilities=plugin_caps,
                        plugin_root=plugin_root,
                    )
                # Load this plugin's session tools WITH its capability gate, so a
                # capability-gated session tool (e.g. the ``scg`` suite's
                # ``scg_*``) surfaces to any session that holds the capability —
                # including via the runtime grant — not only when the
                # client lists the tool in ``allowed_tools``. Same
                # data-driven gate the AgentDefs register through above.
                for entry in pc.session_tool_entries:
                    self._session_tool_registry.load_entry(
                        entry, requires_capabilities=plugin_caps
                    )
            for hooks_json, plugin_root in fan_out.hooks_configs:
                merge_plugin_hooks(self._hook_manager, hooks_json, plugin_root)
            plugin_mcp_servers = fan_out.mcp_servers
        else:
            plugin_mcp_servers = {}

        # Reuse a cached registry across runs whose build inputs (cwd +
        # plugin-contributed MCP servers + MCP-config fingerprint) are identical,
        # instead of rebuilding per query — the per-run rebuild was on the
        # critical path to the first FE-visible event. An explicit
        # ``tool_registry`` (tests, structured runners) still bypasses the cache.
        # The cache key already covers ``(cwd, extra_mcp_servers, mcp-config
        # fingerprint)``, so a run attaching its own servers gets its own entry
        # instead of poisoning the shared one — which is what makes merging a
        # per-run set in here safe at all.
        self._tool_registry = tool_registry or get_or_build_registry(
            cwd=cwd,
            extra_mcp_servers={**plugin_mcp_servers, **self._session_mcp_servers} or None,
        )
        self._context_builder = ContextBuilder(self._session_store)

        # Register lossless micro-compaction as a pre_compact hook.
        from mewbo_core.session.compaction import micro_compact_events

        self._hook_manager.pre_compact.append(micro_compact_events)

    def run(
        self,
        user_query: str,
        *,
        max_iters: int = 3,
        initial_plan: Plan | None = None,
        return_state: bool = False,
        session_id: str | None = None,
        mode: str | None = None,
        should_cancel: Callable[[], bool] | None = None,
        allowed_tools: list[str] | None = None,
        denied_tools: list[str] | None = None,
        strict_tool_scope: bool = False,
        capability_mode: str = "all",
        skill_instructions: str | None = None,
        message_queue: queue.Queue[str] | None = None,
        interrupt_step: threading.Event | None = None,
        user_id: str | None = None,
        source_platform: str | None = None,
        invocation_id: str | None = None,
        extra_session_tools: list[SessionTool] | None = None,
        enable_skills: bool = True,
        project_autoselect: bool = False,
        attachments: list[dict] | None = None,
        user_turn_persisted: bool = False,
    ) -> TaskQueue | tuple[TaskQueue, OrchestrationState]:
        """Run orchestration for a session (sync wrapper around :meth:`arun`).

        Backward-compatible synchronous entry point for CLI / API request
        threads / tests. Owns the event-loop lifecycle via ``asyncio.run``, so
        every existing sync caller behaves exactly as before. Environments that
        already drive an event loop (browser-hosted Pyodide's WebLoop) must call
        :meth:`arun` instead — it awaits the same body with NO nested
        ``asyncio.run`` (the "WebLoop wall").
        """
        return asyncio.run(
            self.arun(
                user_query,
                max_iters=max_iters,
                initial_plan=initial_plan,
                return_state=return_state,
                session_id=session_id,
                mode=mode,
                should_cancel=should_cancel,
                allowed_tools=allowed_tools,
                denied_tools=denied_tools,
                strict_tool_scope=strict_tool_scope,
                capability_mode=capability_mode,
                skill_instructions=skill_instructions,
                message_queue=message_queue,
                interrupt_step=interrupt_step,
                user_id=user_id,
                source_platform=source_platform,
                invocation_id=invocation_id,
                extra_session_tools=extra_session_tools,
                enable_skills=enable_skills,
                project_autoselect=project_autoselect,
                attachments=attachments,
                user_turn_persisted=user_turn_persisted,
            )
        )

    async def arun(
        self,
        user_query: str,
        *,
        max_iters: int = 3,
        initial_plan: Plan | None = None,
        return_state: bool = False,
        session_id: str | None = None,
        mode: str | None = None,
        should_cancel: Callable[[], bool] | None = None,
        allowed_tools: list[str] | None = None,
        denied_tools: list[str] | None = None,
        strict_tool_scope: bool = False,
        capability_mode: str = "all",
        skill_instructions: str | None = None,
        message_queue: queue.Queue[str] | None = None,
        interrupt_step: threading.Event | None = None,
        user_id: str | None = None,
        source_platform: str | None = None,
        invocation_id: str | None = None,
        extra_session_tools: list[SessionTool] | None = None,
        enable_skills: bool = True,
        project_autoselect: bool = False,
        attachments: list[dict] | None = None,
        user_turn_persisted: bool = False,
    ) -> TaskQueue | tuple[TaskQueue, OrchestrationState]:
        """Run orchestration asynchronously (async-first entry point).

        Awaits the orchestration coroutine directly, so an environment that
        already owns a running event loop (browser-hosted Pyodide's WebLoop,
        async test harnesses) can drive a query with native ``await`` and NO
        nested ``asyncio.run``. The sync :meth:`run` wrapper is the CPython
        entry point and delegates here; the two share this single body.
        """
        if session_id is None:
            session_id = self._session_store.create_session()

        # Fold the session's durable signals (tags + merged context + the
        # entry-point surface) into filterable Langfuse tags/metadata once; the
        # seam propagates them to every child observation. Apps write their
        # context/tags before invoking the runtime, so they are present here.
        #
        # The merged context carries only the CLIENT-ADVERTISED capabilities, so
        # overlay the AUGMENTED set (advertised ∪ runtime-provider grants) before
        # deriving — otherwise a capability granted at runtime (the ``scg``
        # provider) would be live in the run yet invisible in the trace's
        # ``capabilities`` facet. ``derive`` stays a pure transform of
        # (tags, context, surface); we only enrich the context it reads.
        derive_context = dict(self._session_store.latest_context(session_id))
        effective_caps = self._session_capabilities(session_id, allowed_tools=allowed_tools)
        if effective_caps:
            derive_context["client_capabilities"] = list(effective_caps)
        # An explicit caller-supplied user_id always wins. Otherwise, the real
        # principal — stamped into the context payload by the API's
        # ``_stamp_principal_subject`` whenever auth is enabled — is the
        # honest identity for tracing. Leave it None (never the session id)
        # when no principal exists, so the seam's own anonymous fallback
        # applies instead of silently aliasing user_id to session_id.
        if user_id is None:
            principal_subject = derive_context.get("principal_subject")
            if isinstance(principal_subject, str) and principal_subject:
                user_id = principal_subject
        provenance = TraceProvenance.derive(
            tags=self._session_store.tags_for_session(session_id),
            context=derive_context,
            surface=source_platform,
        )

        with session_log_context(session_id):
            with langfuse_session_context(
                session_id,
                user_id=user_id,
                invocation_id=invocation_id,
                source_platform=source_platform,
                # The trace name identifies the KIND of turn, never this
                # execution of it: a name carrying a turn index, a session id
                # or the query text mints a fresh name per run and every
                # saved filter, evaluator and dashboard stops matching. The
                # surface is a bounded set, so it stays groupable.
                trace_name=f"turn:{source_platform or 'unknown'}",
                tags=list(provenance.tags),
                metadata=provenance.metadata,
            ):
                return await self._run_with_session_context_async(
                    user_query,
                    max_iters=max_iters,
                    initial_plan=initial_plan,
                    return_state=return_state,
                    session_id=session_id,
                    mode=mode,
                    should_cancel=should_cancel,
                    allowed_tools=allowed_tools,
                    denied_tools=denied_tools,
                    strict_tool_scope=strict_tool_scope,
                    capability_mode=capability_mode,
                    skill_instructions=skill_instructions,
                    message_queue=message_queue,
                    interrupt_step=interrupt_step,
                    extra_session_tools=extra_session_tools,
                    enable_skills=enable_skills,
                    project_autoselect=project_autoselect,
                    attachments=attachments,
                    user_turn_persisted=user_turn_persisted,
                    provenance=provenance,
                )

    async def _run_with_session_context_async(
        self,
        user_query: str,
        *,
        max_iters: int,
        initial_plan: Plan | None,
        return_state: bool,
        session_id: str,
        mode: str | None,
        should_cancel: Callable[[], bool] | None,
        allowed_tools: list[str] | None = None,
        denied_tools: list[str] | None = None,
        strict_tool_scope: bool = False,
        capability_mode: str = "all",
        skill_instructions: str | None = None,
        message_queue: queue.Queue[str] | None = None,
        interrupt_step: threading.Event | None = None,
        extra_session_tools: list[SessionTool] | None = None,
        enable_skills: bool = True,
        project_autoselect: bool = False,
        attachments: list[dict] | None = None,
        user_turn_persisted: bool = False,
        provenance: TraceProvenance | None = None,
    ) -> TaskQueue | tuple[TaskQueue, OrchestrationState]:
        """Run orchestration with Langfuse session context set.

        Async body shared by the sync :meth:`run` wrapper and the async-first
        :meth:`arun`. Every LLM-driven leg is awaited directly (never wrapped
        in ``asyncio.run``) so it composes on a caller-owned running loop.
        """
        state = OrchestrationState(goal=user_query, session_id=session_id)
        resolved_mode = self._resolve_mode(mode)
        state.summary = self._session_store.load_summary(session_id)
        state.tool_results = state.tool_results or []
        state.open_questions = state.open_questions or []
        task_queue: TaskQueue | None = None

        self._hook_manager.run_on_session_start(session_id)
        error_msg: str | None = None
        try:
            # Record the turn — unless the seam that ACCEPTED it already did.
            # ``SessionRuntime.start_async`` writes the ``user`` event the moment
            # it accepts, because this body only reaches here after
            # ``Orchestrator.__init__``'s heavy synchronous setup and the turn
            # would be invisible to every client for that whole window. Suppress,
            # never deduplicate by content: two identical consecutive queries are
            # legitimate, so comparing text would silently drop a real turn.
            # Default ``False`` keeps every direct caller (CLI turn engine,
            # structured runners) writing it here exactly as before.
            if not user_turn_persisted:
                self._session_store.append_user_turn(session_id, user_query, attachments)
            if self._should_update_summary(user_query):
                state.summary = self._update_summary_with_memory(
                    session_id,
                    user_query.strip(),
                )

            updated_summary = await self._maybe_auto_compact_async(session_id)
            if updated_summary:
                state.summary = updated_summary

            # Server-registry slash commands (``/compact``, ``/skills``,
            # ``/tokens``, ``/fork``, ``/tag``, ``/help``) all dispatch
            # through ``mewbo_core.session.commands.execute_command`` so the
            # operation is single-sourced regardless of which UI typed
            # them — CLI ``run_sync`` route, API channel pipeline, or
            # console palette. Per-command rendering still belongs to
            # the calling UI; the orchestrator only short-circuits the
            # tool-use loop and surfaces the rendered ``result.body``.
            stripped_query = user_query.strip()
            parts = stripped_query.split(maxsplit=1)
            if parts and parts[0].startswith("/"):
                cmd_name = parts[0][1:]
                from mewbo_core.session.commands import (
                    COMMANDS,
                    CommandContext,
                    execute_command,
                )

                if cmd_name in COMMANDS:
                    raw_remainder = parts[1] if len(parts) > 1 else ""
                    # ``/compact`` takes the whole remainder as a single
                    # focus directive; other handlers expect token args.
                    cmd_args = (
                        [raw_remainder]
                        if cmd_name == "compact" and raw_remainder
                        else raw_remainder.split()
                    )
                    cmd_ctx = CommandContext(
                        session_id=session_id,
                        session_store=self._session_store,
                        hook_manager=self._hook_manager,
                        model_name=self._model_name,
                    )
                    try:
                        result = await execute_command(cmd_name, cmd_args, cmd_ctx)
                        if cmd_name == "compact":
                            state.summary = self._session_store.load_summary(session_id) or ""
                            state.done_reason = "compacted"
                        else:
                            state.done_reason = f"command:{cmd_name}"
                        state.done = True
                        task_queue = self._build_direct_response(result.body)
                    except Exception as exc:
                        logging.warning("User-initiated /%s failed", cmd_name, exc_info=True)
                        state.done = True
                        state.done_reason = (
                            "compact_failed"
                            if cmd_name == "compact"
                            else f"command_failed:{cmd_name}"
                        )
                        message = (
                            f"Compaction failed: {exc}. Session continues uncompacted."
                            if cmd_name == "compact"
                            else f"/{cmd_name} failed: {exc}. Session continues."
                        )
                        task_queue = self._build_direct_response(message)
                    return (task_queue, state) if return_state else task_queue

            context = self._context_builder.build(
                session_id=session_id,
                user_query=user_query,
                model_name=self._model_name,
            )
            # Always pass the FULL tool spec set to the loop; plan-mode
            # filtering (read-only + configured edit tool + exit_plan_mode)
            # happens inside ``ToolUseLoop._bind_model`` so tools can be
            # re-bound after plan approval without reconstructing specs.
            tool_specs = self._tool_registry.list_specs()
            # Three-state: ``None`` skips the gate entirely (unrestricted); ``[]``
            # is an explicit empty grant and MUST reach ``filter_specs``, which
            # tests it with ``is None`` for the same reason.
            if allowed_tools is not None:
                allowed_tools = self._admit_attached_mcp_tools(allowed_tools)
                if strict_tool_scope:
                    # Strict mode: ``allowed_tools`` is authoritative —
                    # nothing outside it survives, not even built-ins.
                    # Used by wiki-qa so the agent doesn't waste round
                    # trips trying ``aider_shell_tool`` / ``spawn_agent``
                    # / etc. Caller is responsible for including any core
                    # tool it actually needs in ``allowed_tools``.
                    tool_specs = filter_specs(
                        tool_specs,
                        allowed=allowed_tools,
                        denied=denied_tools,
                        capability_mode=capability_mode,
                    )
                else:
                    # Permissive mode (FE default): ``allowed_tools`` only
                    # scopes MCP tools; built-in tools always stay.
                    builtin_ids = [s.tool_id for s in tool_specs if s.kind != "mcp"]
                    tool_specs = filter_specs(
                        tool_specs,
                        allowed=allowed_tools + builtin_ids,
                        denied=denied_tools,
                        capability_mode=capability_mode,
                    )
            elif capability_mode != "all" or denied_tools:
                # No allowlist, but a ROOT capability ceiling applies (a role
                # narrowed a session to ``read_only``) or the caller named an
                # explicit deny. Apply the coarse privilege gate + deny to the
                # registry specs so a read-only session binds only read-tier
                # tools + ``always_load`` — mirroring the per-tier filter a
                # spawned child gets. Guarded so an unrestricted, undenied
                # session skips ``filter_specs`` entirely, and so
                # ``agent.default_denied_tools`` is never applied to a session
                # that named neither a ceiling nor a deny.
                tool_specs = filter_specs(
                    tool_specs, denied=denied_tools, capability_mode=capability_mode
                )

            # Resolve session capabilities once so every downstream lookup
            # (slash-command skill activation, sub-agent catalog, activate_skill
            # tool dispatch) sees the same client-advertised set. Threading
            # allowed_tools derives any product-tool selection into a
            # request-scoped grant too — see _session_capabilities.
            session_caps = self._session_capabilities(session_id, allowed_tools=allowed_tools)

            # Skill invocation detection and hot-reload.
            self._skill_registry.maybe_reload()
            if skill_instructions is None:
                _si, _ts = self._try_skill_invocation(user_query, tool_specs, session_caps)
                if _si is not None:
                    skill_instructions = _si
                if _ts is not None:
                    tool_specs = _ts

            # Unified path: always enter the tool-use loop. Plan mode is
            # enforced inside the loop via tool filtering + path-scoped
            # permission checks + the ``exit_plan_mode`` approval gate.
            if resolved_mode == "plan":
                # Ensure the session's scoped plan directory exists before
                # the model starts so the edit tool can write to plan.md.
                ensure_plan_dir(session_id)
                state.plan_path = plan_file_for(session_id)

            max_depth = int(get_config_value("agent", "max_depth", default=5))
            max_concurrent = int(get_config_value("agent", "max_concurrent", default=20))
            attestation = None
            if bool(get_config_value("agent", "attestation_enabled", default=True)):
                # Seed from the store's last persisted
                # record so a recovered/continued session's chain links
                # continuously instead of restarting at genesis.
                from mewbo_core.agents.attestation import AttestationChain

                seed = self._session_store.last_attestation_hash(session_id)
                attestation = AttestationChain(session_id=session_id, head=seed)
            registry = AgentHypervisor(
                max_concurrent=max_concurrent,
                session_step_budget=self._session_step_budget,
                attestation=attestation,
            )
            root_ctx = AgentContext.root(
                model_name=self._model_name,
                max_depth=max_depth,
                fallback_models=self._fallback_models,
                should_cancel=should_cancel,
                event_logger=lambda event: self._session_store.append_event(session_id, event),
                registry=registry,
                message_queue=message_queue,
                interrupt_step=interrupt_step,
                # Seed the filesystem-containment axis from config;
                # narrowed per sub-agent thereafter. Defaults to "full_access"
                # because the ROOT agent acts directly for the operator —
                # containment attenuates privilege across a spawn, and it is the
                # child (default "workspace_write") that the tier confines.
                workspace_mode=str(
                    get_config_value("agent", "default_workspace_mode", default="full_access")
                ),
                # Seed the delegation-privilege axis from the caller's resolved
                # ceiling. "all" (the default) is a no-op — every child then
                # narrows from here. A role-narrowed session ("read_only" for a
                # viewer) caps the ROOT's own session tools too, since the loop's
                # ``build_for`` reads ``agent_context.capability_mode``.
                capability_mode=capability_mode,
            )
            loop = ToolUseLoop(
                agent_context=root_ctx,
                tool_registry=self._tool_registry,
                permission_policy=self._permission_policy,
                approval_callback=self._approval_callback,
                hook_manager=self._hook_manager,
                safety_plane=self._safety_plane,
                project_instructions=self._project_instructions,
                user_instructions=self._resolve_user_instructions(
                    session_id=session_id,
                    session_caps=session_caps,
                    tool_specs=tool_specs,
                    # The same inputs the loop below is handed, so the tools
                    # named in the operator's template match the ones it builds
                    # — INCLUDING strict_tool_scope, or a permissive FE root's
                    # {{ tools }} catalog drops schedule_trigger the agent holds.
                    allowed_tools=allowed_tools,
                    denied_tools=denied_tools,
                    extra_session_tools=extra_session_tools,
                    provenance=provenance,
                    strict_tool_scope=strict_tool_scope,
                    capability_mode=capability_mode,
                ),
                skill_instructions=skill_instructions,
                skill_registry=self._skill_registry,
                agent_registry=self._agent_registry,
                session_tool_registry=self._session_tool_registry,
                # Forward the caller's allowlist so plugin session tools
                # (e.g. wiki_clone_repo for the wiki indexer) get built for
                # root agents that need them. ``None`` still means "no
                # plugin session tools" for plain user sessions.
                allowed_tools=allowed_tools,
                # Deny wins over everything else the loop's session-tool gates
                # admit — the unconditional auto-surface and the capability
                # auto-surface included. Same list this run's registry specs
                # were just filtered by, so a denied tool id is off BOTH
                # surfaces, not just one of them.
                denied_tools=denied_tools,
                # Whether ``allowed_tools`` is authoritative (strict) or a
                # permissive MCP ceiling — mirrors the ``filter_specs`` branch
                # above so the loop's spawn_agent gate reads the same intent.
                strict_tool_scope=strict_tool_scope,
                cwd=self._cwd,
                session_id=session_id,
                session_capabilities=session_caps,
                extra_session_tools=extra_session_tools,
                enable_skills=enable_skills,
                project_autoselect=project_autoselect,
                # Assembled only for a run that opted in — see
                # ``_build_project_catalog``.
                project_catalog=(
                    self._build_project_catalog() if project_autoselect else None
                ),
                # The loop holds no transcript access, so a project switch
                # cannot read back what it must carry forward without this.
                session_context_reader=lambda: self._last_context(session_id),
            )
            try:
                task_queue, state = await loop.run(
                    user_query,
                    tool_specs=tool_specs,
                    context=context,
                    plan=initial_plan,
                    mode=resolved_mode,
                )
            finally:
                # In the ``finally`` so a run that RAISED still reports the model
                # that was live when it died — that is the case the failure
                # record exists for.
                self._served_model = loop.active_model
                # Snapshot the ownership index BEFORE cleanup force-settles it.
                # Once cleanup has run every handle reports terminal, and the
                # evidence that this run declared itself finished while work it
                # owned was still live is gone with it — so the read has to
                # happen here, not after.
                unsettled_children = await self._unsettled_children(registry)
                # Belt-and-suspenders: ensure all agents cleaned up.
                try:
                    await registry.cleanup(timeout=5.0)
                except Exception:
                    pass
            state.session_id = session_id
            if resolved_mode == "plan":
                state.plan_path = plan_file_for(session_id)

            # Emit assistant response event. Every user turn must have
            # exactly one materialised assistant event in the transcript
            # — if ``task_result`` is empty (e.g. ``max_steps_reached``
            # with no synthesis), write a synthetic closure marker so the
            # UI timeline finalises the turn and the LLM's
            # ``recent_events`` gains narrative closure.
            if task_queue.task_result:
                self._session_store.append_event(
                    session_id,
                    {"type": "assistant", "payload": {"text": task_queue.task_result}},
                )
            else:
                closure = _format_assistant_closure(state.done_reason, task_queue.last_error)
                self._session_store.append_event(
                    session_id,
                    {"type": "assistant", "payload": {"text": closure}},
                )

            self._maybe_generate_title(session_id)

            if not state.done:  # pragma: no cover - defensive guard
                state.done = True
                state.done_reason = "max_iterations_reached"

            # Promise-as-completion gate. A clean terminal is a CLAIM that the
            # work is finished, and the claim is false while runs this session
            # owns are still live. Children are collected at loop teardown, so a
            # handle still non-terminal here was abandoned mid-flight: the run
            # ended, the work did not. Downgrading the reason is what stops that
            # from presenting as success — ``completed`` is the one status that
            # withholds the recovery affordance, and a session with abandoned
            # work is exactly one that needs it.
            if state.done_reason == "completed" and unsettled_children:
                state.done_reason = "unmet_goal"
                if not task_queue.last_error:
                    task_queue.last_error = (
                        f"Run ended while {len(unsettled_children)} owned sub-agent "
                        "run(s) were still live; their work was abandoned rather "
                        "than collected."
                    )

            # Annotated as the TypedDict rather than ``dict[str, object]`` so
            # the checker validates every key against ``CompletionPayload``;
            # a bare literal widens to ``dict[str, str | dict[...]]`` the moment
            # a nested value is added and stops matching the event contract.
            completion_payload: CompletionPayload = {
                "done": state.done,
                "done_reason": state.done_reason,
                "task_result": task_queue.task_result,
            }
            # The loop leaves ``done_reason`` at "completed" for a run that died
            # against a credential, a network path or a quota, so this code is
            # the ONLY thing that carries that fact out of the run. Omitted
            # when absent, so a run that hit no wall carries no key.
            if state.blocked_code:
                completion_payload["blocked_code"] = state.blocked_code
            self._attach_failure_record(completion_payload, task_queue, state)
            self._session_store.append_event(
                session_id,
                {"type": "completion", "payload": completion_payload},
            )

            note = self._promise_note(
                task_queue.task_result or "",
                owned_runs_live=bool(unsettled_children),
            )
            if note is not None:
                self._session_store.append_event(
                    session_id,
                    {"type": "run_note", "payload": {"text": note}},
                )

            updated_summary = await self._maybe_auto_compact_async(session_id)
            if updated_summary:
                state.summary = updated_summary

            return (task_queue, state) if return_state else task_queue
        except Exception as exc:
            logging.exception("Orchestration failed for session {}", session_id)
            if task_queue is None:
                task_queue = TaskQueue(_human_message=user_query, action_steps=[])
            # Classify + clamp ONCE, and let that ONE bounded projection feed
            # every consumer. A raw ``str(exc)`` here is how an upstream HTML
            # error page (thousands of characters, provider-controlled) reached
            # the transcript, the completion payload, the next run's
            # ``recent_events`` bullet list — and, via ``error_msg`` below, the
            # ``on_session_end`` hook, whose string a channel adapter posts
            # verbatim into a forge PR comment or a chat message. That last one
            # is OUTWARD-facing, so it must never carry provider markup.
            run_error = RunError.from_exception(exc, model=self._error_model)
            error_msg = run_error.brief()
            task_queue.last_error = error_msg
            state.done = True
            state.done_reason = "error"
            # Closure marker so the failed turn is always materialised in
            # the UI timeline and the LLM's ``recent_events`` carries
            # narrative closure into the next recovery run. The one-line
            # ``title`` keeps that marker readable — the full diagnostic
            # lives on the completion event's ``error_detail``.
            self._session_store.append_event(
                session_id,
                {
                    "type": "assistant",
                    "payload": {
                        "text": _format_assistant_closure(state.done_reason, run_error.title)
                    },
                },
            )
            failure_payload: CompletionPayload = {
                "done": True,
                "done_reason": state.done_reason,
                "task_result": task_queue.task_result,
                "error": error_msg,
                "last_error": error_msg,
                "error_detail": run_error.model_dump(mode="json"),
            }
            # A run that RAISED can still have hit a wall first — a repo it
            # could not reach, a quota it spent — and that wall is the more
            # actionable half of the story. Carried on this path too, or a
            # blocked run that then died of something else reads as a generic
            # failure with nothing for the user to fix.
            if state.blocked_code:
                failure_payload["blocked_code"] = state.blocked_code
            self._session_store.append_event(
                session_id,
                {"type": "completion", "payload": failure_payload},
            )
            return (task_queue, state) if return_state else task_queue
        finally:
            self._record_outcome_assertions(session_id, error_msg)

    def _record_outcome_assertions(self, session_id: str, error_msg: str | None) -> None:
        """Run the session-end hooks and persist anything they assert.

        A session-end hook is the only component that can see whether the
        session's PURPOSE succeeded — it holds the owning job, while the loop
        holds only its own signals, every one of which can say success while the
        job never reached its terminal state. The assertion lands as its own
        transcript event AFTER the completion, which is where the status
        derivation can consume it without this seam having to reorder the hook
        dispatch ahead of the terminal it describes.

        TOTAL BY CONSTRUCTION, because this runs inside ``run``'s ``finally``: a
        raise here would REPLACE whatever exception the run was already
        propagating, reporting a genuine orchestration failure as a hook bug.
        That also covers a ``HookManager`` predating the return channel — it
        reports ``None``, which is simply no assertion, never an error.
        """
        try:
            assertions = self._hook_manager.run_on_session_end(session_id, error_msg)
        except Exception:  # pragma: no cover - defensive; a finally must not raise
            logging.warning(
                "session-end hooks failed for session {}", session_id, exc_info=True
            )
            return
        if not isinstance(assertions, list):
            return
        for assertion in assertions:
            try:
                self._session_store.append_event(
                    session_id,
                    {
                        "type": "outcome_assertion",
                        "payload": assertion.model_dump(mode="json"),
                    },
                )
            except Exception:  # pragma: no cover - defensive; see above
                logging.warning(
                    "could not persist an outcome assertion for session {}",
                    session_id,
                    exc_info=True,
                )

    def _attach_failure_record(
        self,
        payload: CompletionPayload,
        task_queue: TaskQueue,
        state: OrchestrationState,
    ) -> None:
        """Attach the bounded failure record to a terminal completion payload.

        ``task_queue.last_error`` is a STICKY diagnostic — a mid-run tool
        failure the run continued past still leaves it set. It is CLAMPED
        whenever set, whatever the outcome: its readers do not check
        ``done_reason`` (the scg map-job persists it onto the job record, the
        CLI prints it on /retry|/continue|/edit), so gating the clamp let a run
        that recovered and finished clean carry a raw multi-KB provider page out
        to them.

        ``error`` is WITHHELD on a run whose terminal status is ``completed``,
        because that ONE key is what a client renders as a user-facing error
        card. A tool call that failed and was recovered from is not a session
        failure, and emitting it as one puts an error card under a complete,
        correct answer — which is exactly what it did. This is the same rule the
        no-sticky-string branch below already applied; it is stated once here
        and applied to both.

        The RECORD is still carried on such a run: ``last_error`` and
        ``error_detail`` both ride the payload, and no client renders a card on
        either. That is what keeps the auditability guarantee intact — the runs
        a blanket withhold would silence are overwhelmingly the LAUNDERED ones
        (a halt presenting as success), and dropping the one field able to
        contradict the status is what turns a wrong status into an unfalsifiable
        one. A status is only worth trusting if the record can be used to check
        it, so the record stays and only the render trigger goes.

        The gate reads ``OrchestrationState.terminal_status`` rather than
        comparing ``done_reason`` here: that projection also folds in
        ``verified is False`` and the cancelled/unachieved vocabularies, and a
        second copy of it at this call site could only ever drift from it. Note
        ``done_reason == "error"`` never reaches here — the path that mints it
        raises, and its handler builds its own failure payload.

        ``blocked_code`` is consulted INDEPENDENTLY of that projection, exactly
        as the status layer and the console already consult it. A run that died
        against a credential, a network path or a quota deliberately keeps
        ``done_reason == "completed"`` and carries the wall only in that field,
        so ``terminal_status`` calls it a success — and gating on the projection
        alone withheld the error from precisely the runs a user most needs to
        see, silently, since a client that reads neither field then renders a
        clean success.

        A run that stopped short leaving NO sticky string (a doom-loop halt, a
        spent budget, a failed ground-truth check) still gets a structured
        record here — but only the additive ``error_detail``, never ``error``,
        for the same reason: a halt that produced a wrap-up answer is not an
        error to put in front of a user.
        """
        if task_queue.last_error:
            run_error = RunError.from_message(task_queue.last_error, model=self._error_model)
            brief = run_error.brief()
            task_queue.last_error = brief
            payload["last_error"] = brief
            payload["error_detail"] = run_error.model_dump(mode="json")
            if state.terminal_status() != "completed" or state.blocked_code is not None:
                payload["error"] = brief
        elif state.done_reason in UNACHIEVED_DONE_REASONS:
            payload["error_detail"] = RunError.from_message(
                f"Run ended without reaching its goal ({state.done_reason}).",
                model=self._error_model,
            ).model_dump(mode="json")

    @property
    def _error_model(self) -> str:
        """The model a failure record should name — served if known, else configured.

        ``RunError._build`` falls back to whatever model it is handed whenever
        the message text carries no provider token, so this ONE reader is what
        decides whether a failure blames the model that ran or the one that was
        configured four minutes earlier.
        """
        return self._served_model or self._model_name

    @staticmethod
    async def _unsettled_children(registry: AgentHypervisor) -> list[str]:
        """Ids of agents still non-terminal as the run tears down.

        The ownership index the promise-as-completion gate reads. Must be called
        BEFORE ``registry.cleanup``, which force-settles every handle and erases
        the distinction between "this run finished its work" and "this run
        merely stopped".

        Read-only and total: a registry that cannot answer reports nothing owed
        rather than failing a run that has otherwise succeeded.
        """
        try:
            agents = await registry.list_all()
        except Exception:  # pragma: no cover - defensive; never fail a run here
            return []
        return [h.agent_id for h in agents if h.status in ACTIVE_STATUSES]

    @staticmethod
    def _promise_note(text: str, *, owned_runs_live: bool) -> str | None:
        """Return ONE factual reminder when terminal prose promises future work.

        The prose half of promise-as-completion: a terminal that reads "I'll
        check back on it shortly" and then ends zero seconds later is not
        masking an error, it is fabricating a future. Nothing downstream can
        falsify that from the record, because there is no error to find.

        What this returns is a statement of fact the record already proves — no
        background work is scheduled — and never a verdict on the output. It is
        deliberately silent when owned runs ARE live, because then the promise
        is simply true, and silent on non-matching prose, because a heuristic
        that guesses at intent would invent the very unfalsifiable claim this
        seam exists to remove.
        """
        if owned_runs_live or not text:
            return None
        probe = text.lower()
        if not any(marker in probe for marker in _FUTURE_COMMITMENT_MARKERS):
            return None
        return (
            "This turn ended with no scheduled or running background work. "
            "Any follow-up stated in the response above will not happen on "
            "its own."
        )

    # ------------------------------------------------------------------
    # Session helpers (kept from original)
    # ------------------------------------------------------------------

    def _admit_attached_mcp_tools(self, allowed_tools: list[str]) -> list[str]:
        """Union this run's attached-MCP tool ids into *allowed_tools*.

        A server this run explicitly attached is admitted by the ACT of attaching
        it. Without this the servers are discovered and then filtered straight
        back out — and ``filter_specs`` drops an unrecognised id in SILENCE, so
        the caller would see no error, no tools, and nothing to grep for.

        A caller cannot name the ids itself: they are not knowable until
        discovery has run, which happens in this object's constructor. That is
        the whole reason the grant is expressed as "attach the server" rather
        than as an allowlist entry.

        Returns *allowed_tools* unchanged (same list object) when nothing was
        attached, which is nearly every run.
        """
        attached = self._session_mcp_tool_ids()
        return list(allowed_tools) + attached if attached else allowed_tools

    def _session_mcp_tool_ids(self) -> list[str]:
        """Registry tool ids contributed by this run's own attached MCP servers.

        Empty when nothing was attached, so a run that attaches nothing leaves
        ``allowed_tools`` untouched and never walks the spec list.

        A server that discovered NO tools contributes nothing and is reported —
        it means the server was unreachable or exposes nothing, and the only
        other symptom would be its tools quietly never being callable.
        """
        if not self._session_mcp_servers:
            return []
        ids: list[str] = []
        seen_servers: set[str] = set()
        for spec in self._tool_registry.list_specs():
            server = str(spec.metadata.get("server", ""))
            if server in self._session_mcp_servers:
                ids.append(spec.tool_id)
                seen_servers.add(server)
        silent = sorted(set(self._session_mcp_servers) - seen_servers)
        if silent:
            logging.warning(
                "attached MCP servers contributed no tools: {} "
                "(unreachable, or they expose none)",
                silent,
            )
        return ids

    def _session_capabilities(
        self, session_id: str, *, allowed_tools: list[str] | None = None
    ) -> tuple[str, ...]:
        """Return the capability tuple in effect for *session_id*.

        Reads ``client_capabilities`` from the most recently appended
        ``context`` event (set by the API from the ``X-Mewbo-Capabilities``
        header). Unions in capabilities DERIVED from *allowed_tools*
        — REQUEST-SCOPED, never persisted: when the caller names
        a product ``SessionTool``'s id (e.g. ``wiki_search_pages``) in this
        request's allowlist, ``SessionToolRegistry.capabilities_for`` looks up
        the tool's plugin-manifest ``requires_capabilities`` and unions it in
        for THIS call only. This closes the asymmetry: selecting a
        product tool via the SAME ``context.mcp_tools`` field that already
        gates ``SessionToolRegistry.build_for``'s allowlist now *also* unlocks
        its AgentDef family on the capability-only catalog gate
        (``filter_by_capabilities``). Because the derivation is fresh per
        request rather than written to a context event, omitting the tool on
        a later turn naturally revokes it — unlike the sticky, additive-only
        header path, which this leaves untouched (the two unions
        harmlessly on top of each other). Finally unions in any RUNTIME
        grants registered by a capability library above core
        (``augment_session_capabilities``) — so a capability gated on a live
        predicate (e.g. ``scg`` once the SCG is enabled AND a source is
        mapped) surfaces to an ORDINARY session without the client
        advertising it. Returns an empty tuple on any error or when nothing
        applies.
        """
        from mewbo_core.capabilities import (
            augment_session_capabilities,
            parse_capabilities,
        )

        try:
            events = self._session_store.load_transcript(session_id)
        except Exception:
            events = []
        advertised: object = None
        for event in events:
            if event.get("type") != "context":
                continue
            payload = event.get("payload")
            if isinstance(payload, dict) and "client_capabilities" in payload:
                advertised = payload["client_capabilities"]
        derived = self._session_tool_registry.capabilities_for(allowed_tools or [])
        base = set(parse_capabilities(advertised)) | set(derived)
        return augment_session_capabilities(tuple(sorted(base)))

    def _system_instructions_store(self) -> SystemInstructionsStoreBase | None:
        """Return the custom-instructions store, resolving the factory once.

        Resolution is lazy + sticky: the factory raises when the configured
        driver is ``mongodb`` and Mongo is unreachable, and this feature must
        never be the reason a session refuses to start. A failure is logged once
        and remembered, so a dead connection isn't re-probed every run.
        """
        if self._instructions_store_tried:
            return self._instructions_store
        self._instructions_store_tried = True
        try:
            self._instructions_store = create_system_instructions_store()
        except Exception:
            logging.warning(
                "Custom system instructions unavailable (store unreachable); "
                "running without them.",
                exc_info=True,
            )
            self._instructions_store = None
        return self._instructions_store

    def _resolve_instruction_tools(
        self,
        *,
        tool_specs: list[ToolSpec],
        allowed_tools: list[str] | None,
        session_caps: tuple[str, ...],
        extra_session_tools: list[SessionTool] | None,
        strict_tool_scope: bool,
        capability_mode: str = "all",
        denied_tools: list[str] | None = None,
    ) -> tuple[str, ...]:
        """The tool ids an operator's template sees in ``InstructionContext.tools``.

        The ``ToolRegistry`` specs bound for this run, UNIONED with the session
        tools the root agent will actually be built with. The union is
        load-bearing: the registry specs alone omit every session tool
        (``wiki_*``, ``scg_*``, ``submit_widget``, ``schedule_trigger``), so
        ``'wiki_search' in tools`` silently renders False for an agent that
        genuinely holds it and a template branching on a product tool can never
        fire.

        The session-tool half is resolved through ``SessionToolRegistry.ids_for``,
        the SAME selection ``ToolUseLoop`` builds from, so this list cannot drift
        from what the agent gets. Caller-injected tools (``extra_session_tools``,
        e.g. the structured-response emit tool) are unioned in too — they are
        bound for this run just as truly.

        DELIBERATELY OMITTED: the five internals the loop injects for ITSELF
        (``spawn_agent``/``spawn_agents``, ``update_todos``, ``exit_plan_mode``,
        ``activate_skill``). They are decided INSIDE ``ToolUseLoop``, downstream
        of this call, and several are mode-dependent (``update_todos`` is
        act-mode, ``exit_plan_mode`` is plan-mode), so naming them here would
        trade one lie for another. The field's ``description`` states the
        omission outright — honesty over completeness.

        ``tool_search`` is NOT one of them and IS included: despite being loop
        machinery, it is a genuine ``ToolRegistry`` spec (``always_load``), so it
        arrives through *tool_specs* like any other tool. Listing it as omitted
        would itself have been a lie — the exact failure this method exists to
        fix, one level down.
        """
        ids = {spec.tool_id for spec in tool_specs}
        ids.update(
            self._session_tool_registry.ids_for(
                allowed_tools,
                session_capabilities=session_caps,
                # Same intent the loop's build_for gets, or the operator's
                # {{ tools }} catalog drifts from what the agent holds:
                # a permissive FE root genuinely holds schedule_trigger, and a
                # role-narrowed root drops the write-tier session tools its
                # ``capability_mode`` withholds. ``denied_tools`` for the same
                # reason — the drift law this method exists to enforce.
                denied_tools=denied_tools,
                strict_tool_scope=strict_tool_scope,
                capability_mode=capability_mode,
            )
        )
        ids.update(tool.tool_id for tool in extra_session_tools or [])
        return tuple(sorted(ids))

    def _resolve_user_instructions(
        self,
        *,
        session_id: str,
        session_caps: tuple[str, ...],
        tool_specs: list[ToolSpec],
        allowed_tools: list[str] | None,
        extra_session_tools: list[SessionTool] | None,
        provenance: TraceProvenance | None,
        strict_tool_scope: bool,
        capability_mode: str = "all",
        denied_tools: list[str] | None = None,
    ) -> str | None:
        """Render the operator's custom system instructions for this run.

        The ONE resolution point: read the stored template, render it once
        against this run's :class:`InstructionContext`, and hand the plain string
        to the loop (which then survives compaction and the escalation
        re-render because it lives on the loop instance).

        Fully fail-soft by construction — a missing/disabled doc, an unreachable
        store, or a broken template all yield ``None``, and the run proceeds
        exactly as it would have without the feature. Never raises into a run.
        """
        store = self._system_instructions_store()
        if store is None:
            return None
        try:
            doc = store.get()
        except Exception:
            logging.warning("Could not read the custom system instructions.", exc_info=True)
            return None
        if doc is None or not doc.enabled or not doc.template.strip():
            return None

        surface = provenance.surface if provenance else "unknown"
        context = InstructionContext(
            surface=surface,
            # The enum MEMBER, not its ``.value``. Pydantic coerces the bare
            # string back to the member, so this is runtime-identical — but the
            # field is declared ``SessionOrigin`` and ``describe()`` walks the
            # schema to generate the operator-facing variable reference, so the
            # typed contract has to be honoured at the boundary rather than
            # widened to satisfy a caller.
            origin=provenance.origin if provenance else SessionOrigin.USER,
            is_mobile=is_mobile_surface(surface),
            session_id=session_id,
            model=self._model_name,
            cwd=self._cwd or str(Path.cwd()),
            platform=_platform.system().lower(),
            hostname=socket.gethostname(),
            mewbo_version=get_version(),
            capabilities=tuple(session_caps),
            tools=self._resolve_instruction_tools(
                tool_specs=tool_specs,
                allowed_tools=allowed_tools,
                session_caps=session_caps,
                extra_session_tools=extra_session_tools,
                strict_tool_scope=strict_tool_scope,
                capability_mode=capability_mode,
                denied_tools=denied_tools,
            ),
            # ``project`` is absent for a managed worktree too, not just for an
            # unscoped session: ``TraceProvenance._facets_from_context`` routes a
            # ``managed:<uuid>`` context value to a ``worktree`` facet and never
            # into ``metadata["project"]``. The field's description says so.
            project=provenance.metadata.get("project") if provenance else None,
        )
        rendered = doc.render(context)
        # Persist the render outcome so the settings UI can surface a broken
        # template (which is otherwise invisible: it degrades to no injection
        # rather than to a failure). Best-effort — ``record_error`` swallows.
        store.record_error(doc.id, rendered.error)
        return rendered.text or None

    def _maybe_generate_title(self, session_id: str) -> None:
        """Kick off non-blocking title generation.

        Runs only once per session (guarded by ``load_title`` absence). The
        caller returns immediately; the title appears via a ``title_update``
        event whenever the LLM call finishes. Failures are logged, never
        raised — the first-user-message fallback remains as safety net.

        Two backgrounding strategies keep the sync and async worlds happy:

        * Pyodide / single-threaded runtimes (``sys.platform == "emscripten"``):
          schedule the coroutine on the already-running event loop. The WebLoop
          persists past the per-request lifecycle, so the task survives — and we
          never touch ``threading`` (which is inlined there) nor
          ``asyncio.run`` (illegal inside a running loop).
        * CPython: spawn a daemon thread that owns its own ``asyncio.run``
          lifecycle — identical to the pre-async-surface behaviour.
        """
        if self._session_store.load_title(session_id) is not None:
            return
        if sys.platform == "emscripten":
            try:
                loop = asyncio.get_running_loop()
                task = loop.create_task(self._run_title_generation_async(session_id))
                # Hold a strong reference so asyncio can't GC this
                # fire-and-forget task mid-run; the done-callback discards it
                # once finished (see ``self._background_tasks`` in __init__).
                self._background_tasks.add(task)
                task.add_done_callback(self._background_tasks.discard)
                return
            except RuntimeError:
                pass  # No running loop yet — fall through to the thread path.
        threading.Thread(
            target=self._run_title_generation,
            args=(session_id,),
            name=f"title-gen-{session_id[:8]}",
            daemon=True,
        ).start()

    def _run_title_generation(self, session_id: str) -> None:
        """Sync worker body for background title generation (CPython thread).

        No try/except here: the coroutine body already catches and logs
        every exception internally (see :meth:`_run_title_generation_async`),
        so ``asyncio.run`` never raises out of this call.
        """
        asyncio.run(self._run_title_generation_async(session_id))

    async def _run_title_generation_async(self, session_id: str) -> None:
        """Async title generation — single source of truth for the LLM call."""
        try:
            from mewbo_core.session.title_generator import generate_session_title

            events = self._session_store.load_transcript(session_id)
            title = await generate_session_title(events)
            if not title:
                return
            self._session_store.save_title(session_id, title)
            self._session_store.append_event(
                session_id,
                {"type": "title_update", "payload": {"title": title}},
            )
        except Exception as exc:
            logging.warning("Title generation failed: {}: {}", type(exc).__name__, exc)

    def _maybe_auto_compact(self, session_id: str) -> str | None:
        """Sync wrapper around :meth:`_maybe_auto_compact_async`.

        Kept for callers that drive compaction from a synchronous context
        (test harnesses asserting the thrash-guard short-circuit).
        """
        return asyncio.run(self._maybe_auto_compact_async(session_id))

    async def _maybe_auto_compact_async(self, session_id: str) -> str | None:
        from mewbo_core.session.compact import (
            CompactionMode,
            compact_conversation,
            record_compaction,
        )
        from mewbo_core.session.token_budget import read_last_input_tokens

        raw_events = self._session_store.load_transcript(session_id)
        # Thrash guard: if the most recent event is a compaction marker, the
        # transcript was already summarized this turn (manual /compact or a
        # prior auto cycle). Re-running on stale ``last_input_tokens`` from
        # before the boundary would clobber the fresh summary with a partial
        # one. Skip — the next real LLM call will refresh the budget read.
        if raw_events and raw_events[-1].get("type") == "context_compacted":
            return None

        events = self._hook_manager.run_pre_compact(raw_events)
        summary = self._session_store.load_summary(session_id)
        last_input_tokens = read_last_input_tokens(events)
        budget = get_token_budget(
            events,
            summary,
            self._model_name,
            last_input_tokens=last_input_tokens,
        )
        if not budget.needs_compact:
            return None
        compact_model = self._model_name
        try:
            result = await compact_conversation(events, CompactionMode.PARTIAL)
            compact_model = result.model or self._model_name
            summary = result.summary
            tokens_saved = result.tokens_saved
        except Exception:
            # Structured compaction failed. Do NOT substitute concatenated raw
            # event text — it would poison the context. Skip this cycle; the
            # next turn will try again.
            logging.warning("Structured compaction failed; skipping cycle", exc_info=True)
            return None
        record_compaction(
            self._session_store,
            self._hook_manager,
            session_id,
            summary=summary,
            mode="auto",
            model=compact_model or "",
            tokens_before=budget.total_tokens,
            tokens_saved=tokens_saved,
            events_summarized=len(events),
        )
        return summary

    @staticmethod
    def _reconcile_missing_plugins(plugins_cfg: PluginsConfig) -> None:
        """Ensure all ``enabled_plugins`` exist in the registry.

        On a fresh container or after a volume wipe, enabled plugins may be
        listed in the config but absent from the registry/cache.  This method
        discovers which are missing and attempts to install them from the
        configured marketplaces.  Errors are logged and skipped — session
        startup should not fail because a plugin couldn't be fetched.
        """
        from mewbo_core.tooling.plugins import (
            discover_installed_plugins,
            discover_marketplace_plugins,
            install_plugin,
        )

        cfg = plugins_cfg
        registry_paths = cfg.resolve_registry_paths()
        installed = discover_installed_plugins(registry_paths=registry_paths)
        installed_names = {pc.manifest.name for pc in installed if pc.manifest is not None}
        missing = [
            name.split("@")[0]
            for name in cfg.enabled_plugins
            if name.split("@")[0] not in installed_names
        ]
        if not missing:
            return

        from mewbo_core.common import get_logger

        _log = get_logger(name="core.orchestrator")
        _log.info("Reconciling {} missing plugin(s): {}", len(missing), missing)

        marketplace_dirs = cfg.resolve_marketplace_dirs()
        available = discover_marketplace_plugins(marketplace_dirs=marketplace_dirs)
        available_by_name = {p["name"]: p for p in available}

        for name in missing:
            match = available_by_name.get(name)
            if match is None:
                _log.warning("Plugin '{}' not found in any marketplace — skipping", name)
                continue
            try:
                install_plugin(
                    name,
                    match["marketplace"],
                    marketplace_dirs=marketplace_dirs,
                    install_base=cfg.resolve_install_dir(),
                )
                _log.info("Auto-installed plugin '{}'", name)
            except Exception as exc:
                _log.warning("Failed to auto-install plugin '{}': {}", name, exc)

    @staticmethod
    def _should_update_summary(text: str) -> bool:
        lowered = text.lower()
        keywords = [
            "remember",
            "note this",
            "save this",
            "pin this",
            "keep this",
            "magic number",
            "magic numbers",
        ]
        return any(keyword in lowered for keyword in keywords)

    def _update_summary_with_memory(self, session_id: str, text: str) -> str:
        summary = self._session_store.load_summary(session_id) or ""
        new_line = f"Memory: {text}"
        lines = [line for line in summary.splitlines() if line.strip()] if summary else []
        if new_line not in lines:
            lines.append(new_line)
        updated = "\n".join(lines[-10:]).strip()
        self._session_store.save_summary(session_id, updated)
        return updated

    @staticmethod
    def _build_direct_response(message: str) -> TaskQueue:
        task_queue = TaskQueue(action_steps=[])
        task_queue.task_result = message
        return task_queue

    def _try_skill_invocation(
        self,
        user_query: str,
        tool_specs: list,
        session_capabilities: tuple[str, ...] = (),
    ) -> tuple[str | None, list | None]:
        """Detect ``/skill-name args`` in the query and activate the skill.

        Honours capability gating so a slash command for a gated skill is
        inert in sessions that haven't advertised the matching capability —
        same semantics as the LLM's ``activate_skill`` tool dispatch.

        Returns ``(skill_instructions, scoped_tool_specs)`` on match,
        or ``(None, None)`` if the query is not a skill invocation.
        """
        query = user_query.strip()
        if not query.startswith("/"):
            return None, None

        parts = query.split(None, 1)
        name = parts[0].lstrip("/")
        args = parts[1] if len(parts) > 1 else ""

        skill = self._skill_registry.get(name, session_capabilities)
        if skill is None:
            return None, None

        logging.info("Activating skill '{}' with args '{}'", name, args)
        instructions, scoped_specs = activate_skill(skill, args, tool_specs, cwd=self._cwd)
        return instructions, scoped_specs

    def _last_context(self, session_id: str) -> dict[str, object]:
        """The session's effective context: the most-recent payload, VERBATIM.

        Deliberately NOT ``SessionStoreBase.latest_context``, which FOLDS every
        context event. The two are different questions and this is the one the
        consumers of a re-written context event ask: the API's
        ``_load_last_context`` and the console's ``getLastContext`` both reverse-
        scan for the newest payload and use it as-is. A writer that wants to be
        neutral for them has to carry forward what THEY would have read, and a
        fold would hand back fields an earlier turn deliberately cleared.

        The scan is ``latest_event_of_type`` on the store, beside
        ``merge_context_events`` — the two reducers are a pair and both live
        there, so the API's copy and this one ask one implementation rather than
        each carrying its own. ``O(1)`` on the Mongo driver,
        ``O(one session)`` on the base; bounded by the TYPE, never by a count,
        because the newest context event sits arbitrarily far back after a long
        run and a window that missed it would hand the writer an empty payload to
        carry forward — silently clearing the very fields it exists to preserve.

        Returns a COPY: the caller merges into it before persisting the result as
        the next context event.
        """
        event = self._session_store.latest_event_of_type(session_id, "context")
        payload = event.get("payload") if event else None
        return dict(payload) if isinstance(payload, dict) else {}

    @staticmethod
    def _build_project_catalog() -> ProjectCatalog:
        """Assemble the catalog the project-selection tools read and resolve through.

        Built here rather than in ``__init__`` because it is only ever needed by
        a run that opted into project autoselect: constructing it opens the
        project and repository stores, and an ordinary run has no reason to
        touch either. Each store is optional — a backend that refuses to open
        costs that ONE section of the catalog, exactly as a backend that raises
        while LISTING does, rather than leaving the model with no projects at
        all.

        ``checkout_locator`` is deliberately left unset. Matching a repository
        identity against a managed project's git remotes is the API app's job,
        above core in the DAG, so a core-only caller gets registered
        repositories listed WITHOUT a path — the honest answer here rather than
        a guessed one.
        """
        from mewbo_core.workspaces.project_catalog import ProjectCatalog
        from mewbo_core.workspaces.project_store import ProjectStoreBase, create_project_store
        from mewbo_core.workspaces.repository_store import (
            RepositoryStoreBase,
            create_repository_store,
        )

        project_store: ProjectStoreBase | None = None
        repository_store: RepositoryStoreBase | None = None
        try:
            project_store = create_project_store()
        except Exception as exc:  # noqa: BLE001 - one dead store, not a dead catalog
            logging.warning("Project store unavailable for the project catalog: {}", exc)
        try:
            repository_store = create_repository_store()
        except Exception as exc:  # noqa: BLE001 - one dead store, not a dead catalog
            logging.warning("Repository store unavailable for the project catalog: {}", exc)
        return ProjectCatalog(
            configured=get_config().projects,
            project_store=project_store,
            repository_store=repository_store,
        )

    @staticmethod
    def _resolve_mode(mode: str | None) -> str:
        """Resolve the orchestration mode.

        Only the explicit ``mode`` parameter is honoured — keyword heuristics
        on the user query have been removed because they produced fragile,
        surprising behaviour (accidentally entering plan mode on innocent
        phrasing). Clients must pass ``mode="plan"`` explicitly to opt in.
        """
        if mode in {"plan", "act"}:
            return mode
        return "act"

__init__(*, model_name: str | None = None, fallback_models: tuple[str, ...] | None = None, session_store: SessionStoreBase | None = None, tool_registry: ToolRegistry | None = None, permission_policy: PermissionPolicy | None = None, approval_callback: Callable[[ActionStep], bool] | None = None, hook_manager: HookManager | None = None, cwd: str | None = None, session_step_budget: int = 0, system_instructions_store: SystemInstructionsStoreBase | None = None, session_mcp_servers: dict[str, dict] | None = None) -> None

Initialize orchestration dependencies.

session_mcp_servers are MCP servers a caller attaches to THIS run, in the standard {name: {command|url, …}} config shape. They are merged into the registry build below alongside the plugin-contributed ones, and their discovered tool ids are admitted through this run's allowed_tools (:meth:_session_mcp_tool_ids).

It is deliberately NOT named extra_mcp_servers, even though that is the parameter it feeds. The two carry different grants, and reusing one name for both senses is the trap the house rules name outright: a plugin-contributed server's tools are subject to allowed_tools like any other registry tool, while a server named HERE is admitted by the act of naming it — the caller attaching a server for one run IS the grant, and it has no other way to express one, since a server's tool ids are not knowable until discovery has run. Widening the existing parameter's meaning instead would silently admit every plugin server's tools into every scoped session in the deployment.

Discovery is a network/subprocess cost and it belongs HERE rather than at whichever caller assembled the config: this constructor already runs on the background run thread, which is the offline side of the acceptance/execution boundary. A caller resolving the ids itself would be paying for discovery on its own request path.

Source code in packages/mewbo_core/src/mewbo_core/loop/orchestrator.py
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
def __init__(
    self,
    *,
    model_name: str | None = None,
    fallback_models: tuple[str, ...] | None = None,
    session_store: SessionStoreBase | None = None,
    tool_registry: ToolRegistry | None = None,
    permission_policy: PermissionPolicy | None = None,
    approval_callback: Callable[[ActionStep], bool] | None = None,
    hook_manager: HookManager | None = None,
    cwd: str | None = None,
    session_step_budget: int = 0,
    system_instructions_store: SystemInstructionsStoreBase | None = None,
    session_mcp_servers: dict[str, dict] | None = None,
) -> None:
    """Initialize orchestration dependencies.

    *session_mcp_servers* are MCP servers a caller attaches to THIS run, in
    the standard ``{name: {command|url, …}}`` config shape. They are merged
    into the registry build below alongside the plugin-contributed ones, and
    their discovered tool ids are admitted through this run's
    ``allowed_tools`` (:meth:`_session_mcp_tool_ids`).

    **It is deliberately NOT named ``extra_mcp_servers``, even though that is
    the parameter it feeds.** The two carry different grants, and reusing one
    name for both senses is the trap the house rules name outright: a
    plugin-contributed server's tools are subject to ``allowed_tools`` like
    any other registry tool, while a server named HERE is admitted by the act
    of naming it — the caller attaching a server for one run IS the grant, and
    it has no other way to express one, since a server's tool ids are not
    knowable until discovery has run. Widening the existing parameter's
    meaning instead would silently admit every plugin server's tools into
    every scoped session in the deployment.

    Discovery is a network/subprocess cost and it belongs HERE rather than at
    whichever caller assembled the config: this constructor already runs on
    the background run thread, which is the offline side of the
    acceptance/execution boundary. A caller resolving the ids itself would be
    paying for discovery on its own request path.
    """
    self._cwd = cwd
    self._session_mcp_servers = dict(session_mcp_servers or {})
    self._session_step_budget = session_step_budget
    # Resolved LAZILY on first use (``_instructions_store``): the factory
    # raises when the configured driver is mongodb and Mongo is unreachable,
    # and a custom-instructions store being down must never stop a session
    # from starting. ``_instructions_store_tried`` makes the failure sticky
    # so we don't retry a dead connection once per run.
    self._instructions_store: SystemInstructionsStoreBase | None = system_instructions_store
    self._instructions_store_tried = system_instructions_store is not None
    # Strong references for fire-and-forget background tasks scheduled on
    # an already-running loop (the emscripten branch of
    # ``_maybe_generate_title``) so asyncio can't GC them mid-run — the
    # done-callback discards its own entry once finished. Mirrors
    # ``AgentHandle.asyncio_task`` (hypervisor.py) / ``_lifecycle_tasks``
    # (spawn_agent.py).
    self._background_tasks: set[asyncio.Task] = set()
    self._model_name = (
        model_name
        or get_config_value("llm", "action_plan_model")
        or get_config_value("llm", "default_model", default="gpt-5.2")
    )
    self._fallback_models = (
        fallback_models
        if fallback_models is not None
        else tuple(effective_fallback_models())
    )
    self._session_store = session_store or create_session_store()
    self._permission_policy = permission_policy or load_permission_policy()
    self._approval_callback = approval_callback or approval_callback_from_config()
    self._hook_manager = hook_manager or default_hook_manager()

    self._project_instructions = discover_project_instructions(cwd)
    # Loaded exactly once, here, from operator-owned config — never from a
    # tool a model can call. ``None`` when the deployment-wide switch is off
    # (the default): every call site below tests ``is not None`` and does
    # nothing else, so an unconfigured install never stats ``.mewbo/``,
    # never parses a document and never builds a rule.
    self._safety_plane = SafetyPlane.load(
        cwd, enabled=bool(get_config_value("safety", "enabled", default=False))
    )
    self._skill_registry = SkillRegistry()
    self._skill_registry.load(cwd)

    # Plugin loading (before load_registry so MCP servers are collected first).
    # Uses the shared load_all_plugin_components() so the same logic is
    # reused by the API /skills and /tools endpoints (DRY).
    self._agent_registry = None
    self._session_tool_registry = SessionToolRegistry()
    # schedule_trigger rides the ordinary SessionToolRegistry so
    # a spawned sub-agent whose allowlist admits it can bind it — the old
    # root-only extra_session_tools seam structurally could not. The factory
    # exists only once the app pushed its store+policy; None (CLI/tests) ⇒
    # the tool is simply absent, exactly as before.
    from mewbo_core.triggers.session_tool import schedule_trigger_factory

    _trigger_factory = schedule_trigger_factory()
    if _trigger_factory is not None:
        self._session_tool_registry.register(_trigger_factory)
    plugins_cfg = get_config().plugins

    if plugins_cfg.enabled:
        # Reconcile missing plugins on fresh containers / volume wipes.
        if plugins_cfg.enabled_plugins:
            self._reconcile_missing_plugins(plugins_cfg)

        from pathlib import Path as _Path

        from mewbo_core.agents.agent_registry import AgentRegistry, parse_agent_file
        from mewbo_core.hooks import merge_plugin_hooks
        from mewbo_core.tooling.plugins import load_all_plugin_components

        fan_out = load_all_plugin_components()
        self._agent_registry = AgentRegistry()

        self._skill_registry.load_plugin_components(fan_out)
        for pc in fan_out.components:
            if pc.manifest is None:
                continue
            plugin_source = f"plugin:{pc.manifest.name}"
            plugin_caps = pc.manifest.requires_capabilities
            plugin_root = pc.manifest.install_path
            for af in pc.agent_files:
                agent_def = parse_agent_file(_Path(af), source=plugin_source)
                if agent_def is None:
                    continue
                self._agent_registry.register(
                    agent_def,
                    capabilities=plugin_caps,
                    plugin_root=plugin_root,
                )
            # Load this plugin's session tools WITH its capability gate, so a
            # capability-gated session tool (e.g. the ``scg`` suite's
            # ``scg_*``) surfaces to any session that holds the capability —
            # including via the runtime grant — not only when the
            # client lists the tool in ``allowed_tools``. Same
            # data-driven gate the AgentDefs register through above.
            for entry in pc.session_tool_entries:
                self._session_tool_registry.load_entry(
                    entry, requires_capabilities=plugin_caps
                )
        for hooks_json, plugin_root in fan_out.hooks_configs:
            merge_plugin_hooks(self._hook_manager, hooks_json, plugin_root)
        plugin_mcp_servers = fan_out.mcp_servers
    else:
        plugin_mcp_servers = {}

    # Reuse a cached registry across runs whose build inputs (cwd +
    # plugin-contributed MCP servers + MCP-config fingerprint) are identical,
    # instead of rebuilding per query — the per-run rebuild was on the
    # critical path to the first FE-visible event. An explicit
    # ``tool_registry`` (tests, structured runners) still bypasses the cache.
    # The cache key already covers ``(cwd, extra_mcp_servers, mcp-config
    # fingerprint)``, so a run attaching its own servers gets its own entry
    # instead of poisoning the shared one — which is what makes merging a
    # per-run set in here safe at all.
    self._tool_registry = tool_registry or get_or_build_registry(
        cwd=cwd,
        extra_mcp_servers={**plugin_mcp_servers, **self._session_mcp_servers} or None,
    )
    self._context_builder = ContextBuilder(self._session_store)

    # Register lossless micro-compaction as a pre_compact hook.
    from mewbo_core.session.compaction import micro_compact_events

    self._hook_manager.pre_compact.append(micro_compact_events)

arun(user_query: str, *, max_iters: int = 3, initial_plan: Plan | None = None, return_state: bool = False, session_id: str | None = None, mode: str | None = None, should_cancel: Callable[[], bool] | None = None, allowed_tools: list[str] | None = None, denied_tools: list[str] | None = None, strict_tool_scope: bool = False, capability_mode: str = 'all', skill_instructions: str | None = None, message_queue: queue.Queue[str] | None = None, interrupt_step: threading.Event | None = None, user_id: str | None = None, source_platform: str | None = None, invocation_id: str | None = None, extra_session_tools: list[SessionTool] | None = None, enable_skills: bool = True, project_autoselect: bool = False, attachments: list[dict] | None = None, user_turn_persisted: bool = False) -> TaskQueue | tuple[TaskQueue, OrchestrationState] async

Run orchestration asynchronously (async-first entry point).

Awaits the orchestration coroutine directly, so an environment that already owns a running event loop (browser-hosted Pyodide's WebLoop, async test harnesses) can drive a query with native await and NO nested asyncio.run. The sync :meth:run wrapper is the CPython entry point and delegates here; the two share this single body.

Source code in packages/mewbo_core/src/mewbo_core/loop/orchestrator.py
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
async def arun(
    self,
    user_query: str,
    *,
    max_iters: int = 3,
    initial_plan: Plan | None = None,
    return_state: bool = False,
    session_id: str | None = None,
    mode: str | None = None,
    should_cancel: Callable[[], bool] | None = None,
    allowed_tools: list[str] | None = None,
    denied_tools: list[str] | None = None,
    strict_tool_scope: bool = False,
    capability_mode: str = "all",
    skill_instructions: str | None = None,
    message_queue: queue.Queue[str] | None = None,
    interrupt_step: threading.Event | None = None,
    user_id: str | None = None,
    source_platform: str | None = None,
    invocation_id: str | None = None,
    extra_session_tools: list[SessionTool] | None = None,
    enable_skills: bool = True,
    project_autoselect: bool = False,
    attachments: list[dict] | None = None,
    user_turn_persisted: bool = False,
) -> TaskQueue | tuple[TaskQueue, OrchestrationState]:
    """Run orchestration asynchronously (async-first entry point).

    Awaits the orchestration coroutine directly, so an environment that
    already owns a running event loop (browser-hosted Pyodide's WebLoop,
    async test harnesses) can drive a query with native ``await`` and NO
    nested ``asyncio.run``. The sync :meth:`run` wrapper is the CPython
    entry point and delegates here; the two share this single body.
    """
    if session_id is None:
        session_id = self._session_store.create_session()

    # Fold the session's durable signals (tags + merged context + the
    # entry-point surface) into filterable Langfuse tags/metadata once; the
    # seam propagates them to every child observation. Apps write their
    # context/tags before invoking the runtime, so they are present here.
    #
    # The merged context carries only the CLIENT-ADVERTISED capabilities, so
    # overlay the AUGMENTED set (advertised ∪ runtime-provider grants) before
    # deriving — otherwise a capability granted at runtime (the ``scg``
    # provider) would be live in the run yet invisible in the trace's
    # ``capabilities`` facet. ``derive`` stays a pure transform of
    # (tags, context, surface); we only enrich the context it reads.
    derive_context = dict(self._session_store.latest_context(session_id))
    effective_caps = self._session_capabilities(session_id, allowed_tools=allowed_tools)
    if effective_caps:
        derive_context["client_capabilities"] = list(effective_caps)
    # An explicit caller-supplied user_id always wins. Otherwise, the real
    # principal — stamped into the context payload by the API's
    # ``_stamp_principal_subject`` whenever auth is enabled — is the
    # honest identity for tracing. Leave it None (never the session id)
    # when no principal exists, so the seam's own anonymous fallback
    # applies instead of silently aliasing user_id to session_id.
    if user_id is None:
        principal_subject = derive_context.get("principal_subject")
        if isinstance(principal_subject, str) and principal_subject:
            user_id = principal_subject
    provenance = TraceProvenance.derive(
        tags=self._session_store.tags_for_session(session_id),
        context=derive_context,
        surface=source_platform,
    )

    with session_log_context(session_id):
        with langfuse_session_context(
            session_id,
            user_id=user_id,
            invocation_id=invocation_id,
            source_platform=source_platform,
            # The trace name identifies the KIND of turn, never this
            # execution of it: a name carrying a turn index, a session id
            # or the query text mints a fresh name per run and every
            # saved filter, evaluator and dashboard stops matching. The
            # surface is a bounded set, so it stays groupable.
            trace_name=f"turn:{source_platform or 'unknown'}",
            tags=list(provenance.tags),
            metadata=provenance.metadata,
        ):
            return await self._run_with_session_context_async(
                user_query,
                max_iters=max_iters,
                initial_plan=initial_plan,
                return_state=return_state,
                session_id=session_id,
                mode=mode,
                should_cancel=should_cancel,
                allowed_tools=allowed_tools,
                denied_tools=denied_tools,
                strict_tool_scope=strict_tool_scope,
                capability_mode=capability_mode,
                skill_instructions=skill_instructions,
                message_queue=message_queue,
                interrupt_step=interrupt_step,
                extra_session_tools=extra_session_tools,
                enable_skills=enable_skills,
                project_autoselect=project_autoselect,
                attachments=attachments,
                user_turn_persisted=user_turn_persisted,
                provenance=provenance,
            )

run(user_query: str, *, max_iters: int = 3, initial_plan: Plan | None = None, return_state: bool = False, session_id: str | None = None, mode: str | None = None, should_cancel: Callable[[], bool] | None = None, allowed_tools: list[str] | None = None, denied_tools: list[str] | None = None, strict_tool_scope: bool = False, capability_mode: str = 'all', skill_instructions: str | None = None, message_queue: queue.Queue[str] | None = None, interrupt_step: threading.Event | None = None, user_id: str | None = None, source_platform: str | None = None, invocation_id: str | None = None, extra_session_tools: list[SessionTool] | None = None, enable_skills: bool = True, project_autoselect: bool = False, attachments: list[dict] | None = None, user_turn_persisted: bool = False) -> TaskQueue | tuple[TaskQueue, OrchestrationState]

Run orchestration for a session (sync wrapper around :meth:arun).

Backward-compatible synchronous entry point for CLI / API request threads / tests. Owns the event-loop lifecycle via asyncio.run, so every existing sync caller behaves exactly as before. Environments that already drive an event loop (browser-hosted Pyodide's WebLoop) must call :meth:arun instead — it awaits the same body with NO nested asyncio.run (the "WebLoop wall").

Source code in packages/mewbo_core/src/mewbo_core/loop/orchestrator.py
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
def run(
    self,
    user_query: str,
    *,
    max_iters: int = 3,
    initial_plan: Plan | None = None,
    return_state: bool = False,
    session_id: str | None = None,
    mode: str | None = None,
    should_cancel: Callable[[], bool] | None = None,
    allowed_tools: list[str] | None = None,
    denied_tools: list[str] | None = None,
    strict_tool_scope: bool = False,
    capability_mode: str = "all",
    skill_instructions: str | None = None,
    message_queue: queue.Queue[str] | None = None,
    interrupt_step: threading.Event | None = None,
    user_id: str | None = None,
    source_platform: str | None = None,
    invocation_id: str | None = None,
    extra_session_tools: list[SessionTool] | None = None,
    enable_skills: bool = True,
    project_autoselect: bool = False,
    attachments: list[dict] | None = None,
    user_turn_persisted: bool = False,
) -> TaskQueue | tuple[TaskQueue, OrchestrationState]:
    """Run orchestration for a session (sync wrapper around :meth:`arun`).

    Backward-compatible synchronous entry point for CLI / API request
    threads / tests. Owns the event-loop lifecycle via ``asyncio.run``, so
    every existing sync caller behaves exactly as before. Environments that
    already drive an event loop (browser-hosted Pyodide's WebLoop) must call
    :meth:`arun` instead — it awaits the same body with NO nested
    ``asyncio.run`` (the "WebLoop wall").
    """
    return asyncio.run(
        self.arun(
            user_query,
            max_iters=max_iters,
            initial_plan=initial_plan,
            return_state=return_state,
            session_id=session_id,
            mode=mode,
            should_cancel=should_cancel,
            allowed_tools=allowed_tools,
            denied_tools=denied_tools,
            strict_tool_scope=strict_tool_scope,
            capability_mode=capability_mode,
            skill_instructions=skill_instructions,
            message_queue=message_queue,
            interrupt_step=interrupt_step,
            user_id=user_id,
            source_platform=source_platform,
            invocation_id=invocation_id,
            extra_session_tools=extra_session_tools,
            enable_skills=enable_skills,
            project_autoselect=project_autoselect,
            attachments=attachments,
            user_turn_persisted=user_turn_persisted,
        )
    )

mewbo_core.loop.task_master

Task planning and orchestration loop for Mewbo.

generate_action_plan(user_query: str, model_name: str | None = None, tool_registry: ToolRegistry | None = None, session_summary: str | None = None, recent_events: list[EventRecord] | None = None, selected_events: list[EventRecord] | None = None, *, mode: str = 'act', feedback: str | None = None) -> Plan

Generate a plan for a user query.

Source code in packages/mewbo_core/src/mewbo_core/loop/task_master.py
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
def generate_action_plan(
    user_query: str,
    model_name: str | None = None,
    tool_registry: ToolRegistry | None = None,
    session_summary: str | None = None,
    recent_events: list[EventRecord] | None = None,
    selected_events: list[EventRecord] | None = None,
    *,
    mode: str = "act",
    feedback: str | None = None,
) -> Plan:
    """Generate a plan for a user query."""
    tool_registry = tool_registry or load_registry()
    resolved_model = cast(
        str,
        model_name
        or get_config_value("llm", "action_plan_model")
        or get_config_value("llm", "default_model", default="gpt-5.2"),
    )
    context = _build_context_snapshot(
        session_summary,
        recent_events,
        selected_events,
        resolved_model,
    )
    return Planner(tool_registry).generate(
        user_query,
        resolved_model,
        context=context,
        mode=mode,
        feedback=feedback,
    )

orchestrate_session(user_query: str, model_name: str | None = None, fallback_models: tuple[str, ...] | None = None, max_iters: int = 3, initial_plan: Plan | None = None, return_state: bool = False, session_id: str | None = None, session_store: SessionStoreBase | None = None, tool_registry: ToolRegistry | None = None, permission_policy: PermissionPolicy | None = None, approval_callback: Callable[[ActionStep], bool] | None = None, hook_manager: HookManager | None = None, mode: str | None = None, should_cancel: Callable[[], bool] | None = None, allowed_tools: list[str] | None = None, denied_tools: list[str] | None = None, strict_tool_scope: bool = False, capability_mode: str = 'all', skill_instructions: str | None = None, message_queue: queue.Queue[str] | None = None, interrupt_step: threading.Event | None = None, cwd: str | None = None, session_step_budget: int = 0, user_id: str | None = None, source_platform: str | None = None, invocation_id: str | None = None, extra_session_tools: list[SessionTool] | None = None, enable_skills: bool = True, project_autoselect: bool = False, attachments: list[dict] | None = None, user_turn_persisted: bool = False, session_mcp_servers: dict[str, dict] | None = None) -> TaskQueue | tuple[TaskQueue, OrchestrationState]

Run the orchestration loop synchronously.

Thin sync wrapper over :func:orchestrate_session_async — mirrors Orchestrator.run() = asyncio.run(self.arun(...)) one layer up, so there is exactly one place (here) that constructs the Orchestrator and forwards every kwarg for both the sync and async entry points.

Source code in packages/mewbo_core/src/mewbo_core/loop/task_master.py
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
def orchestrate_session(
    user_query: str,
    model_name: str | None = None,
    fallback_models: tuple[str, ...] | None = None,
    max_iters: int = 3,
    initial_plan: Plan | None = None,
    return_state: bool = False,
    session_id: str | None = None,
    session_store: SessionStoreBase | None = None,
    tool_registry: ToolRegistry | None = None,
    permission_policy: PermissionPolicy | None = None,
    approval_callback: Callable[[ActionStep], bool] | None = None,
    hook_manager: HookManager | None = None,
    mode: str | None = None,
    should_cancel: Callable[[], bool] | None = None,
    allowed_tools: list[str] | None = None,
    denied_tools: list[str] | None = None,
    strict_tool_scope: bool = False,
    capability_mode: str = "all",
    skill_instructions: str | None = None,
    message_queue: queue.Queue[str] | None = None,
    interrupt_step: threading.Event | None = None,
    cwd: str | None = None,
    session_step_budget: int = 0,
    user_id: str | None = None,
    source_platform: str | None = None,
    invocation_id: str | None = None,
    extra_session_tools: list[SessionTool] | None = None,
    enable_skills: bool = True,
    project_autoselect: bool = False,
    attachments: list[dict] | None = None,
    user_turn_persisted: bool = False,
    session_mcp_servers: dict[str, dict] | None = None,
) -> TaskQueue | tuple[TaskQueue, OrchestrationState]:
    """Run the orchestration loop synchronously.

    Thin sync wrapper over :func:`orchestrate_session_async` — mirrors
    ``Orchestrator.run() = asyncio.run(self.arun(...))`` one layer up, so
    there is exactly one place (here) that constructs the ``Orchestrator``
    and forwards every kwarg for both the sync and async entry points.
    """
    return asyncio.run(
        orchestrate_session_async(
            user_query,
            model_name=model_name,
            fallback_models=fallback_models,
            max_iters=max_iters,
            initial_plan=initial_plan,
            return_state=return_state,
            session_id=session_id,
            session_store=session_store,
            tool_registry=tool_registry,
            permission_policy=permission_policy,
            approval_callback=approval_callback,
            hook_manager=hook_manager,
            mode=mode,
            should_cancel=should_cancel,
            allowed_tools=allowed_tools,
            denied_tools=denied_tools,
            strict_tool_scope=strict_tool_scope,
            capability_mode=capability_mode,
            skill_instructions=skill_instructions,
            message_queue=message_queue,
            interrupt_step=interrupt_step,
            cwd=cwd,
            session_step_budget=session_step_budget,
            user_id=user_id,
            source_platform=source_platform,
            invocation_id=invocation_id,
            extra_session_tools=extra_session_tools,
            enable_skills=enable_skills,
            project_autoselect=project_autoselect,
            attachments=attachments,
            user_turn_persisted=user_turn_persisted,
            session_mcp_servers=session_mcp_servers,
        )
    )

orchestrate_session_async(user_query: str, model_name: str | None = None, fallback_models: tuple[str, ...] | None = None, max_iters: int = 3, initial_plan: Plan | None = None, return_state: bool = False, session_id: str | None = None, session_store: SessionStoreBase | None = None, tool_registry: ToolRegistry | None = None, permission_policy: PermissionPolicy | None = None, approval_callback: Callable[[ActionStep], bool] | None = None, hook_manager: HookManager | None = None, mode: str | None = None, should_cancel: Callable[[], bool] | None = None, allowed_tools: list[str] | None = None, denied_tools: list[str] | None = None, strict_tool_scope: bool = False, capability_mode: str = 'all', skill_instructions: str | None = None, message_queue: queue.Queue[str] | None = None, interrupt_step: threading.Event | None = None, cwd: str | None = None, session_step_budget: int = 0, user_id: str | None = None, source_platform: str | None = None, invocation_id: str | None = None, extra_session_tools: list[SessionTool] | None = None, enable_skills: bool = True, project_autoselect: bool = False, attachments: list[dict] | None = None, user_turn_persisted: bool = False, session_mcp_servers: dict[str, dict] | None = None) -> TaskQueue | tuple[TaskQueue, OrchestrationState] async

Run the orchestration loop on an already-running event loop.

Async mirror of :func:orchestrate_session for environments (Pyodide's WebLoop, async test harnesses) where asyncio.run cannot be used because the loop is already active. Same signature and semantics as the sync entry point — it awaits :meth:Orchestrator.arun instead of calling :meth:Orchestrator.run, so no nested asyncio.run is created.

Source code in packages/mewbo_core/src/mewbo_core/loop/task_master.py
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
async def orchestrate_session_async(
    user_query: str,
    model_name: str | None = None,
    fallback_models: tuple[str, ...] | None = None,
    max_iters: int = 3,
    initial_plan: Plan | None = None,
    return_state: bool = False,
    session_id: str | None = None,
    session_store: SessionStoreBase | None = None,
    tool_registry: ToolRegistry | None = None,
    permission_policy: PermissionPolicy | None = None,
    approval_callback: Callable[[ActionStep], bool] | None = None,
    hook_manager: HookManager | None = None,
    mode: str | None = None,
    should_cancel: Callable[[], bool] | None = None,
    allowed_tools: list[str] | None = None,
    denied_tools: list[str] | None = None,
    strict_tool_scope: bool = False,
    capability_mode: str = "all",
    skill_instructions: str | None = None,
    message_queue: queue.Queue[str] | None = None,
    interrupt_step: threading.Event | None = None,
    cwd: str | None = None,
    session_step_budget: int = 0,
    user_id: str | None = None,
    source_platform: str | None = None,
    invocation_id: str | None = None,
    extra_session_tools: list[SessionTool] | None = None,
    enable_skills: bool = True,
    project_autoselect: bool = False,
    attachments: list[dict] | None = None,
    user_turn_persisted: bool = False,
    session_mcp_servers: dict[str, dict] | None = None,
) -> TaskQueue | tuple[TaskQueue, OrchestrationState]:
    """Run the orchestration loop on an already-running event loop.

    Async mirror of :func:`orchestrate_session` for environments (Pyodide's
    WebLoop, async test harnesses) where ``asyncio.run`` cannot be used because
    the loop is already active. Same signature and semantics as the sync entry
    point — it awaits :meth:`Orchestrator.arun` instead of calling
    :meth:`Orchestrator.run`, so no nested ``asyncio.run`` is created.
    """
    return await Orchestrator(
        model_name=model_name,
        fallback_models=fallback_models,
        session_store=session_store,
        tool_registry=tool_registry,
        permission_policy=permission_policy,
        approval_callback=approval_callback,
        hook_manager=hook_manager,
        cwd=cwd,
        session_step_budget=session_step_budget,
        session_mcp_servers=session_mcp_servers,
    ).arun(
        user_query,
        max_iters=max_iters,
        initial_plan=initial_plan,
        return_state=return_state,
        session_id=session_id,
        mode=mode,
        should_cancel=should_cancel,
        allowed_tools=allowed_tools,
        denied_tools=denied_tools,
        strict_tool_scope=strict_tool_scope,
        capability_mode=capability_mode,
        skill_instructions=skill_instructions,
        message_queue=message_queue,
        interrupt_step=interrupt_step,
        user_id=user_id,
        source_platform=source_platform,
        invocation_id=invocation_id,
        extra_session_tools=extra_session_tools,
        enable_skills=enable_skills,
        project_autoselect=project_autoselect,
        attachments=attachments,
        user_turn_persisted=user_turn_persisted,
    )

mewbo_core.loop.tool_use_loop

Async tool-use conversation loop with sub-agent support.

ResilienceNote dataclass

Bounded, factual record of this run's LLM retry/fallback events.

The model is otherwise BLIND to its own retries: RetryStrategy.run has only two outward channels — emit (event log / console / CLI) and raise — and never touches the message list, so a run that quietly re-drives a failing model, or escalates down its ladder, never sees any of it. This note carries the SAME grounded facts the event log already holds — which model was tried, how many attempts, the error class, whether a switch occurred — into a dedicated system-prompt slot.

Holds only the most recent max_events records (hot in-process runtime state — a plain dataclass, no trust boundary). :meth:render returns the whole note, or "" when nothing has gone wrong yet, so a clean run pays nothing and the slot stays empty.

Source code in packages/mewbo_core/src/mewbo_core/loop/tool_use_loop.py
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
@dataclass
class ResilienceNote:
    """Bounded, factual record of this run's LLM retry/fallback events.

    The model is otherwise BLIND to its own retries: ``RetryStrategy.run`` has
    only two outward channels — ``emit`` (event log / console / CLI) and
    ``raise`` — and never touches the message list, so a run that quietly
    re-drives a failing model, or escalates down its ladder, never sees any of
    it. This note carries the SAME grounded facts the event log already holds —
    which model was tried, how many attempts, the error class, whether a switch
    occurred — into a dedicated system-prompt slot.

    Holds only the most recent ``max_events`` records (hot in-process runtime
    state — a plain dataclass, no trust boundary). :meth:`render` returns the
    whole note, or ``""`` when nothing has gone wrong yet, so a clean run pays
    nothing and the slot stays empty.
    """

    max_events: int = DEFAULT_RESILIENCE_NOTE_EVENTS
    _events: list[str] = field(default_factory=list)

    def record(self, event: Event) -> None:
        """Append a compact summary for a retry/fallback event; ignore the rest."""
        etype = event.get("type")
        if etype not in ("llm_retry", "llm_fallback"):
            return
        payload = event.get("payload") or {}
        self._events.append(self._summarize(etype, payload))
        if len(self._events) > self.max_events:
            del self._events[: len(self._events) - self.max_events]

    @staticmethod
    def _summarize(etype: str, payload: dict[str, Any]) -> str:
        """One factual line for a retry or a fallback event."""
        if etype == "llm_retry":
            return (
                f"retried {payload.get('model', '?')} "
                f"(attempt {payload.get('attempt', '?')}/{payload.get('max_attempts', '?')}, "
                f"{payload.get('error_type', '?')})"
            )
        return (
            f"switched {payload.get('from_model', '?')} -> {payload.get('to_model', '?')} "
            f"({payload.get('reason', '?')})"
        )

    def render(self) -> str:
        """The full note text, or ``""`` when no resilience event has occurred."""
        if not self._events:
            return ""
        lines = "\n".join(f"- {entry}" for entry in self._events)
        return f"{_RESILIENCE_NOTE_HEADER}\n{lines}"

record(event: Event) -> None

Append a compact summary for a retry/fallback event; ignore the rest.

Source code in packages/mewbo_core/src/mewbo_core/loop/tool_use_loop.py
175
176
177
178
179
180
181
182
183
def record(self, event: Event) -> None:
    """Append a compact summary for a retry/fallback event; ignore the rest."""
    etype = event.get("type")
    if etype not in ("llm_retry", "llm_fallback"):
        return
    payload = event.get("payload") or {}
    self._events.append(self._summarize(etype, payload))
    if len(self._events) > self.max_events:
        del self._events[: len(self._events) - self.max_events]

render() -> str

The full note text, or "" when no resilience event has occurred.

Source code in packages/mewbo_core/src/mewbo_core/loop/tool_use_loop.py
199
200
201
202
203
204
def render(self) -> str:
    """The full note text, or ``""`` when no resilience event has occurred."""
    if not self._events:
        return ""
    lines = "\n".join(f"- {entry}" for entry in self._events)
    return f"{_RESILIENCE_NOTE_HEADER}\n{lines}"

ToolBatch dataclass

A batch of tool calls with shared concurrency mode.

Source code in packages/mewbo_core/src/mewbo_core/loop/tool_use_loop.py
330
331
332
333
334
335
@dataclass
class ToolBatch:
    """A batch of tool calls with shared concurrency mode."""

    calls: list[Any]
    concurrent: bool

ToolCallResult dataclass

Result of executing a single tool call.

Source code in packages/mewbo_core/src/mewbo_core/loop/tool_use_loop.py
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
@dataclass(frozen=True)
class ToolCallResult:
    """Result of executing a single tool call."""

    tool_call_id: str
    tool_id: str
    content: str
    success: bool
    # Structured-envelope facts, present only for a session tool that returned
    # one. ``blocked_code`` is the envelope code when it names a condition no
    # retry can clear (credentials, reachability, permission, quota);
    # ``permanence`` is the tool's own retry verdict. Both default to ``None``,
    # so every other execution path is unchanged.
    blocked_code: str | None = None
    permanence: str | None = None
    # Image parts a multimodal tool returned (a device screenshot today).
    # ``content`` stays the STRING every existing consumer reads — the cap, the
    # ANSI strip, the event snapshot and the compaction summary are all
    # unchanged and still string-only. These parts bypass all of it and are
    # spliced into the ``ToolMessage`` alongside the text, which is what puts
    # the image inside the provider's native ``tool_result`` block.
    images: tuple[dict[str, Any], ...] = ()

ToolUseLoop

Async tool-use conversation loop.

Each instance owns one conversation with one LLM. Sub-agents are created by spawning new ToolUseLoop instances via the spawn_agent internal tool.

Source code in packages/mewbo_core/src/mewbo_core/loop/tool_use_loop.py
 348
 349
 350
 351
 352
 353
 354
 355
 356
 357
 358
 359
 360
 361
 362
 363
 364
 365
 366
 367
 368
 369
 370
 371
 372
 373
 374
 375
 376
 377
 378
 379
 380
 381
 382
 383
 384
 385
 386
 387
 388
 389
 390
 391
 392
 393
 394
 395
 396
 397
 398
 399
 400
 401
 402
 403
 404
 405
 406
 407
 408
 409
 410
 411
 412
 413
 414
 415
 416
 417
 418
 419
 420
 421
 422
 423
 424
 425
 426
 427
 428
 429
 430
 431
 432
 433
 434
 435
 436
 437
 438
 439
 440
 441
 442
 443
 444
 445
 446
 447
 448
 449
 450
 451
 452
 453
 454
 455
 456
 457
 458
 459
 460
 461
 462
 463
 464
 465
 466
 467
 468
 469
 470
 471
 472
 473
 474
 475
 476
 477
 478
 479
 480
 481
 482
 483
 484
 485
 486
 487
 488
 489
 490
 491
 492
 493
 494
 495
 496
 497
 498
 499
 500
 501
 502
 503
 504
 505
 506
 507
 508
 509
 510
 511
 512
 513
 514
 515
 516
 517
 518
 519
 520
 521
 522
 523
 524
 525
 526
 527
 528
 529
 530
 531
 532
 533
 534
 535
 536
 537
 538
 539
 540
 541
 542
 543
 544
 545
 546
 547
 548
 549
 550
 551
 552
 553
 554
 555
 556
 557
 558
 559
 560
 561
 562
 563
 564
 565
 566
 567
 568
 569
 570
 571
 572
 573
 574
 575
 576
 577
 578
 579
 580
 581
 582
 583
 584
 585
 586
 587
 588
 589
 590
 591
 592
 593
 594
 595
 596
 597
 598
 599
 600
 601
 602
 603
 604
 605
 606
 607
 608
 609
 610
 611
 612
 613
 614
 615
 616
 617
 618
 619
 620
 621
 622
 623
 624
 625
 626
 627
 628
 629
 630
 631
 632
 633
 634
 635
 636
 637
 638
 639
 640
 641
 642
 643
 644
 645
 646
 647
 648
 649
 650
 651
 652
 653
 654
 655
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
2028
2029
2030
2031
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
2100
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
2115
2116
2117
2118
2119
2120
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
2143
2144
2145
2146
2147
2148
2149
2150
2151
2152
2153
2154
2155
2156
2157
2158
2159
2160
2161
2162
2163
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
2178
2179
2180
2181
2182
2183
2184
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
2204
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
2226
2227
2228
2229
2230
2231
2232
2233
2234
2235
2236
2237
2238
2239
2240
2241
2242
2243
2244
2245
2246
2247
2248
2249
2250
2251
2252
2253
2254
2255
2256
2257
2258
2259
2260
2261
2262
2263
2264
2265
2266
2267
2268
2269
2270
2271
2272
2273
2274
2275
2276
2277
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
2291
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
2308
2309
2310
2311
2312
2313
2314
2315
2316
2317
2318
2319
2320
2321
2322
2323
2324
2325
2326
2327
2328
2329
2330
2331
2332
2333
2334
2335
2336
2337
2338
2339
2340
2341
2342
2343
2344
2345
2346
2347
2348
2349
2350
2351
2352
2353
2354
2355
2356
2357
2358
2359
2360
2361
2362
2363
2364
2365
2366
2367
2368
2369
2370
2371
2372
2373
2374
2375
2376
2377
2378
2379
2380
2381
2382
2383
2384
2385
2386
2387
2388
2389
2390
2391
2392
2393
2394
2395
2396
2397
2398
2399
2400
2401
2402
2403
2404
2405
2406
2407
2408
2409
2410
2411
2412
2413
2414
2415
2416
2417
2418
2419
2420
2421
2422
2423
2424
2425
2426
2427
2428
2429
2430
2431
2432
2433
2434
2435
2436
2437
2438
2439
2440
2441
2442
2443
2444
2445
2446
2447
2448
2449
2450
2451
2452
2453
2454
2455
2456
2457
2458
2459
2460
2461
2462
2463
2464
2465
2466
2467
2468
2469
2470
2471
2472
2473
2474
2475
2476
2477
2478
2479
2480
2481
2482
2483
2484
2485
2486
2487
2488
2489
2490
2491
2492
2493
2494
2495
2496
2497
2498
2499
2500
2501
2502
2503
2504
2505
2506
2507
2508
2509
2510
2511
2512
2513
2514
2515
2516
2517
2518
2519
2520
2521
2522
2523
2524
2525
2526
2527
2528
2529
2530
2531
2532
2533
2534
2535
2536
2537
2538
2539
2540
2541
2542
2543
2544
2545
2546
2547
2548
2549
2550
2551
2552
2553
2554
2555
2556
2557
2558
2559
2560
2561
2562
2563
2564
2565
2566
2567
2568
2569
2570
2571
2572
2573
2574
2575
2576
2577
2578
2579
2580
2581
2582
2583
2584
2585
2586
2587
2588
2589
2590
2591
2592
2593
2594
2595
2596
2597
2598
2599
2600
2601
2602
2603
2604
2605
2606
2607
2608
2609
2610
2611
2612
2613
2614
2615
2616
2617
2618
2619
2620
2621
2622
2623
2624
2625
2626
2627
2628
2629
2630
2631
2632
2633
2634
2635
2636
2637
2638
2639
2640
2641
2642
2643
2644
2645
2646
2647
2648
2649
2650
2651
2652
2653
2654
2655
2656
2657
2658
2659
2660
2661
2662
2663
2664
2665
2666
2667
2668
2669
2670
2671
2672
2673
2674
2675
2676
2677
2678
2679
2680
2681
2682
2683
2684
2685
2686
2687
2688
2689
2690
2691
2692
2693
2694
2695
2696
2697
2698
2699
2700
2701
2702
2703
2704
2705
2706
2707
2708
2709
2710
2711
2712
2713
2714
2715
2716
2717
2718
2719
2720
2721
2722
2723
2724
2725
2726
2727
2728
2729
2730
2731
2732
2733
2734
2735
2736
2737
2738
2739
2740
2741
2742
2743
2744
2745
2746
2747
2748
2749
2750
2751
2752
2753
2754
2755
2756
2757
2758
2759
2760
2761
2762
2763
2764
2765
2766
2767
2768
2769
2770
2771
2772
2773
2774
2775
2776
2777
2778
2779
2780
2781
2782
2783
2784
2785
2786
2787
2788
2789
2790
2791
2792
2793
2794
2795
2796
2797
2798
2799
2800
2801
2802
2803
2804
2805
2806
2807
2808
2809
2810
2811
2812
2813
2814
2815
2816
2817
2818
2819
2820
2821
2822
2823
2824
2825
2826
2827
2828
2829
2830
2831
2832
2833
2834
2835
2836
2837
2838
2839
2840
2841
2842
2843
2844
2845
2846
2847
2848
2849
2850
2851
2852
2853
2854
2855
2856
2857
2858
2859
2860
2861
2862
2863
2864
2865
2866
2867
2868
2869
2870
2871
2872
2873
2874
2875
2876
2877
2878
2879
2880
2881
2882
2883
2884
2885
2886
2887
2888
2889
2890
2891
2892
2893
2894
2895
2896
2897
2898
2899
2900
2901
2902
2903
2904
2905
2906
2907
2908
2909
2910
2911
2912
2913
2914
2915
2916
2917
2918
2919
2920
2921
2922
2923
2924
2925
2926
2927
2928
2929
2930
2931
2932
2933
2934
2935
2936
2937
2938
2939
2940
2941
2942
2943
2944
2945
2946
2947
2948
2949
2950
2951
2952
2953
2954
2955
2956
2957
2958
2959
2960
2961
2962
2963
2964
2965
2966
2967
2968
2969
2970
2971
2972
2973
2974
2975
2976
2977
2978
2979
2980
2981
2982
2983
2984
2985
2986
2987
2988
2989
2990
2991
2992
2993
2994
2995
2996
2997
2998
2999
3000
3001
3002
3003
3004
3005
3006
3007
3008
3009
3010
3011
3012
3013
3014
3015
3016
3017
3018
3019
3020
3021
3022
3023
3024
3025
3026
3027
3028
3029
3030
3031
3032
3033
3034
3035
3036
3037
3038
3039
3040
3041
3042
3043
3044
3045
3046
3047
3048
3049
3050
3051
3052
3053
3054
3055
3056
3057
3058
3059
3060
3061
3062
3063
3064
3065
3066
3067
3068
3069
3070
3071
3072
3073
3074
3075
3076
3077
3078
3079
3080
3081
3082
3083
3084
3085
3086
3087
3088
3089
3090
3091
3092
3093
3094
3095
3096
3097
3098
3099
3100
3101
3102
3103
3104
3105
3106
3107
3108
3109
3110
3111
3112
3113
3114
3115
3116
3117
3118
3119
3120
3121
3122
3123
3124
3125
3126
3127
3128
3129
3130
3131
3132
3133
3134
3135
3136
3137
3138
3139
3140
3141
3142
3143
3144
3145
3146
3147
3148
3149
3150
3151
3152
3153
3154
3155
3156
3157
3158
3159
3160
3161
3162
3163
3164
3165
3166
3167
3168
3169
3170
3171
3172
3173
3174
3175
3176
3177
3178
3179
3180
3181
3182
3183
3184
3185
3186
3187
3188
3189
3190
3191
3192
3193
3194
3195
3196
3197
3198
3199
3200
3201
3202
3203
3204
3205
3206
3207
3208
3209
3210
3211
3212
3213
3214
3215
3216
3217
3218
3219
3220
3221
3222
3223
3224
3225
3226
3227
3228
3229
3230
3231
3232
3233
3234
3235
3236
3237
3238
3239
3240
3241
3242
3243
3244
3245
3246
3247
3248
3249
3250
3251
3252
3253
3254
3255
3256
3257
3258
3259
3260
3261
3262
3263
3264
3265
3266
3267
3268
3269
3270
3271
3272
3273
3274
3275
3276
3277
3278
3279
3280
3281
3282
3283
3284
3285
3286
3287
3288
3289
3290
3291
3292
3293
3294
3295
3296
3297
3298
3299
3300
3301
3302
3303
3304
3305
3306
3307
3308
3309
3310
3311
3312
3313
3314
3315
3316
3317
3318
3319
3320
3321
3322
3323
3324
3325
3326
3327
3328
3329
3330
3331
3332
3333
3334
3335
3336
3337
3338
3339
3340
3341
3342
3343
3344
3345
3346
3347
3348
3349
3350
3351
3352
3353
3354
3355
3356
3357
3358
3359
3360
3361
3362
3363
3364
3365
3366
3367
3368
3369
3370
3371
3372
3373
3374
3375
3376
3377
3378
3379
3380
3381
3382
3383
3384
3385
3386
3387
3388
3389
3390
3391
3392
3393
3394
3395
3396
3397
3398
3399
3400
3401
3402
3403
3404
3405
3406
3407
3408
3409
3410
3411
3412
3413
3414
3415
3416
3417
3418
3419
3420
3421
3422
3423
3424
3425
3426
3427
3428
3429
3430
3431
3432
3433
3434
3435
3436
3437
3438
3439
3440
3441
3442
3443
3444
3445
3446
3447
3448
3449
3450
3451
3452
3453
3454
3455
3456
3457
3458
3459
3460
3461
3462
3463
3464
3465
3466
3467
3468
3469
3470
3471
3472
3473
3474
3475
3476
3477
3478
3479
3480
3481
3482
3483
3484
3485
3486
3487
3488
3489
3490
3491
3492
3493
3494
3495
3496
3497
3498
3499
3500
3501
3502
3503
3504
3505
3506
3507
3508
3509
3510
3511
3512
3513
3514
3515
3516
3517
3518
3519
3520
3521
3522
3523
3524
3525
3526
3527
3528
3529
3530
3531
3532
3533
3534
3535
3536
3537
3538
3539
3540
3541
3542
3543
3544
3545
3546
3547
3548
3549
3550
3551
3552
3553
3554
3555
3556
3557
3558
3559
3560
3561
3562
3563
3564
3565
3566
3567
3568
3569
3570
3571
3572
3573
3574
3575
3576
3577
3578
3579
3580
3581
3582
3583
3584
3585
3586
3587
3588
3589
3590
3591
3592
3593
3594
3595
3596
3597
3598
3599
3600
3601
3602
3603
3604
3605
3606
3607
3608
3609
3610
3611
3612
3613
3614
3615
3616
3617
3618
3619
3620
3621
3622
3623
3624
3625
3626
3627
3628
3629
3630
3631
3632
3633
3634
3635
3636
3637
3638
3639
3640
3641
3642
3643
3644
3645
3646
3647
3648
3649
3650
3651
3652
3653
3654
3655
3656
3657
3658
3659
3660
3661
3662
3663
3664
3665
3666
3667
3668
3669
3670
3671
3672
3673
3674
3675
3676
3677
3678
3679
3680
3681
3682
3683
3684
3685
3686
3687
3688
3689
3690
3691
3692
3693
3694
3695
3696
3697
3698
3699
3700
3701
3702
3703
3704
3705
3706
3707
3708
3709
3710
3711
3712
3713
3714
3715
3716
3717
3718
3719
3720
3721
3722
3723
3724
3725
3726
3727
3728
3729
3730
3731
3732
3733
3734
3735
3736
3737
3738
3739
3740
3741
3742
3743
3744
3745
3746
3747
3748
3749
3750
3751
3752
3753
3754
3755
3756
3757
3758
3759
3760
3761
3762
3763
3764
3765
3766
3767
3768
3769
3770
3771
3772
3773
3774
3775
3776
3777
3778
3779
3780
3781
3782
3783
3784
3785
3786
3787
3788
3789
3790
3791
3792
3793
3794
3795
3796
3797
3798
3799
3800
3801
3802
3803
3804
3805
3806
3807
3808
3809
3810
3811
3812
3813
3814
3815
3816
3817
3818
3819
3820
3821
3822
3823
3824
3825
3826
3827
3828
3829
3830
3831
3832
3833
3834
3835
3836
3837
3838
3839
3840
3841
3842
3843
3844
3845
3846
3847
3848
3849
3850
3851
3852
3853
3854
3855
3856
3857
3858
3859
3860
3861
3862
3863
3864
3865
3866
3867
3868
3869
3870
3871
3872
3873
3874
3875
3876
3877
3878
3879
3880
3881
3882
3883
3884
3885
3886
3887
3888
3889
3890
3891
3892
3893
3894
3895
3896
3897
3898
3899
3900
3901
3902
3903
3904
3905
3906
3907
3908
3909
3910
3911
3912
3913
3914
3915
3916
3917
3918
3919
3920
3921
3922
3923
3924
3925
3926
3927
3928
3929
3930
3931
3932
3933
3934
3935
3936
3937
3938
3939
3940
3941
3942
3943
3944
3945
3946
3947
3948
3949
3950
3951
3952
3953
3954
3955
3956
3957
3958
3959
3960
3961
3962
3963
3964
3965
3966
3967
3968
3969
3970
3971
3972
3973
3974
3975
3976
3977
3978
3979
3980
3981
3982
3983
3984
3985
3986
3987
3988
3989
3990
3991
3992
3993
3994
3995
3996
3997
3998
3999
4000
4001
4002
4003
4004
4005
4006
4007
4008
4009
4010
4011
4012
4013
4014
4015
4016
4017
4018
4019
4020
4021
4022
4023
4024
4025
4026
4027
4028
4029
4030
4031
4032
4033
4034
4035
4036
4037
4038
4039
4040
4041
4042
4043
4044
4045
4046
4047
4048
4049
4050
4051
4052
4053
4054
4055
4056
4057
4058
4059
4060
4061
4062
4063
4064
4065
4066
4067
4068
4069
4070
4071
4072
4073
4074
4075
4076
4077
4078
4079
4080
4081
4082
4083
4084
4085
4086
4087
4088
4089
4090
4091
4092
4093
4094
4095
4096
4097
4098
4099
4100
4101
4102
4103
4104
4105
4106
4107
4108
4109
4110
4111
4112
4113
4114
4115
4116
4117
4118
4119
4120
4121
4122
4123
4124
4125
4126
4127
4128
4129
4130
4131
4132
4133
4134
4135
4136
4137
4138
4139
4140
4141
4142
4143
4144
4145
4146
4147
4148
4149
4150
4151
4152
4153
4154
4155
4156
4157
4158
4159
4160
4161
4162
4163
4164
4165
4166
4167
4168
4169
4170
4171
4172
4173
4174
4175
4176
4177
4178
4179
4180
4181
4182
4183
4184
4185
4186
4187
4188
4189
4190
4191
4192
4193
4194
4195
4196
4197
4198
4199
4200
4201
4202
4203
4204
4205
4206
4207
4208
4209
4210
4211
4212
4213
4214
4215
4216
4217
4218
4219
4220
4221
4222
4223
4224
4225
4226
4227
4228
4229
4230
4231
4232
4233
4234
4235
4236
4237
4238
4239
4240
4241
4242
4243
4244
4245
4246
4247
4248
4249
4250
4251
4252
4253
4254
4255
4256
4257
4258
4259
4260
4261
4262
4263
4264
4265
4266
4267
4268
4269
4270
4271
4272
4273
4274
4275
4276
4277
4278
4279
4280
4281
4282
4283
4284
4285
4286
4287
4288
4289
4290
4291
4292
4293
4294
4295
4296
4297
4298
4299
4300
4301
4302
4303
4304
4305
4306
4307
4308
4309
4310
4311
4312
4313
4314
4315
4316
4317
4318
4319
4320
4321
4322
4323
4324
4325
4326
4327
4328
4329
4330
4331
4332
4333
4334
4335
4336
4337
4338
4339
4340
4341
4342
4343
4344
4345
4346
4347
4348
4349
4350
4351
4352
4353
4354
4355
4356
4357
4358
4359
4360
4361
4362
4363
4364
4365
4366
4367
4368
4369
4370
4371
4372
4373
4374
4375
4376
4377
4378
4379
4380
4381
4382
4383
4384
4385
4386
4387
4388
4389
4390
4391
4392
4393
4394
4395
4396
4397
4398
4399
4400
4401
4402
4403
4404
4405
4406
4407
4408
4409
4410
4411
4412
4413
4414
4415
4416
4417
4418
4419
4420
4421
4422
4423
4424
4425
4426
4427
4428
4429
4430
4431
4432
4433
4434
4435
4436
4437
4438
4439
4440
4441
4442
4443
4444
4445
4446
4447
4448
4449
4450
4451
4452
4453
4454
4455
4456
4457
4458
4459
4460
4461
4462
4463
4464
4465
4466
4467
4468
4469
4470
4471
4472
4473
4474
4475
4476
4477
4478
4479
4480
4481
4482
4483
4484
4485
4486
4487
4488
4489
4490
4491
4492
4493
4494
4495
4496
4497
4498
4499
4500
4501
4502
4503
4504
4505
4506
4507
4508
4509
4510
4511
4512
4513
4514
4515
4516
4517
4518
4519
4520
4521
4522
4523
4524
4525
4526
4527
4528
4529
4530
4531
4532
4533
4534
4535
4536
4537
4538
4539
4540
4541
4542
4543
4544
4545
4546
4547
4548
4549
4550
4551
4552
4553
4554
4555
4556
4557
4558
4559
4560
4561
4562
4563
4564
4565
4566
4567
4568
4569
4570
4571
4572
4573
4574
4575
4576
4577
4578
4579
4580
4581
4582
4583
4584
4585
4586
4587
4588
4589
4590
4591
4592
4593
4594
4595
4596
4597
4598
4599
4600
4601
4602
4603
4604
4605
4606
4607
4608
4609
4610
4611
4612
4613
4614
4615
4616
4617
4618
4619
4620
4621
4622
4623
4624
4625
4626
4627
4628
4629
4630
4631
4632
4633
4634
4635
4636
4637
4638
4639
4640
4641
4642
4643
4644
4645
4646
4647
4648
4649
4650
4651
4652
4653
4654
4655
4656
4657
4658
4659
4660
4661
4662
4663
4664
4665
4666
4667
4668
4669
4670
4671
4672
4673
4674
4675
4676
4677
4678
4679
4680
4681
4682
4683
4684
4685
4686
4687
4688
4689
4690
4691
4692
4693
4694
4695
4696
4697
4698
4699
4700
4701
4702
4703
4704
4705
4706
4707
4708
4709
4710
4711
4712
4713
4714
4715
4716
4717
4718
4719
4720
4721
4722
4723
4724
4725
4726
4727
4728
4729
4730
4731
4732
4733
4734
4735
4736
4737
4738
4739
4740
4741
4742
4743
4744
4745
4746
4747
4748
4749
4750
4751
4752
4753
4754
4755
4756
4757
4758
4759
4760
4761
4762
4763
4764
4765
4766
4767
4768
4769
4770
4771
4772
4773
4774
4775
4776
4777
4778
4779
4780
4781
4782
4783
4784
4785
4786
4787
4788
4789
4790
4791
4792
4793
4794
4795
4796
4797
4798
4799
4800
4801
4802
4803
4804
4805
4806
4807
4808
4809
4810
4811
4812
4813
4814
4815
4816
4817
4818
4819
4820
4821
4822
4823
4824
4825
4826
4827
4828
4829
4830
4831
4832
4833
4834
4835
4836
4837
4838
4839
4840
4841
4842
4843
4844
4845
4846
4847
4848
4849
4850
4851
4852
4853
4854
4855
4856
4857
4858
4859
4860
4861
4862
4863
4864
4865
4866
4867
4868
4869
4870
4871
4872
4873
4874
4875
4876
4877
4878
4879
4880
4881
4882
4883
4884
4885
4886
4887
4888
4889
4890
4891
4892
4893
4894
4895
4896
4897
4898
4899
4900
4901
4902
4903
4904
class ToolUseLoop:
    """Async tool-use conversation loop.

    Each instance owns one conversation with one LLM. Sub-agents are
    created by spawning new ``ToolUseLoop`` instances via the
    ``spawn_agent`` internal tool.
    """

    def __init__(
        self,
        *,
        agent_context: AgentContext,
        tool_registry: ToolRegistry,
        permission_policy: PermissionPolicy,
        approval_callback: Callable[[ActionStep], bool] | None = None,
        hook_manager: HookManager,
        safety_plane: SafetyPlane | None = None,
        project_instructions: str | None = None,
        user_instructions: str | None = None,
        skill_instructions: str | None = None,
        skill_registry: Any = None,
        agent_registry: Any = None,
        session_tool_registry: SessionToolRegistry | None = None,
        allowed_tools: list[str] | None = None,
        denied_tools: list[str] | None = None,
        strict_tool_scope: bool = False,
        cwd: str | None = None,
        session_id: str | None = None,
        session_capabilities: tuple[str, ...] = (),
        extra_session_tools: list[SessionTool] | None = None,
        enable_skills: bool = True,
        project_autoselect: bool = False,
        project_catalog: ProjectCatalog | None = None,
        session_context_reader: Callable[[], dict[str, object]] | None = None,
        contract: DelegationContract | None = None,
        verification: CommandVerification | None = None,
        verifier_runner: VerifierRunner | None = None,
        watchdog_sleeper: Callable[[float], Awaitable[None]] | None = None,
    ) -> None:
        """Initialize the tool-use loop.

        Args:
            agent_context: Required — carries model, cancel, logger, registry.
            tool_registry: Registered tools available to this agent.
            permission_policy: Permission rules for tool execution.
            approval_callback: Optional callback for ASK decisions (None for sub-agents).
            hook_manager: Lifecycle hooks.
            safety_plane: Operator-owned tool-call gate and session observer,
                built once by the orchestrator from config + ``.mewbo/`` and
                forwarded by value to every child loop. ``None`` (the default)
                when the deployment-wide switch is off — every safety call
                site is a single ``is not None`` check away from a no-op.
            project_instructions: CLAUDE.md / AGENTS.md content discovered at session start.
            user_instructions: Operator-authored custom instructions, ALREADY RENDERED
                by ``Orchestrator`` from the stored template (``system_instructions/``).
                A plain string — the loop does no Jinja and no store work. Lives on the
                instance, not on a per-call arg, so it survives the in-place system-prompt
                re-render that a model escalation performs.
            skill_instructions: Pre-rendered skill body (from user /skill invocation).
            skill_registry: SkillRegistry for auto-invocation catalog + activate_skill handling.
            agent_registry: AgentRegistry for agent type catalog + spawn_agent type lookup.
            session_tool_registry: Registry of plugin-contributed session-tool
                factories.  Each matching factory (filtered by ``allowed_tools``)
                is instantiated for this agent and added to ``self._session_tools``.
            allowed_tools: The agent's allowlist used to filter which session
                tools the plugin registry should build for this agent.  ``None``
                means "no plugin session tools" (root agents get only the
                built-in ``ExitPlanModeTool``).
            denied_tools: Session-tool ids withheld regardless of which gate in
                ``SessionToolRegistry.build_for`` would otherwise admit them —
                the unconditional auto-surface and the capability auto-surface
                included, and it beats a NAMED ``allowed_tools`` entry too.
                Deliberately NOT three-state like ``allowed_tools``: deny is
                purely subtractive, so ``None`` and ``[]`` are the same "nothing
                denied" set. ``None`` (the default) changes nothing for every
                existing session.
            strict_tool_scope: Whether ``allowed_tools`` is AUTHORITATIVE for this
                agent. ``True`` (spawned leaf sub-agents, wiki-qa/search runs) —
                the allowlist is the whole tool scope, so it also gates
                ``spawn_agent``. ``False`` (the FE default) — ``allowed_tools`` is
                only a PERMISSIVE ceiling over MCP tools (``context.mcp_tools``);
                built-ins and the internal ``spawn_agent`` are NOT scoped by it,
                mirroring the orchestrator's permissive ``filter_specs`` branch.
            cwd: Working directory for this agent (project root).
            session_id: Session identifier — used for plan-mode path scoping.
            session_capabilities: Client-advertised capability tuple from the
                ``X-Mewbo-Capabilities`` header (persisted on the session
                context event). Used to filter capability-gated agents and
                skills out of the system-prompt catalogs and ``activate_skill``
                / ``spawn_agent`` lookups.
            extra_session_tools: Caller-injected ``SessionTool`` instances
                appended to ``self._session_tools`` without a plugin manifest
                (e.g. the structured-response ``emit_result`` tool).
            enable_skills: When ``False``, the auto-invocable ``activate_skill``
                schema is never injected even if the registry holds skills — a
                headless product drive (search/wiki) can opt out so it doesn't
                burn its first step activating a host ``~/.claude`` skill it
                never intended to expose. Default ``True`` (unchanged behavior).
            project_autoselect: When ``True`` (and this is the ROOT agent), bind
                ``list_projects`` / ``switch_project`` so the model can choose
                which workspace to work in. Default ``False`` binds neither.
            project_catalog: The catalog those two tools read and resolve
                through, built by the caller (``Orchestrator``) because
                assembling it needs the config plus the project/repository
                stores. ``None`` disables the pair regardless of the flag.
            session_context_reader: Reads back the session's CURRENT effective
                context — the most-recent ``context`` payload — so a project
                switch can carry it forward instead of replacing it. Injected
                because this loop has no transcript access at all: its only
                context channel is the write-only ``event_logger``, and the
                truncated ``recent_events`` window it does hold would produce a
                merge that silently drops anything older than the window.
                ``None`` degrades to writing the two switch keys alone.
            contract: The spawner's declared ``DelegationContract`` for THIS
                agent — ``None``/disabled leaves the child unbounded.
                Checked in the run loop's budget block
                IN ADDITION TO (never instead of) the shared session budget.
            verification: The spawner's declared ground-truth completion check
                for THIS agent — ``None`` leaves completion ungated. Only
                RUN when the two-gate ``self._verification_active`` holds
                (master switch on AND a spec AND an execute/all capability
                mode); otherwise it is carried but inert.
            verifier_runner: Injected ``VerifierRunner`` (defaults to
                ``CommandVerifierRunner``). A test passes a recording fake so
                the gate is exercised without a real subprocess.
            watchdog_sleeper: Injected wait between watchdog sweeps (defaults
                to ``asyncio.sleep``). A test drives one sweep by passing a
                fake, so it never has to patch the stdlib module every other
                coroutine in the process also awaits through.
        """
        self._ctx = agent_context
        # The run's stop signal, wrapping the SAME predicate the between-turns
        # check has always read. It exists so an in-flight model call or tool
        # execution can be raced against a stop instead of outliving it — see
        # ``cancellation.py`` for why polling alone bounded the stop by a turn.
        self._cancellation = CancellationSignal(agent_context.should_cancel)
        self._contract = contract
        # The model whose per-model prompt overrides + tool variant are ACTIVE.
        # Starts at the configured primary; the fallback ladder promotes it
        # to the escalated model on a sticky switch (see ``_apply_model_escalation``)
        # so the heal becomes behavioural, not just a model swap.
        self._active_model = agent_context.model_name
        self._enable_skills = enable_skills
        self._tool_registry = tool_registry
        self._permission_policy = permission_policy
        self._approval_callback = approval_callback
        self._hook_manager = hook_manager
        # Forwarded BY VALUE from the orchestrator that built the root loop —
        # never re-read from config or disk here. A model switch mid-session
        # re-renders ``messages[0]`` at the safe turn boundary but never touches
        # this reference, and ``_build_child_loop`` (spawn_agent.py) passes the
        # SAME object to every descendant, so an agent editing or deleting the
        # on-disk document mid-session changes nothing about the plane judging
        # its own session.
        self._safety_plane = safety_plane
        self._safety_started_at: float | None = None
        self._safety_tool_calls = 0
        self._safety_write_calls = 0
        self._project_instructions = project_instructions
        self._user_instructions = user_instructions
        self._skill_instructions = skill_instructions
        self._skill_registry = skill_registry
        self._agent_registry = agent_registry
        self._session_tool_registry = session_tool_registry
        self._cwd = cwd
        self._session_id = session_id
        self._session_capabilities = session_capabilities
        # Retained so the tool ceiling can reach the tools this loop injects
        # OUTSIDE ``filter_specs`` — see :meth:`_loop_injected_admitted`.
        self._allowed_tools = allowed_tools
        self._denied_tools = denied_tools
        self._strict_tool_scope = strict_tool_scope
        self._project_catalog = project_catalog
        self._session_context_reader = session_context_reader
        # The catalog key of the project this loop is currently in. ``None``
        # until the first switch — construction supplies a directory, never a
        # key — so it reports "no previous project", never a guessed one.
        self._project_key: str | None = None

        # Deferred-tool partitioning state. ``run()`` stamps all four from the
        # specs it is handed; they are initialized here because
        # ``rebind_workspace`` re-derives the bound set from them and would
        # otherwise depend on the attribute-creation order inside ``run`` — a
        # dependency that fails as an ``AttributeError`` swallowed into a
        # "could not switch" refusal, i.e. a workspace half moved.
        self._tool_specs_full: list[ToolSpec] = []
        self._tool_search_enabled: bool = False
        self._deferred_ids: set[str] = set()
        self._last_active_ids: set[str] = set()
        # A workspace switch performed mid-turn, applied at the NEXT turn
        # boundary (the ``_requested_switch`` pattern): the tool mutates the
        # loop's own state immediately, but the transcript tail stays untouched
        # until the turn ends. ``None`` whenever no switch is pending.
        self._pending_workspace_bind: tuple[list[dict[str, Any]], Any] | None = None

        # Filesystem-containment firebreak. Built ONCE here from
        # three inputs: the enforcement kill-switch (``agent.workspace_
        # enforcement``, staged OFF by default), this agent's narrowed
        # ``workspace_mode``, and the workspace ``cwd``. It is a non-None
        # ``WorkspaceContainment`` ONLY when all three admit containment
        # (enforcement on AND a restrictive tier AND a real cwd) — so with the
        # flag off, a full_access tier, or no cwd, ``self._containment is None``
        # and every path resolves to the plain tenant union.
        # ``self._containment is not None`` is therefore the single "containment
        # active" predicate the loop's root-injection + tool-execution seams read.
        self._containment: WorkspaceContainment | None = self._build_containment(cwd)

        # Verifier-gated completion. Two-gate arming computed ONCE (mirrors the
        # write-progress signal): a ground-truth check only gates an agent that
        # (a) has a spec, (b) runs under the master switch, and (c) could
        # plausibly ACT — capability_mode ∈ {execute, all}. A read-only child,
        # the disabled default, or the staged-off switch leaves the gate inert,
        # so every natural completion is accepted unchecked. The runner
        # is injected (default ``CommandVerifierRunner``) so a test drives the
        # gate with a recording fake and never spawns a real subprocess.
        # The watchdog's poll wait, injected as a collaborator. A caller that
        # must drive a sweep without waiting for one replaces THIS, rather
        # than ``asyncio.sleep`` on the stdlib module — that object is shared
        # by every module and every running loop in the process, so patching
        # it reaches coroutines this loop has nothing to do with.
        self._watchdog_sleep = watchdog_sleeper or asyncio.sleep
        self._verification = verification
        self._verifier_runner: VerifierRunner = verifier_runner or CommandVerifierRunner()
        self._verification_active = verification is not None and CommandVerification.gate_active(
            enabled=bool(get_config_value("agent", "verification_enabled", default=False)),
            capability_mode=agent_context.capability_mode,
        )
        # Latches for the run: whether a check has already passed (never re-run
        # once green) and how many failed re-drives remain.
        self._verify_passed = False
        self._verify_retries_left = int(
            get_config_value("agent", "verification_max_retries", default=2)
        )

        # Dedup cache for read_file: prevents redundant reads when the
        # same file + range hasn't changed on disk (mtime check).
        self._file_read_cache: dict[str, _CachedFileRead] = {}

        # Takes stale images out of history when the list is compacted. A
        # plain field rather than a knob — nothing has asked to tune it.
        self._image_history = ImageHistoryStrip()

        # Plan-mode state (mutable across the loop's lifetime).
        self._current_mode: str = "act"
        # Authoritative token count from the most recent LLM response's
        # usage_metadata.input_tokens. Zero until the first call lands.
        self._last_input_tokens: int = 0

        # In-flight LLM call, for the liveness leg. ``None`` whenever no call is
        # outstanding; a monotonic timestamp while one is. The retry strategy
        # already bounds each attempt with ``asyncio.wait_for``, but that bound
        # can only fire if the awaited coroutine reaches a cancellation point —
        # a provider read wedged below the event loop never does, and one such
        # call sat silent for 24 minutes having emitted ``llm_call_start`` and
        # no end. Nothing noticed it live: the sweepers that would have run once
        # each, at process boot.
        self._llm_call_started_at: float | None = None
        self._llm_call_step: int = 0

        # How many tool schemas the last ``_bind_model`` actually bound, stamped
        # onto ``llm_call_start`` so the bound surface is queryable after the
        # fact. Stored by the binder rather than recomputed at the emit site:
        # a second derivation of "what is bound" is free to drift from the real
        # one, which is the very failure this records.
        self._bound_tool_count: int = 0

        # Error-visibility seam: a bounded, factual note of this run's LLM
        # retry/fallback events, injected as its OWN system-prompt slot so the
        # model can see its own retries. ``_active_resilience_note``
        # tracks the text last baked into ``messages[0]`` so the loop re-renders
        # the prompt only when the note actually changes (a clean run never
        # pays; a healing run re-renders at most once per new event).
        self._resilience_note = ResilienceNote()
        self._active_resilience_note: str = ""

        # Self-steering routing state. The live per-run ``RetryStrategy`` (set in
        # ``run``) is what the ``model_control`` tool reuses for its switch
        # budget + cooldown; ``_requested_switch`` is a model the tool asked to
        # switch to, applied at the NEXT turn boundary (like a sticky fallback,
        # transcript tail untouched); ``_live_messages`` lets the continuity-lock
        # guardrail inspect the in-flight transcript.
        self._retry_strategy: RetryStrategy | None = None
        self._requested_switch: str | None = None
        self._live_messages: list[BaseMessage] | None = None

        # Create SpawnAgentTool when this agent can spawn children — gated on
        # BOTH depth (``can_spawn``) AND tool scope. spawn_agent is injected
        # here rather than through ``filter_specs``, so an explicit ``tools:``
        # allowlist that omits it would otherwise be silently bypassed: a leaf
        # agent scoped to build-and-submit (the st-widget-builder) could still
        # delegate, and misread its own errors as "delegate to a scoped agent",
        # recursing into copies of itself. An ABSENT
        # allowlist (``None``) stays unrestricted (root / ad-hoc spawns); an
        # EMPTY one grants nothing, delegation included. That distinction is
        # load-bearing, not pedantry: a role-bounded viewer's composed allowlist
        # omits the spawn family precisely to disable delegation, and can compose
        # down to empty — reading empty as "unrestricted" would hand delegation
        # back to the principal the ceiling exists to deny.
        #
        # The allowlist gate applies ONLY when the scope is STRICT (a spawned
        # leaf / wiki-qa / search run — where ``allowed_tools`` is the whole
        # authoritative tool scope). Under a PERMISSIVE scope (the FE default),
        # ``allowed_tools`` is ``context.mcp_tools`` — a ceiling over MCP tools
        # only, which never lists the internal spawn_agent; built-ins stay (see
        # the orchestrator's permissive ``filter_specs`` branch), so spawn_agent
        # must stay too. Without this carve-out every console/Aura session that
        # advertised MCP tools had root delegation silently disabled (session
        # 04ea546e…: the st-widget-builder skill mandates spawn_agent, which was
        # unreachable, so the agent violated the skill and built the widget itself).
        spawn_in_scope = (
            not strict_tool_scope
            or allowed_tools is None
            or bool({"spawn_agent", "spawn_agents"} & set(allowed_tools))
        ) and not self._ctx.atomic  # Atomic is a hard firebreak
        self._spawn_agent_tool: Any = None
        if agent_context.can_spawn and spawn_in_scope:
            from mewbo_core.agents.spawn_agent import SpawnAgentTool

            # Sub-agents inherit parent's approval
            # policy so they can execute write/edit/shell tools.
            self._spawn_agent_tool = SpawnAgentTool(
                agent_context=agent_context,
                tool_registry=tool_registry,
                permission_policy=permission_policy,
                approval_callback=approval_callback,
                hook_manager=hook_manager,
                project_instructions=project_instructions,
                user_instructions=user_instructions,
                cwd=cwd,
                agent_registry=agent_registry,
                session_tool_registry=session_tool_registry,
                session_capabilities=session_capabilities,
                enable_skills=enable_skills,
            )

        # Assemble session tools — per-agent stateful handlers that carry
        # their own schema, dispatch, and run-termination flag. The core's
        # built-in ``ExitPlanModeTool`` is always attached to root agents
        # with a session id; plugin-contributed tools are selected by EITHER
        # the agent's ``allowed_tools`` allowlist OR a capability gate (a
        # factory whose ``requires_capabilities`` ⊆ ``session_capabilities``),
        # so a runtime-granted capability surfaces its tools to the
        # root agent without the client listing them explicitly.
        self._session_tools: list[SessionTool] = []
        if agent_context.depth == 0 and session_id is not None:
            # Deliberately NOT ceiling-checked: ``exit_plan_mode`` is the only
            # way out of plan mode, so withholding it from a strict scope that
            # failed to name it would leave the run with no exit at all. It is a
            # structural terminator, never a surface an agent wanders into.
            self._session_tools.append(
                ExitPlanModeTool(
                    session_id=session_id,
                    event_logger=agent_context.event_logger,
                )
            )
            # Authoritative live todos (act mode): terminal-free, re-emits the
            # FULL statused list as ONE ``todos`` event on each call. Attached
            # inline (not via the plugin factory) so it bypasses the plugin
            # factory's allowlist gate — hence the explicit ceiling check here,
            # or an AgentDef's authoritative ``tools:`` under-states what its
            # agent actually holds.
            if self._loop_injected_admitted("update_todos"):
                self._session_tools.append(
                    UpdateTodosTool(
                        session_id=session_id,
                        event_logger=agent_context.event_logger,
                        agent_id=agent_context.agent_id,
                    )
                )
            # Project selection (opt-in). Root-only is deliberate and enforced
            # by the enclosing ``depth == 0`` branch: a child is HANDED a
            # workspace by its parent, and letting it re-scope the session would
            # move the ground under every sibling still working in the old
            # directory. Same inline attachment as update_todos, hence the same
            # explicit ceiling check — these bypass the plugin factory's
            # allowlist gate, so an AgentDef's authoritative ``tools:`` would
            # otherwise under-state what its agent holds.
            if project_autoselect and project_catalog is not None:
                from mewbo_core.workspaces.project_switch import (
                    LIST_PROJECTS_TOOL_ID,
                    SWITCH_PROJECT_TOOL_ID,
                    ListProjectsTool,
                    SwitchProjectTool,
                )

                if self._loop_injected_admitted(LIST_PROJECTS_TOOL_ID):
                    self._session_tools.append(ListProjectsTool(catalog=project_catalog))
                if self._loop_injected_admitted(SWITCH_PROJECT_TOOL_ID):
                    self._session_tools.append(
                        SwitchProjectTool(
                            catalog=project_catalog,
                            rebind=self.rebind_workspace,
                        )
                    )
        if session_tool_registry is not None and session_id is not None:
            self._session_tools.extend(
                session_tool_registry.build_for(
                    allowed_tools,
                    session_id=session_id,
                    event_logger=agent_context.event_logger,
                    session_capabilities=session_capabilities,
                    # Deny wins over everything else this call admits —
                    # unconditional, the capability auto-surface, even a named
                    # allowlist entry. Same list ``ids_for`` (the operator-facing
                    # catalog) is given, so the two selections cannot drift.
                    denied_tools=denied_tools,
                    # df875 law: a PERMISSIVE allowlist (FE mcp_tools) is
                    # only an MCP ceiling, so an unconditional tool
                    # (schedule_trigger) still surfaces; a STRICT AgentDef scope
                    # must name it. Mirrors the spawn_in_scope gate above.
                    strict_tool_scope=strict_tool_scope,
                    # Delegation privilege ceiling: the loop's effective
                    # (already-narrowed) capability_mode gates SESSION tools too,
                    # not just registry tools — else a read_only spawn would still
                    # receive write-tier session actions (submit/mint/commit/arm).
                    # Root is always "all" (no-op); a narrowed sub-agent attenuates.
                    capability_mode=agent_context.capability_mode,
                )
            )
        # Caller-injected session tools (e.g. the structured-response emit
        # tool, and every client-declared device tool) — no plugin manifest
        # needed. They terminate / dispatch through the same machinery as
        # plugin tools.
        #
        # They ARE gated on ``capability_mode``, through the same
        # ``capability_mode_admits`` predicate ``build_for`` uses above. This
        # append sits AFTER build_for's gates, so without this it is a hole in
        # the privilege ceiling: a ``read_only`` sub-agent would be handed
        # every client-declared device tool, which now includes a shell at
        # shell UID. A tool that declares no ``capability`` is treated as
        # ``execute`` — session tools are actions, and an undeclared tier must
        # fail closed under a restrictive mode rather than open.
        if extra_session_tools:
            self._session_tools.extend(
                tool
                for tool in extra_session_tools
                if capability_mode_admits(
                    agent_context.capability_mode,
                    getattr(tool, "capability", None) or "execute",
                )
            )

        # Self-steering model control. Bound whenever the operator opts into
        # self-steering fallback — unlike a task tool it is resilience
        # infrastructure (peer of the automatic fallback ladder), so it is NOT
        # ceiling-checked against ``allowed_tools``: a child pinned by its
        # AgentDef to a model the key rejects (the live failure) must be able to
        # switch even though that AgentDef never named the tool, and is attached
        # at EVERY depth, not just the root. Off by default (self_steering=False)
        # ⇒ zero extra surface, so it never widens a strict scope uninvited.
        if session_id is not None and bool(
            get_config_value("llm", "fallback", "self_steering", default=False)
        ):
            from mewbo_core.tooling.model_control import ModelControlTool

            self._session_tools.append(
                ModelControlTool(
                    session_id=session_id,
                    agent_id=agent_context.agent_id,
                    depth=agent_context.depth,
                    get_active_model=lambda: self._active_model,
                    get_ladder=lambda: [self._ctx.model_name, *self._ctx.fallback_models],
                    get_strategy=lambda: self._retry_strategy,
                    has_unanswered_tool_use=self._has_dangling_tool_use,
                    apply_switch=self._request_model_switch,
                    # Route through ``_emit_event`` (not the raw sink) so a
                    # deliberate switch's ``llm_fallback`` is captured by the
                    # resilience note too, keeping it the complete switch record.
                    event_logger=self._emit_event,
                    get_step=lambda: self._llm_call_step,
                    max_switches=int(
                        get_config_value("llm", "fallback", "max_switches", default=2)
                    ),
                    allow_upgrade=bool(
                        get_config_value("llm", "fallback", "allow_upgrade", default=False)
                    ),
                )
            )

    # ------------------------------------------------------------------
    # Public API
    # ------------------------------------------------------------------

    @property
    def active_model(self) -> str:
        """The model this loop last generated with — the one that SERVED the run.

        Read by the orchestrator, which otherwise builds its failure record from
        the model frozen at construction: a run that escalated A -> B -> C then
        died named A, and an operator reading that benches a model that had not
        been active for minutes. Equal to the configured model until a sticky
        escalation promotes it, so the ordinary path is unchanged.
        """
        return self._active_model

    async def run(
        self,
        user_query: str,
        *,
        tool_specs: list[ToolSpec],
        context: ContextSnapshot | None = None,
        plan: Plan | None = None,
        mode: str = "act",
    ) -> tuple[TaskQueue, OrchestrationState]:
        """Run the async tool-use loop and return TaskQueue + OrchestrationState."""
        state = OrchestrationState(goal=user_query)
        # Plan-mode is enforced via: (1) filtered tool schema at bind time,
        # (2) path-scoped permission check on edits, (3) the exit_plan_mode
        # approval gate. The loop flips ``_current_mode`` to ``"act"`` after
        # the user approves a plan and re-binds tools.
        self._current_mode = mode if mode in {"plan", "act"} else "act"
        if self._current_mode == "plan" and self._session_id is not None:
            state.plan_path = plan_file_for(self._session_id)
            ensure_plan_dir(self._session_id)
        # Propagate plan context so children inherit session and mode.
        if self._spawn_agent_tool is not None:
            self._spawn_agent_tool.session_id = self._session_id
            self._spawn_agent_tool.parent_mode = self._current_mode
            # The EFFECTIVE set, not the deferral-active subset: deferral strips
            # schemas from the initial bind and re-fetches them through
            # tool_search, so a child seeded from it would lose tools this agent
            # genuinely holds. Children narrow this set; they never widen it.
            self._spawn_agent_tool.parent_tool_specs = list(tool_specs)
        executed_steps: list[ActionStep] = []
        tool_outputs: list[str] = []
        last_error: str | None = None
        final_response: str | None = None

        # Register self in the hypervisor registry.
        # Reuse the handle created by SpawnAgentTool when one already exists for
        # this agent_id — avoids overwriting it and losing the reference held by
        # the lifecycle manager (which later stores AgentResult on the handle).
        existing = await self._ctx.registry.get(self._ctx.agent_id)
        if existing is not None:
            handle = existing
        else:
            handle = AgentHandle(
                agent_id=self._ctx.agent_id,
                parent_id=self._ctx.parent_id,
                depth=self._ctx.depth,
                model_name=self._ctx.model_name,
                task_description=user_query[:200],
                # A spawned child's handle is normally
                # pre-registered (and pre-stamped) by SpawnAgentTool before its
                # loop ever runs, so this branch is the exception (a directly
                # constructed loop, e.g. the root). Stamping ``self._contract``
                # here too keeps the handle self-consistent with whatever this
                # loop was actually built with, instead of silently reading
                # back the disabled default.
                contract=self._contract or DelegationContract(),
            )
            await self._ctx.registry.register(handle)
        handle.status = "running"

        # Background watchdog for stall detection.
        # Code-level reflex — zero token cost. Root only.
        watchdog_task: asyncio.Task[None] | None = None
        if self._ctx.depth == 0:
            watchdog_task = asyncio.create_task(self._watchdog())

        # Langfuse context managers — initialized in try, cleaned in finally.
        agent_span: Any = None
        _agent_span_cm: Any = None
        _propagate_cm: Any = None

        try:
            # Global eye for root agent
            agent_tree = ""
            if self._ctx.depth == 0:
                agent_tree = await self._ctx.registry.render_agent_tree(
                    exclude_agent_id=self._ctx.agent_id,
                )
            # Deferred-tool partitioning. When ``agent.tool_search.mode`` is
            # 'on', MCP / metadata.deferred specs are stripped from the
            # initial bind and surfaced by name only via the
            # ``<available-deferred-tools>`` block. The model fetches the
            # schemas it needs through ``tool_search``; the per-turn re-bind
            # hook below grows the bound list as tools are discovered.
            self._tool_search_enabled = self._is_tool_search_enabled(tool_specs)
            self._tool_specs_full = list(tool_specs)
            self._deferred_ids = (
                {s.tool_id for s in tool_specs if is_deferred(s)}
                if self._tool_search_enabled
                else set()
            )
            active_specs = self._select_active_specs(tool_specs, discovered=set())
            messages = self._build_messages(
                user_query,
                context,
                plan,
                agent_tree=agent_tree,
            )
            # Expose the live transcript so the model_control continuity-lock
            # guardrail can inspect it (a mutable reference — it sees appends).
            self._live_messages = messages
            tool_schemas = self._build_tool_schemas_for_mode(
                active_specs,
                self._current_mode,
            )
            model = self._bind_model(tool_schemas)
            self._last_active_ids = {s.tool_id for s in active_specs}

            langfuse_handler = build_langfuse_handler(
                # The loop does NOT author trace identity. The real session and
                # principal arrive from the enclosing session context via
                # ``propagate_attributes`` and override anything the handler
                # carries — so a value invented here is not an override, it is
                # contradictory garbage sitting in every observation's metadata.
                # An empty string attaches no metadata key at all, which is the
                # honest reading of "this seam does not know". The session id is
                # passed only where the loop genuinely holds one.
                user_id="",
                session_id=self._session_id or "",
                trace_name="mewbo-tool-use",
                version=get_version(),
                release=get_config_value("runtime", "envmode", default="Not Specified"),
            )
            invoke_config: dict[str, Any] = {}
            if langfuse_handler is not None:
                invoke_config["callbacks"] = [langfuse_handler]
                metadata = getattr(langfuse_handler, "langfuse_metadata", None)
                if isinstance(metadata, dict) and metadata:
                    invoke_config["metadata"] = metadata

            # -- Langfuse: agent-level span + attribute propagation --------
            # Typed ``agent`` so the trace renders as an agent graph, and named
            # from the AgentDef rather than the runtime handle id: a per-run hex
            # id is unbounded cardinality and folds every aggregation into
            # one-bucket-per-run. The handle id keeps its place in metadata.
            _agent_def_name = await self._agent_def_name()
            _agent_span_cm = langfuse_trace_span(
                f"invoke_agent {_agent_def_name}",
                as_type="agent",
                attributes={
                    "gen_ai.operation.name": "invoke_agent",
                    "gen_ai.agent.name": _agent_def_name,
                },
                metadata={
                    "agentid": self._ctx.agent_id[:12],
                    "model": self._ctx.model_name,
                    "depth": str(self._ctx.depth),
                    "mode": self._current_mode,
                },
                input_data={"task": user_query[:200]},
            )
            agent_span = _agent_span_cm.__enter__()
            _propagate_cm = langfuse_propagate(
                tags=[
                    "mewbo-tool-use",
                    f"model:{self._ctx.model_name}",
                    f"depth:{self._ctx.depth}",
                ]
            )
            _propagate_cm.__enter__()

            turns = 0
            self._emit_safety_disclosure()
            # One atomic resilience strategy per run — holds the retry budget,
            # circuit breaker and policy knobs; survives every turn. Retained on
            # the instance so the model_control tool reuses THIS run's budget +
            # circuit breaker for its own switch guardrails.
            retry_strategy = RetryStrategy.from_config()
            self._retry_strategy = retry_strategy
            doom_guard = DoomLoopGuard.from_config(
                extra_poll_rules=self._poll_class_rules(tool_specs)
            )
            # Two-gate arming, computed ONCE for the run: the write-progress
            # signal is only meaningful for an agent that could plausibly
            # WRITE at all — narrowed by its own capability_mode AND actually
            # holding a write-tier tool. Anything else (read_only children, a
            # session with no write tools bound) leaves the signal permanently
            # inert.
            write_capable = self._ctx.capability_mode in {"execute", "all"} and any(
                s.capability_tier() == "write" for s in tool_specs
            )
            write_progress = WriteProgressSignal.from_config(write_capable=write_capable)
            write_progress_stamped = False
            # One-shot latch so the per-agent budget warning
            # (distinct from the session-wide ``loop.budget_warning`` above)
            # fires exactly once per run, not on every turn inside headroom.
            contract_step_warned = False
            # ``(tool_id, code)`` of the most recent blocked-class envelope
            # error that has NOT since been cleared. "Unrecovered" is judged per
            # tool: a later SUCCESS from the same tool means the blocked
            # operation went through after all, and only that clears it —
            # an unrelated tool succeeding says nothing about whether the repo
            # ever became reachable.
            last_blocked: tuple[str, str] | None = None
            # Consecutive promise-as-completion refusals (see the gate below).
            promise_nudges = 0
            # Consecutive required-terminal refusals (see the gate below).
            required_terminal_nudges = 0
            while not state.done:
                # Check cancellation. This is the between-turns read; the same
                # signal also guards the model call and each tool batch below,
                # so a stop never waits out whichever of those is in flight.
                if self._cancellation.requested:
                    state.done = True
                    state.done_reason = "canceled"
                    break

                # Safety-plane observer — the between-turns half of the plane.
                # A scalar check, so its cost does not grow with run length.
                if self._check_safety_turn(turns):
                    state.done = True
                    state.done_reason = "safety_blocked"
                    break

                # Check for interrupt (root agent only).
                if self._ctx.interrupt_step is not None and self._ctx.interrupt_step.is_set():
                    self._ctx.interrupt_step.clear()
                    messages.append(
                        HumanMessage(
                            content=get_prompt_registry().render("loop.interrupt_marker")
                        )
                    )

                # Drain any queued user steering messages (root agent only).
                if self._ctx.message_queue is not None:
                    while not self._ctx.message_queue.empty():
                        try:
                            msg = self._ctx.message_queue.get_nowait()
                            messages.append(HumanMessage(content=msg))
                        except _queue_mod.Empty:
                            break

                # Apply a deliberate model switch the model_control tool
                # requested last turn. Delegated to the SAME escalation path a
                # sticky fallback uses (promote active model / re-render prompt /
                # re-derive edit tool / rebind), applied at this turn boundary so
                # the transcript tail is untouched. Pin it on the strategy too, so
                # ``_order_models`` keeps the chosen model at the chain head
                # instead of a prior sticky pin reordering it back out.
                if self._requested_switch is not None:
                    _switch_target = self._requested_switch
                    self._requested_switch = None
                    retry_strategy._pinned_model = _switch_target
                    tool_schemas, model = self._apply_model_escalation(
                        _switch_target,
                        messages,
                        context=context,
                        plan=plan,
                        agent_tree=agent_tree,
                        tool_schemas=tool_schemas,
                        model=model,
                    )

                # Adopt a workspace switch ``switch_project`` performed mid-turn.
                # ``rebind_workspace`` already moved every piece of loop state
                # and built the new binding; what it deliberately left is the
                # transcript rewrite, taken here at the turn boundary exactly as
                # a requested model switch is. Re-rendering ``messages[0]`` is
                # what carries the new project's instructions, environment block
                # and git context into the model's next generation.
                if self._pending_workspace_bind is not None:
                    tool_schemas, model = self._pending_workspace_bind
                    self._pending_workspace_bind = None
                    messages[0] = SystemMessage(
                        content=self._render_system_prompt(context, plan, agent_tree)
                    )
                    self._active_resilience_note = self._resilience_note.render()

                # Keep the resilience note current in the system prompt: re-render
                # ``messages[0]`` only when the note text changed since it was last
                # baked in, so a clean run never pays and a healing run re-renders
                # at most once per new retry/fallback event.
                _note = self._resilience_note.render()
                if _note != self._active_resilience_note:
                    self._active_resilience_note = _note
                    messages[0] = SystemMessage(
                        content=self._render_system_prompt(context, plan, agent_tree)
                    )

                # The span stays — it is the natural parent of this turn's
                # generation and tool calls — but its NAME must not carry the
                # turn counter: a per-execution integer in a name is unbounded
                # cardinality, so "how long does a step take" had as many
                # groups as the longest run has turns. The index is a metadata
                # field, which is where a filter can still reach it.
                with langfuse_trace_span(
                    "agent_step",
                    metadata={
                        "turn": str(turns),
                        "model": self._ctx.model_name,
                    },
                ) as span:
                    if span is not None:
                        try:
                            span.update_trace(input={"turn": turns, "message_count": len(messages)})
                        except Exception:
                            pass

                    # Graduated enforcement: warn as the budget nears, then force
                    # ONE wrap-up turn at exhaustion so an unbounded
                    # fan-out can't run away — but the agent still gets to
                    # answer instead of a bare halt.
                    if self._ctx.registry.budget_exhausted():
                        final_response = await self._budget_wrapup_turn(
                            "budget_exhausted",
                            state=state,
                            messages=messages,
                            tool_outputs=tool_outputs,
                            invoke_config=invoke_config,
                        )
                        break
                    if self._ctx.registry.budget_warning():
                        messages.append(
                            SystemMessage(
                                content=get_prompt_registry().render("loop.budget_warning")
                            )
                        )

                    # DelegationContract. A per-agent bound
                    # LAYERED UNDER the session budget just checked above:
                    # checked here regardless (a spawner's ceiling applies even
                    # when the shared pool has headroom left). Disabled
                    # contracts (the default) skip this entirely.
                    if self._contract is not None and self._contract.enabled:
                        contract_over = False
                        step_state = await self._ctx.registry.agent_step_state(
                            self._ctx.agent_id
                        )
                        if step_state == "over":
                            contract_over = True
                        elif step_state == "warn" and not contract_step_warned:
                            contract_step_warned = True
                            messages.append(
                                SystemMessage(
                                    content=get_prompt_registry().render(
                                        "loop.agent_budget_warning"
                                    )
                                )
                            )
                        if not contract_over:
                            token_state = await self._ctx.registry.agent_token_state(
                                self._ctx.agent_id
                            )
                            if token_state == "over":
                                contract_over = True
                        if contract_over:
                            final_response = await self._budget_wrapup_turn(
                                "halted_agent_budget",
                                state=state,
                                messages=messages,
                                tool_outputs=tool_outputs,
                                invoke_config=invoke_config,
                            )
                            break

                    # Heartbeat events so clients (console/CLI) can distinguish
                    # "waiting on LLM" from a silent hang. ``bound_tools`` is the
                    # size of the surface this call carries — the transcript
                    # recorded nothing about the bound set anywhere, so a step
                    # that silently bound a collapsed one was only diagnosable by
                    # reading the model's behaviour back. A COUNT is deliberate:
                    # the full name list on every step of every run is payload
                    # bloat on a hot path, and a drop shows up in the count.
                    self._emit_event(
                        {
                            "type": "llm_call_start",
                            "payload": {
                                "agent_id": self._ctx.agent_id,
                                "depth": self._ctx.depth,
                                "step": turns,
                                "model": self._active_model,
                                "bound_tools": self._bound_tool_count,
                            },
                        }
                    )
                    # Own clock for the successful ``llm_call_end`` payload's
                    # ``duration_ms``, separate from the liveness leg below: that
                    # one is re-armed per retry/fallback attempt
                    # (``_invoke_with_resilience``), so it cannot bracket the
                    # whole logical call the way this single capture does.
                    _llm_call_t0 = _time.monotonic()
                    # Arm the liveness leg for exactly the window this call is
                    # outstanding; the ``finally`` disarms it on every exit so a
                    # completed call can never read as a wedged one.
                    self._llm_call_started_at = _time.monotonic()
                    self._llm_call_step = turns
                    try:
                        # Guarded: the resilience ladder can spend minutes across
                        # retries and fallbacks, so an unguarded await would keep
                        # a stop invisible until the ladder settled.
                        response, _final_model = await self._cancellation.guard(
                            self._invoke_with_resilience(
                                primary_model=model,
                                messages=messages,
                                tool_schemas=tool_schemas,
                                turns=turns,
                                invoke_config=invoke_config,
                                strategy=retry_strategy,
                            )
                        )
                    except RunCancelled:
                        state.done = True
                        state.done_reason = "canceled"
                        break
                    except LlmResilienceExhausted as exhausted:
                        # Clean halt: surface a true failure (never masked as
                        # "completed") so the FE can offer one-click recovery.
                        self._emit_event(
                            {
                                "type": "llm_call_end",
                                "payload": {
                                    "agent_id": self._ctx.agent_id,
                                    "depth": self._ctx.depth,
                                    "step": turns,
                                    "success": False,
                                    "model": (
                                        exhausted.models_tried[-1]
                                        if exhausted.models_tried
                                        else self._ctx.model_name
                                    ),
                                    "error_type": exhausted.last_error_type,
                                    "reason": exhausted.reason,
                                },
                            }
                        )
                        if span is not None:
                            try:
                                span.update(
                                    level="ERROR",
                                    # Same substitution as the completion string:
                                    # ``str(TimeoutError())`` is empty, and this
                                    # span write would otherwise carry a void
                                    # status_message for the very failure it marks.
                                    status_message=LlmResilienceExhausted.describe_error(
                                        exhausted.last_error
                                    ),
                                    metadata={
                                        "errortype": exhausted.last_error_type,
                                        "models_tried": ",".join(exhausted.models_tried),
                                        "reason": exhausted.reason,
                                    },
                                )
                            except Exception:
                                pass
                        record_span_exception(
                            span,
                            exhausted.last_error,
                            attributes={
                                "errortype": exhausted.last_error_type or "",
                                "reason": exhausted.reason or "",
                            },
                        )
                        raise
                    finally:
                        self._llm_call_started_at = None
                    # Fallback ladder: if the resilience strategy escalated
                    # to (and pinned) a different model, re-render the system
                    # prompt + re-derive the edit-tool variant against THAT model
                    # so the heal is behavioural, not just a model swap.
                    tool_schemas, model = self._apply_model_escalation(
                        _final_model,
                        messages,
                        context=context,
                        plan=plan,
                        agent_tree=agent_tree,
                        tool_schemas=tool_schemas,
                        model=model,
                    )
                    _step_usage = getattr(response, "usage_metadata", None)
                    _h_ref = await self._ctx.registry.get(self._ctx.agent_id)
                    # LangChain ``UsageMetadata`` exposes provider cache and
                    # reasoning subtotals (Anthropic + OpenAI normalised):
                    #   input_token_details.cache_creation — written to cache
                    #     this call (Anthropic 5-min: 1.25× input price)
                    #   input_token_details.cache_read — served from cache
                    #     (Anthropic: 0.1× input; OpenAI: 0.5× input)
                    #   output_token_details.reasoning — extended-thinking /
                    #     o1 hidden tokens (billed as output)
                    # Capturing them per call lets clients show fresh-vs-
                    # cached breakdown and an honest billable signal that
                    # accounts for cache discounts.
                    _in_det = _step_usage.get("input_token_details") or {} if _step_usage else {}
                    _out_det = _step_usage.get("output_token_details") or {} if _step_usage else {}
                    self._emit_event(
                        {
                            "type": "llm_call_end",
                            "payload": {
                                "agent_id": self._ctx.agent_id,
                                "depth": self._ctx.depth,
                                "step": turns,
                                "success": True,
                                "model": _final_model,
                                "input_tokens": (
                                    _step_usage.get("input_tokens", 0) if _step_usage else 0
                                ),
                                "output_tokens": (
                                    _step_usage.get("output_tokens", 0) if _step_usage else 0
                                ),
                                "cache_creation_input_tokens": int(
                                    _in_det.get("cache_creation", 0) or 0
                                ),
                                "cache_read_input_tokens": int(_in_det.get("cache_read", 0) or 0),
                                "reasoning_output_tokens": int(_out_det.get("reasoning", 0) or 0),
                                "cumulative_input_tokens": (_h_ref.input_tokens if _h_ref else 0),
                                "cumulative_output_tokens": (_h_ref.output_tokens if _h_ref else 0),
                                "duration_ms": int((_time.monotonic() - _llm_call_t0) * 1000),
                            },
                        }
                    )
                    # Strip thinking blocks from the response before appending
                    # to the conversation history.  Anthropic requires a
                    # ``signature`` field on thinking blocks when replayed,
                    # but proxies (LiteLLM) may not preserve it.
                    raw = getattr(response, "content", None)
                    if isinstance(raw, list):
                        # A reasoning model's answer arrives as a BARE STRING
                        # element alongside ``{"type": "thinking", ...}``
                        # dicts, not as a proper content part. Left as-is, a
                        # strict OpenAI-shaped backend (a self-hosted Ollama
                        # behind LiteLLM) rejects the replayed history with
                        # 400 "invalid message format" the moment the turn is
                        # replayed on a later request.
                        sanitized: list[str | dict[Any, Any]] = [
                            block if isinstance(block, dict) else {"type": "text", "text": block}
                            for block in raw
                            if not (isinstance(block, dict) and block.get("type") == "thinking")
                        ]
                        # Never leave empty-string assistant content.
                        # Empty assistant turns in
                        # history cause extended-thinking models to
                        # hallucinate framework-style placeholders.
                        if not sanitized:
                            sanitized = [{"type": "text", "text": _NO_CONTENT_PLACEHOLDER}]
                        response = AIMessage(
                            content=sanitized,
                            tool_calls=response.tool_calls,
                            additional_kwargs=response.additional_kwargs,
                            usage_metadata=response.usage_metadata,
                            id=response.id,
                        )
                    elif response.tool_calls and (not raw or not str(raw).strip()):
                        # The proxy (LiteLLM) strips thinking blocks itself
                        # and returns ``content=""`` (a STRING, not a list).
                        # Without this branch, the empty string survives
                        # sanitisation and gets replayed in history, causing
                        # the model to hallucinate placeholder meta-text.
                        response = AIMessage(
                            content=_NO_CONTENT_PLACEHOLDER,
                            tool_calls=response.tool_calls,
                            additional_kwargs=response.additional_kwargs,
                            usage_metadata=response.usage_metadata,
                            id=response.id,
                        )
                    messages.append(response)

                    if not response.tool_calls:
                        # Text response — the model claims completion.
                        content = self._extract_text_content(getattr(response, "content", ""))
                        # Verifier gate: run the ground-truth check BEFORE
                        # accepting the claim (only when armed and not already
                        # green). A failure with a retry left injects the
                        # grounded verifier output and ``continue``s — funnelling
                        # back through the top-of-loop budget checks FIRST, so
                        # retries are bounded by BOTH verification_max_retries AND
                        # the step/wall budget. Exhausted retries accept the text
                        # but flag it honestly (``verification_failed``); a pass
                        # falls through to the normal ``completed`` accept.
                        if self._verification_active and not self._verify_passed:
                            state.verify_attempts += 1
                            outcome = await self._run_verifier(
                                step=turns, attempt=state.verify_attempts
                            )
                            if outcome.passed:
                                self._verify_passed = True
                                state.verified = True
                            elif self._verify_retries_left > 0:
                                self._verify_retries_left -= 1
                                messages.append(
                                    SystemMessage(
                                        content=get_prompt_registry().render(
                                            "loop.verification_failed",
                                            output=outcome.feedback,
                                        )
                                    )
                                )
                                continue
                            else:
                                final_response = content
                                tool_outputs.append(content)
                                state.done = True
                                state.done_reason = "verification_failed"
                                state.verified = False
                                break
                        # Promise-as-completion gate (root only). A clean
                        # terminal declared while the root's OWN background runs
                        # are still live is a promise about future work — "I'll
                        # check back shortly" with an unfinished probe fleet
                        # behind it — not a completion. Refuse it here, upstream
                        # of the orchestrator's honesty downgrade, and send the
                        # model to await its runs. ``collect_running`` is keyed on
                        # THIS agent so it counts only owned CHILDREN — the seam
                        # deliberately does NOT read the session-wide ownership
                        # index, because the root's own handle is still
                        # ``running`` here (it is marked done only in the loop's
                        # finally), so that index would count the root itself and
                        # refuse every terminal. Bounded, so a model that will not
                        # wait is eventually let through rather than burning the
                        # whole budget spinning.
                        if self._ctx.depth == 0 and promise_nudges < _PROMISE_GATE_MAX_NUDGES:
                            live_owned = await self._ctx.registry.collect_running(
                                self._ctx.agent_id
                            )
                            if live_owned:
                                promise_nudges += 1
                                messages.append(
                                    SystemMessage(
                                        content=get_prompt_registry().render(
                                            "loop.agents_still_running",
                                            count=len(live_owned),
                                            ids=", ".join(
                                                h.agent_id[:8] for h in live_owned
                                            ),
                                        )
                                    )
                                )
                                continue
                        # Required-terminal gate (root only). A session tool can
                        # declare that ending the run REQUIRES a call to it —
                        # ``wiki_emit_answer`` is the canonical case: a composed
                        # answer delivered as plain text instead of through the
                        # tool is discarded downstream, and the run has no other
                        # channel to tell the user it happened. Refuse the clean
                        # terminal while a declared obligation is unmet and nudge
                        # the model to call it; bounded the same way the promise
                        # gate above is, so a model that genuinely cannot emit is
                        # eventually let through rather than spinning the whole
                        # budget here. UNLIKE the promise gate, exhaustion here
                        # stamps the honest ``unmet_goal`` rather than a claimed
                        # completion the user never saw.
                        if self._ctx.depth == 0:
                            unmet_terminals = self._unmet_required_terminals()
                            if unmet_terminals:
                                if required_terminal_nudges < _REQUIRED_TERMINAL_MAX_NUDGES:
                                    required_terminal_nudges += 1
                                    messages.append(
                                        SystemMessage(
                                            content=get_prompt_registry().render(
                                                "loop.required_terminal_missing",
                                                tool_ids=", ".join(unmet_terminals),
                                            )
                                        )
                                    )
                                    continue
                                final_response = content
                                tool_outputs.append(content)
                                state.done = True
                                state.done_reason = "unmet_goal"
                                if last_blocked is not None:
                                    state.blocked_code = last_blocked[1]
                                break
                        final_response = content
                        tool_outputs.append(content)
                        state.done = True
                        state.done_reason = "completed"
                        # The accept is NOT unconditional: a run whose clone never
                        # authenticated would otherwise stamp a clean completion
                        # because its LAST turn happened to be text. Consult the
                        # unrecovered blocked-class error instead — it says the
                        # goal was unreachable this run whatever the closing
                        # prose claims. Carried as its OWN field, never as a new
                        # done_reason: that vocabulary is a wire contract shared
                        # with every client, and the status layer owns the
                        # mapping from this code to a user-facing state.
                        if last_blocked is not None:
                            state.blocked_code = last_blocked[1]
                        break

                    # The model is repeating the same tool + input with no
                    # progress. Halt cleanly and hand back for one-click
                    # recovery instead of executing the same call again.
                    doom_guard.observe(response.tool_calls)
                    if doom_guard.is_stuck():
                        # Keep the transcript valid for resume: the AIMessage
                        # with tool_calls was already appended, so synthesize
                        # interrupted results for its dangling calls.
                        repair_tool_pairing(messages)
                        _repeated = response.tool_calls[0].get("name", "tool")
                        last_error = (
                            f"Halted: model repeated '{_repeated}' "
                            f"{doom_guard.threshold}x with identical input and result "
                            "(no progress)."
                        )
                        halt_payload: RecoveryHaltPayload = {
                            "action": "halt_no_progress",
                            "agent_id": self._ctx.agent_id,
                            "depth": self._ctx.depth,
                            "step": turns,
                            "tool": _repeated,
                        }
                        self._emit_event({"type": "recovery", "payload": halt_payload})
                        state.done = True
                        state.done_reason = "halted_no_progress"
                        break

                    # Emit intermediate text as agent_message for trace logs.
                    text_content = self._extract_text_content(getattr(response, "content", ""))
                    if text_content:
                        self._emit_event(
                            {
                                "type": "agent_message",
                                "payload": {
                                    "text": text_content,
                                    "agent_id": self._ctx.agent_id,
                                    "depth": self._ctx.depth,
                                },
                            }
                        )

                    # Execute tool calls with concurrency-aware partitioning.
                    # Exclusive tools run alone; concurrent-safe tools are gathered.
                    specs_map = {s.tool_id: s for s in tool_specs}
                    batches = self._partition_tool_calls(response.tool_calls, specs_map)
                    results: list[ToolCallResult] = []
                    # Guarded per batch: a step can hold several long tool calls,
                    # and an unguarded loop here made the stop wait for ALL of
                    # them. Breaking out mid-step leaves the already-executed
                    # calls' ``tool_result`` events in the transcript and the rest
                    # absent, which is the honest record of what actually ran —
                    # nothing downstream re-drives the message list, because the
                    # run terminates without another model call.
                    cancelled_mid_step = False
                    for batch in batches:
                        try:
                            if batch.concurrent:
                                batch_results = await self._cancellation.guard(
                                    asyncio.gather(
                                        *[
                                            self._safe_execute(tc, tool_specs)
                                            for tc in batch.calls
                                        ],
                                    )
                                )
                                results.extend(batch_results)
                            else:
                                results.append(
                                    await self._cancellation.guard(
                                        self._safe_execute(batch.calls[0], tool_specs)
                                    )
                                )
                        except RunCancelled:
                            cancelled_mid_step = True
                            break
                    if cancelled_mid_step:
                        state.done = True
                        state.done_reason = "canceled"
                        break

                    # Feed results back to the doom guard so "no progress" means
                    # same input AND same outcome — a call whose result advances
                    # (e.g. check_agents as children finish) is healthy progress.
                    doom_guard.record_result(results)

                    # Track the blocked-class condition across turns so the
                    # completion seam can tell "finished" from "gave up".
                    for _r in results:
                        if _r.blocked_code:
                            last_blocked = (_r.tool_id, _r.blocked_code)
                        elif (
                            _r.success
                            and last_blocked is not None
                            and _r.tool_id == last_blocked[0]
                        ):
                            last_blocked = None

                    # Write-progress signal: distinct from the doom guard —
                    # counts consecutive steps for a write-capable agent with
                    # no WRITE-tier tool execution (read/execute/search/
                    # unknown-tier steps only). Observe-only: crossing the
                    # threshold always emits telemetry; the optional reminder
                    # never names the signal or its criteria.
                    write_progress.observe(results, specs_map)
                    if write_progress.threshold_crossed():
                        self._emit_event(
                            {
                                "type": "write_progress_signal",
                                "payload": {
                                    "agent_id": self._ctx.agent_id,
                                    "depth": self._ctx.depth,
                                    "step": turns,
                                    "steps_since_write": write_progress.steps_since_write,
                                    "threshold": write_progress.threshold,
                                },
                            }
                        )
                        if write_progress.reminder_enabled:
                            messages.append(
                                SystemMessage(
                                    content=get_prompt_registry().render(
                                        "loop.task_objective_reminder", goal=state.goal
                                    )
                                )
                            )
                    elif write_progress.exhausted() and not write_progress_stamped:
                        # Go quiet after the last event — flag it once for the
                        # parent's global eye instead of firing forever.
                        write_progress_stamped = True
                        _h_write_progress = await self._ctx.registry.get(self._ctx.agent_id)
                        if _h_write_progress:
                            _h_write_progress.progress_note = (
                                "No write-tier tool call in the last "
                                f"{write_progress.steps_since_write} steps."
                            )

                    for tool_call, result in zip(response.tool_calls, results):
                        messages.append(
                            ToolMessage(
                                # A multimodal result becomes a list of content
                                # parts, which LiteLLM translates into a native
                                # ``tool_result`` carrying the text block then
                                # the image block. Everything else stays the
                                # plain string it always was, so an ordinary
                                # result's cache prefix is byte-identical.
                                content=_tool_message_content(result),
                                tool_call_id=result.tool_call_id,
                            )
                        )
                        tool_outputs.append(f"{result.tool_id}: {result.content}")
                        if not result.success:
                            last_error = result.content
                        # Track as ActionStep for TaskQueue compatibility.
                        action_step = self._tool_call_to_action_step(tool_call)
                        mock = get_mock_speaker()
                        action_step.result = mock(content=result.content)
                        executed_steps.append(action_step)

                    # Episodic plan-mode: a session tool (e.g. exit_plan_mode)
                    # signals the loop to terminate so the thread exits
                    # cleanly. Approval/rejection happens out-of-band via
                    # SessionRuntime. Materialise the list so every tool's
                    # flag is consumed — ``any`` short-circuits and would
                    # leave a second tool's flag set for the next turn.
                    # ``terminal_reason()`` lets each tool declare the right
                    # done_reason (default: "awaiting_approval"; emit tool: "completed").
                    term_tools = [
                        (t, t.should_terminate_run()) for t in self._session_tools
                    ]
                    terminating = [t for t, flag in term_tools if flag]
                    if terminating:
                        state.done = True
                        state.done_reason = terminating[0].terminal_reason()
                        break

                    await self._ctx.registry.update_step(
                        self._ctx.agent_id,
                        results[-1].tool_id if results else "",
                    )
                    turns += 1

                    # Auto-update progress
                    # for parent monitoring. Zero token cost — direct write.
                    if self._ctx.depth > 0 and results:
                        last = results[-1]
                        note = f"turn {turns}: {last.tool_id}"
                        snippet = last.content[:100] if last.content else ""
                        if last.success:
                            note += f" -> {snippet}"
                        else:
                            note += f" -> FAILED: {snippet}"
                        handle = await self._ctx.registry.get(
                            self._ctx.agent_id,
                        )
                        if handle:
                            handle.progress_note = note

                    # Inject failure feedback so the model can adapt.
                    failures = [r for r in results if not r.success]
                    if failures:
                        messages.append(
                            SystemMessage(
                                content=f"{len(failures)}/{len(results)} tool call(s)"
                                " failed this step — adapt your approach."
                            )
                        )

                    # Re-bind newly discovered deferred tools. Discovery is
                    # derived from the message history each turn, so this is
                    # compaction-resilient: whatever survives compaction
                    # still drives the bound set on the next iteration.
                    if self._tool_search_enabled and self._deferred_ids:
                        discovered = self._discovered_from_messages(messages)
                        active_specs = self._select_active_specs(
                            self._tool_specs_full, discovered=discovered
                        )
                        new_active_ids = {s.tool_id for s in active_specs}
                        if new_active_ids != self._last_active_ids:
                            tool_schemas = self._build_tool_schemas_for_mode(
                                active_specs, self._current_mode
                            )
                            model = self._bind_model(tool_schemas)
                            self._last_active_ids = new_active_ids

                    # Proactive mid-loop compaction check.
                    if self._should_compact_messages(messages):
                        _compact_info = await self._compact_messages(messages)
                        if _compact_info:
                            await self._ctx.registry.record_compaction(
                                self._ctx.agent_id,
                            )
                            self._emit_event(
                                {
                                    "type": "context_compacted",
                                    "payload": {
                                        **_compact_info,
                                        "agent_id": self._ctx.agent_id,
                                        "depth": self._ctx.depth,
                                        "mode": "mid_loop",
                                        "turn": turns,
                                    },
                                }
                            )

                    if span is not None:
                        try:
                            span.update_trace(
                                output={
                                    "tool_calls": len(response.tool_calls),
                                    "turns": turns,
                                }
                            )
                        except Exception:
                            pass

            # Safety net: currently unreachable (all loop exits set
            # state.done=True), but retained as defensive code for future
            # exit paths that may break without setting state.done.
            if not state.done and final_response is None and messages:
                # Inject child results at synthesis.
                # Root may have non-blocking children still running — give them
                # a brief grace period, then include available results.
                if self._ctx.depth == 0:
                    running = await self._ctx.registry.collect_running(
                        self._ctx.agent_id,
                    )
                    if running:
                        await asyncio.sleep(2.0)
                    completed = await self._ctx.registry.collect_completed(
                        self._ctx.agent_id,
                    )
                    if completed:
                        result_lines = []
                        for h in completed:
                            r = h.result
                            if r:
                                result_lines.append(
                                    f"[{h.agent_id[:8]}] {r.status}: {r.summary or r.content[:300]}"
                                )
                        if result_lines:
                            messages.append(
                                SystemMessage(
                                    content=get_prompt_registry().render(
                                        "loop.agent_results_header",
                                        joined="\n".join(result_lines),
                                    ),
                                )
                            )
                    still_running = await self._ctx.registry.collect_running(
                        self._ctx.agent_id,
                    )
                    if still_running:
                        ids = ", ".join(h.agent_id[:8] for h in still_running)
                        messages.append(
                            SystemMessage(
                                content=get_prompt_registry().render(
                                    "loop.agents_still_running",
                                    count=len(still_running),
                                    ids=ids,
                                ),
                            )
                        )

                final_response = await self._unbound_wrapup_invoke(
                    prompt_key="loop.final_answer_synthesis",
                    model_name=self._ctx.model_name,
                    messages=messages,
                    tool_outputs=tool_outputs,
                    invoke_config=invoke_config,
                )

            if not state.done:
                state.done = True
                state.done_reason = "completed"

            # Forced closing summary for a CHILD that finished via store
            # side-effects. The heaviest sub-agents did all their real work
            # through tool writes and then stopped on empty text, so the parent
            # — which projects ``task_result`` as the child's summary — received
            # nothing, and a 3.3M-token probe reached its parent blank. One
            # unbound wrap-up turn forces a compressed summary before stop.
            # Root is exempt: its empty terminal is a user-facing turn, not a
            # summary owed upstream. A cancel is exempt too — it is not a
            # completion to summarize. The budget/synthesis paths already filled
            # ``final_response``, so the empty-guard skips them (no double turn).
            if (
                self._ctx.depth > 0
                and state.done_reason != "canceled"
                and not (final_response or "").strip()
            ):
                final_response = await self._unbound_wrapup_invoke(
                    prompt_key="loop.final_answer_synthesis",
                    model_name=self._active_model,
                    messages=messages,
                    tool_outputs=tool_outputs,
                    invoke_config=invoke_config,
                )

            # Build TaskQueue for compatibility with CLI / API consumers.
            plan_steps = list(plan.steps) if plan and plan.steps else []
            task_queue = TaskQueue(
                plan_steps=plan_steps,
                action_steps=executed_steps,
            )
            # task_result is the LLM's final synthesized text only.
            task_queue.task_result = (final_response or "").strip()
            task_queue.last_error = last_error
            state.tool_results = tool_outputs
            # Every OTHER terminal path (halt, budget wrap-up, verifier
            # exhaustion, plan-mode exit) carries the same fact — a blocked run
            # is blocked regardless of which exit it took.
            if state.blocked_code is None and last_blocked is not None:
                state.blocked_code = last_blocked[1]

        finally:
            # Close Langfuse agent span and propagation context.
            if agent_span is not None:
                try:
                    agent_span.update(
                        output={
                            "total_steps": turns,
                            "done_reason": state.done_reason or "unknown",
                        }
                    )
                except Exception:
                    pass
            if _propagate_cm is not None:
                try:
                    _propagate_cm.__exit__(None, None, None)
                except Exception:  # pragma: no cover - defensive
                    pass
            if _agent_span_cm is not None:
                try:
                    _agent_span_cm.__exit__(None, None, None)
                except Exception:  # pragma: no cover - defensive
                    pass

            if watchdog_task is not None and not watchdog_task.done():
                watchdog_task.cancel()

            # Wait for lifecycle managers to complete cleanup.
            if self._spawn_agent_tool is not None:
                await self._spawn_agent_tool.await_lifecycle_managers(timeout=3.0)

            # Cleanup: cancel any child agent that has not reached a terminal.
            # Reads ACTIVE_STATUSES rather than comparing against ``running``
            # alone: a capacity-deferred child sits at ``submitted`` until a slot
            # frees, so a sweep filtering on the one literal walked past exactly
            # the children that had never got to run.
            children = await self._ctx.registry.list_children(self._ctx.agent_id)
            for child in children:
                if child.status in ACTIVE_STATUSES:
                    await self._ctx.registry.cancel_agent(child.agent_id)

            # The terminal is PROJECTED, never re-derived here. ``state.done``
            # means the loop stopped, not that the task succeeded — it is True on
            # every exit path, cancellation and doom-halt included — so deriving
            # a status from it directly reported a clean success for a run the
            # user stopped. ``terminal_status()`` is the authority, and this is
            # the only mark a ROOT ever gets; a child was corrected a moment
            # later on the spawn path, which is why only the root kept the wrong
            # value permanently and why the child had a window where a reader
            # disagreed with the final answer. Projecting here settles both.
            await self._ctx.registry.mark_done(
                self._ctx.agent_id,
                state.terminal_status(),
            )

        return task_queue, state

    # ------------------------------------------------------------------
    # Error-isolated task wrapper
    # ------------------------------------------------------------------

    # A wait/sync tool (``check_agents``) bounds its OWN wait internally on the
    # requested ``timeout``; the outer execution ceiling must never clip that
    # below what the model asked for. Headroom on top of the requested wait for
    # the post-wait agent-tree render + result collection.
    _WAIT_TOOL_RENDER_MARGIN_S = 30.0

    def _get_tool_timeout(self, tool_name: str) -> float:
        """Get timeout for a tool. Uses spec.timeout, falls back to 120s."""
        spec = self._tool_registry.get_spec(tool_name) if self._tool_registry else None
        if spec:
            return spec.timeout
        return 120.0

    def llm_call_stall_age(self, now: float) -> float | None:
        """Seconds the in-flight LLM call has been outstanding, else ``None``.

        Pure read over ``(_llm_call_started_at, now)`` with no clock of its own,
        so the liveness rule is exercisable at any age without waiting for one
        — the watchdog supplies ``time.monotonic()``, a test supplies a number.
        """
        started = self._llm_call_started_at
        if started is None:
            return None
        return max(0.0, now - started)

    def _build_containment(self, cwd: str | None) -> WorkspaceContainment | None:
        """The filesystem-containment firebreak for *cwd*, or ``None`` when inert.

        The ONE place the three admitting conditions are spelled out — the
        enforcement kill-switch (``agent.workspace_enforcement``, ON by default),
        this agent's narrowed ``workspace_mode``, and a real workspace. A root
        agent fails the second condition (it runs ``full_access``), so the
        containment a run builds is the SPAWNED child's, which is the boundary
        the tiers exist to draw. Extracted
        from the constructor so :meth:`rebind_workspace` rebuilds it under the
        SAME predicate rather than a re-typed copy: a containment left rooted at
        the previous project is a jail around the wrong directory, and a
        re-typed predicate that drifts by one condition is how it would end up
        rooted at neither.
        """
        if (
            cwd
            and self._ctx.workspace_mode != "full_access"
            and bool(get_config_value("agent", "workspace_enforcement", default=False))
        ):
            return WorkspaceContainment(mode=self._ctx.workspace_mode, root=cwd)
        return None

    def rebind_workspace(self, entry: ProjectEntry) -> dict[str, object]:
        """Re-point this loop at *entry*'s directory. THE one mutation seam.

        Everything derived from the working directory moves together here,
        because a partial switch is worse than none: the agent believes it moved
        and half the machinery did not. Returns the facts ``switch_project``
        renders back to the model.

        The step that is easiest to miss is the SPAWN seam.
        :class:`~mewbo_core.agents.spawn_agent.SpawnAgentTool` holds its OWN copy of the
        workspace and forwards it verbatim to every child loop it builds, so a
        switch that updates only ``self._cwd`` leaves every sub-agent spawned
        afterwards working in the OLD project — silently, and with results that
        look plausible right up until they are applied to the wrong tree.

        Two deliberate limits, stated so nobody reads them as bugs:

        - The spec set NARROWS to what this run already held. A switch must
          never widen the tool ceiling the caller granted at run start, so a new
          project's own MCP servers are not admitted mid-run; a fresh run
          against that project resolves them normally, and the ``context`` event
          written here is what makes that fresh run land in the right place.
        - Skills ACCUMULATE rather than being replaced. The registry also holds
          plugin-contributed and user-level skills that a fresh scan of the new
          directory alone would drop, and losing those costs more than carrying
          the previous project's.
        """
        new_cwd = entry.path
        if not new_cwd:
            raise ValueError(f"Project '{entry.key}' has no directory to switch into.")

        # Captured before anything moves. ``_project_key`` is ``None`` until the
        # first switch: this loop is handed a DIRECTORY at construction and never
        # a catalog key (the app resolves the key to a path before calling), so
        # there is genuinely no prior key to report on the first move.
        previous_cwd = self._cwd
        previous_project = self._project_key
        self._project_key = entry.key

        self._cwd = new_cwd
        self._containment = self._build_containment(new_cwd)

        # Resolved BEFORE the spawn seam is re-pointed: children inherit the
        # instructions as a plain string captured at build time, so handing the
        # tool the previous project's text would pin every future child to it.
        self._project_instructions = discover_project_instructions(new_cwd)
        if self._spawn_agent_tool is not None:
            self._spawn_agent_tool.rebind_cwd(
                new_cwd, project_instructions=self._project_instructions
            )

        # Cached by cwd, so switching back to a project visited earlier in the
        # run costs a dict lookup rather than a second registry build.
        self._tool_registry = get_or_build_registry(cwd=new_cwd)
        previous_ids = {spec.tool_id for spec in self._tool_specs_full}
        self._tool_specs_full = [
            spec for spec in self._tool_registry.list_specs() if spec.tool_id in previous_ids
        ]

        skills = 0
        if self._enable_skills and self._skill_registry is not None:
            self._skill_registry.load(new_cwd)
            skills = len(self._skill_registry.list_all())

        # Bind now so the tool can report a truthful count, but leave the
        # system-prompt re-render and the swap of the live model to the turn
        # boundary — the same split ``_requested_switch`` makes, which keeps the
        # transcript tail untouched while a tool call is still in flight.
        active_specs = self._select_active_specs(
            self._tool_specs_full,
            discovered=self._discovered_from_messages(self._live_messages or []),
        )
        tool_schemas = self._build_tool_schemas_for_mode(active_specs, self._current_mode)
        self._pending_workspace_bind = (tool_schemas, self._bind_model(tool_schemas))
        self._last_active_ids = {spec.tool_id for spec in active_specs}

        # A plain ``context`` event carrying the two keys the API's session-cwd
        # resolver already walks backwards for. Re-engagement via /query,
        # /message, the diff endpoints and SessionSpec reconstruction therefore
        # all follow the switch with no new resolver and no new event type;
        # inventing one would leave every one of them reading the project the
        # session STARTED in.
        #
        # It carries the PREVIOUS payload forward because a context event is
        # read two incompatible ways in this tree. One camp folds every event
        # (``SessionStoreBase.merge_context_events``); the other takes the most
        # recent payload VERBATIM — the API's ``_load_last_context``, which
        # `/message` re-engage and `/recover` read for the model, the mode, the
        # tool grants and the step budget, and the console's ``getLastContext``,
        # which the model pill and the composer hydrate from. Emitting the two
        # switch keys ALONE would therefore not merely mislabel those readers,
        # it would BLANK them. The same shape has bitten here before: a
        # model-only context event once left every verbatim reader seeing a
        # session with no tool ceiling, no playbook and no cwd.
        #
        # Carrying the LAST payload (not a fold of all of them) is what makes
        # this exactly neutral: before the switch the effective context was that
        # payload, after it is that payload plus the two keys the switch really
        # changed. A fold would instead resurrect a field an earlier turn had
        # deliberately cleared — clients omit a cleared field rather than
        # sending null, which is precisely why the verbatim readers are verbatim.
        payload: dict[str, object] = {"project": entry.key, "cwd": new_cwd}
        if self._session_context_reader is not None:
            try:
                previous = self._session_context_reader()
            except Exception as exc:  # noqa: BLE001 - never fail a switch over carry-forward
                logging.warning("Could not read session context to carry forward: {}", exc)
            else:
                if isinstance(previous, dict):
                    payload = {**previous, **payload}
        self._emit_event({"type": "context", "payload": payload})

        return {
            "cwd": new_cwd,
            "previous_cwd": previous_cwd,
            "previous_project": previous_project,
            "project_instructions_found": bool(self._project_instructions),
            "bound_tools": self._bound_tool_count,
            "skills": skills,
        }

    def _loop_injected_admitted(self, tool_id: str) -> bool:
        """Whether the tool ceiling admits a tool this loop injects itself.

        The loop binds several tools OUTSIDE ``filter_specs`` by design — the
        spawn family, ``activate_skill``, the agent-management tools, the
        inline session tools. By design also meant outside the allowlist, so a
        strictly-scoped agent whose AgentDef named five tools was in fact
        handed those five plus everything here. That surplus is what a
        wandering run wanders into: across the observed burner sessions the
        winners and the burners ran the SAME model on the SAME task, and what
        separated them was reachable tool surface, not capability.

        The gate is the SAME shape as the spawn seam's (``spawn_in_scope``) and
        obeys the same three-state law: ``None`` is unrestricted, ``[]`` grants
        nothing, a list grants exactly its members. Crucially it fires ONLY
        under STRICT scope — a permissive console/mobile root always sends a
        large ``context.mcp_tools`` list that is a ceiling over MCP tools alone
        and never names a built-in, so treating it as authoritative here would
        strip every such session's built-ins. That is the same trap that once
        silently broke the mobile alarm flow.
        """
        if not self._strict_tool_scope or self._allowed_tools is None:
            return True
        return tool_id in self._allowed_tools

    def _poll_class_rules(self, tool_specs: list[ToolSpec]) -> list[PollClassRule]:
        """Poll-class rules for this run, gathered from what each tool declares.

        Two declaration seams, because this loop binds two populations that
        never meet: a registry tool declares ``ToolSpec.poll`` /
        ``poll_when_args``, and a session tool (absent from the registry
        entirely) declares ``poll_class`` / ``poll_when_args``. Both are read
        via ``getattr`` — ``SessionTool`` is a structural Protocol whose
        defaults a standalone implementer does not inherit, and these are
        optional attributes rather than contract members.

        Resolving per run from declarations is what keeps core out of it: no
        tool id appears here, so the next self-polling tool is added by
        declaring it on that tool, in whichever package owns it. A nested run's
        status probe answering "processing" twice in under a second is honest
        waiting, and no repetition-based detector can tell that from a stuck
        loop unless the tool says which call is which.
        """
        rules: list[PollClassRule] = []
        for source in (*tool_specs, *self._session_tools):
            tool_id = getattr(source, "tool_id", "")
            if not tool_id:
                continue
            when_args = tuple(getattr(source, "poll_when_args", ()) or ())
            if when_args:
                rules.append(
                    PollClassRule(tool_id=tool_id, when_args=frozenset(when_args))
                )
            # ``poll`` on a registry spec, ``poll_class`` on a session tool —
            # either spelling means "every call to me is a poll".
            elif getattr(source, "poll", False) or getattr(source, "poll_class", False):
                rules.append(PollClassRule(tool_id=tool_id))
        return rules

    def _result_char_cap(self, tool_id: str) -> int:
        """The model-facing result cap for *tool_id* (0 = uncapped).

        Resolution order — session-tool declaration, then registry spec, then
        the 2000 default. The session-tool arm is what this exists for: a
        session tool is never registered in the stateless ``ToolRegistry``, so
        ``get_spec`` returns ``None`` for EVERY one of them and the default
        applied to the whole class by accident rather than by any decision. The
        number is calibrated for unbounded shell/MCP output and is wrong for a
        curated first-party payload: a repo-manifest scan reached the model as
        22 of ~1,900 paths — 1.14% — while the next step of its own playbook
        required picking files out of that manifest.

        The ONE resolver behind both the model-facing truncation and the event
        summary, so the store can never claim a cap the model did not read.
        """
        session_tool = self._session_tool(tool_id)
        if session_tool is not None:
            # ``SessionTool`` is a structural Protocol: a standalone implementer
            # inherits no defaults from it, so the declaration must be read
            # defensively rather than assumed present on the instance.
            declared = getattr(session_tool, "max_result_chars", None)
            if isinstance(declared, int) and not isinstance(declared, bool):
                return declared
            return DEFAULT_SESSION_TOOL_MAX_RESULT_CHARS
        spec = self._tool_registry.get_spec(tool_id) if self._tool_registry else None
        if spec is not None:
            return spec.max_result_chars
        # The THIRD population, and the one that had no way to declare at all:
        # ``_bind_model`` binds several tools directly, so they are neither a
        # registry spec nor a ``SessionTool`` instance and fell through to the
        # 2000-char default that the arm above exists to keep off curated
        # payloads. Silence still means 2000 — correct for the small results the
        # rest of that population returns — but a tool whose result is authored
        # content now declares, on the module that owns it.
        return LOOP_INJECTED_RESULT_CAPS.get(tool_id, 2000)

    @staticmethod
    def _windowed(text: str, cap: int) -> str:
        """Fit *text* into *cap* characters, keeping BOTH ends over the middle.

        Head-only truncation (``text[:cap]``) discards the one part that answers
        the question: a command's verdict — the traceback, the failure summary,
        the exit banner — is at the END. Keeping only the head delivers the part
        the caller already knew (what was invoked) and drops the part it asked
        for, which is why so much tooling here pipes through ``tail``/``grep``
        before the output is ever returned.

        The head is kept because it identifies WHAT ran — a command echo, a
        config banner, the first error in a run of many — and the tail gets the
        larger share because that is where outcomes land. The omission is stated
        with its size rather than implied, so a reader can tell a windowed
        result from a complete one without re-deriving the cap.
        """
        if cap <= 0 or len(text) <= cap:
            return text
        marker_width = 64
        budget = max(0, cap - marker_width)
        head_chars = budget // 4
        tail_chars = budget - head_chars
        omitted = len(text) - head_chars - tail_chars
        marker = f"\n[... {omitted} characters omitted ...]\n"
        if tail_chars <= 0:
            return text[:cap]
        return f"{text[:head_chars]}{marker}{text[len(text) - tail_chars :]}"

    @staticmethod
    def _windowed_items(items: list[Any], budget: int) -> list[Any]:
        """Fit a LIST field into *budget* characters by dropping whole elements.

        A listing's bulk is its element COUNT, not any one element's length, so
        the honest degradation is fewer entries plus a count of what was
        dropped. A character cut through the middle of the serialized list
        would yield a truncated final path — a filename that does not exist,
        which is worse than a missing one.

        Head-kept, unlike ``_windowed``: a listing is ordered and a caller pages
        forward from where it stops, so the tail carries no verdict the way a
        command's stdout does.
        """
        # +4 per element covers the quoting and separator it costs once the list
        # is serialized; approximate, and ``_fitted_json`` re-fits against the
        # measured overshoot anyway.
        def cost(item: Any) -> int:
            return len(str(item)) + 4

        if sum(cost(item) for item in items) <= budget:
            return items
        # The marker is itself an element and must be paid for out of the same
        # budget — omit that and a list fitted to exactly the budget overshoots
        # by the marker's width, which sends the whole payload to the
        # ``{"truncated": true}`` wrapper this exists to avoid.
        marker_width = 64
        allowance = max(0, budget - marker_width)
        kept: list[Any] = []
        spent = 0
        for item in items:
            spent += cost(item)
            if spent > allowance:
                break
            kept.append(item)
        omitted = len(items) - len(kept)
        return [*kept, f"[... {omitted} of {len(items)} entries omitted ...]"]

    def _fitted_field(self, key: str, value: Any, budget: int) -> Any:
        """Fit ONE field of a dict result, by the kind of bulk it carries."""
        if key not in _WINDOWED_RESULT_FIELDS:
            return value
        if isinstance(value, str):
            return self._windowed(value, budget) if len(value) > budget else value
        if isinstance(value, list):
            return self._windowed_items(value, budget)
        return value

    def _fitted_json(self, content: dict[str, Any], cap: int) -> str:
        r"""Serialize *content* into at most *cap* characters, still PARSEABLE.

        ``_windowed`` fits a FIELD, but the cap applies to the SERIALIZED
        envelope, which is always larger — sibling keys, quoting, and the ``\\n``
        and ``\\t`` escapes of a source file all land outside the window. So a
        field windowed to exactly the cap still overshoots, and cutting the JSON
        string to recover leaves an unterminated value inside an unclosed
        object: the model reads damage rather than a bounded result, and loses
        the trailing metadata (a file's ``total_lines``) that is precisely what
        would let it ask for the rest.

        The corrective pass is MEASURED, not estimated. After one serialization
        the overshoot is known exactly, and dropping that many raw characters
        can only shrink the escaped form by at least as much, so a second pass
        settles it — hence a bounded loop rather than a convergence check.

        A payload whose bulk is not in a windowable field cannot be fitted this
        way at all. It is wrapped in one valid object instead of being handed
        over broken, which loses the structure but keeps the result parseable.
        """
        budget = cap
        serialized = ""
        for _ in range(2):
            fitted = {
                key: self._fitted_field(key, value, budget) for key, value in content.items()
            }
            serialized = json.dumps(fitted, ensure_ascii=False, default=str)
            if len(serialized) <= cap:
                return serialized
            budget -= len(serialized) - cap
            if budget <= 0:
                break
        # Half the cap because JSON escaping can double a string's length in the
        # worst case, so the wrapper stays within the same order as the cap.
        return json.dumps(
            {"truncated": True, "content": self._windowed(serialized, max(cap // 2, 1))},
            ensure_ascii=False,
        )

    def _session_tool(self, tool_id: str) -> SessionTool | None:
        """The session-tool instance bound as *tool_id* for this agent, or ``None``.

        Session tools bypass the stateless ``ToolRegistry``, so every structural
        ``getattr`` convention they declare (``max_result_chars``,
        ``result_headline``) has to find the INSTANCE first — one lookup, not one
        per convention.
        """
        return next((t for t in self._session_tools if t.tool_id == tool_id), None)

    def _unmet_required_terminals(self) -> list[str]:
        """Tool ids of every bound session tool whose terminal obligation is unmet.

        ``required_terminal`` / ``terminal_satisfied()`` are read defensively
        via ``getattr`` + ``callable`` — the same convention as
        ``max_result_chars`` and ``execution_timeout`` — because ``SessionTool``
        is a structural Protocol whose standalone implementers inherit no
        defaults from it. A tool that declares neither is never even
        considered, so this is a no-op for every session tool that came before
        this gate existed.

        A tool that declares ``required_terminal`` truthy but exposes no
        callable ``terminal_satisfied`` (or one that raises) counts as UNMET
        rather than being skipped: the tool asserted an obligation exists, and
        a broken or missing attestation must not silently waive it. The raise
        case is logged and swallowed — this hook runs at the natural-completion
        seam and must never be what crashes a turn.
        """
        unmet: list[str] = []
        for tool in self._session_tools:
            if not getattr(tool, "required_terminal", False):
                continue
            check = getattr(tool, "terminal_satisfied", None)
            satisfied = False
            if callable(check):
                try:
                    satisfied = bool(check())
                except Exception as exc:  # noqa: BLE001 — a declaration hook must not crash a turn
                    logging.warning(
                        "session tool {} terminal_satisfied() raised: {}",
                        getattr(tool, "tool_id", "?"),
                        exc,
                    )
                    satisfied = False
            if not satisfied:
                unmet.append(getattr(tool, "tool_id", "?"))
        return unmet

    def _declared_headline(self, tool_id: str) -> str | None:
        """Drain the one-line headline *tool_id* declared for the call just done.

        The same ``getattr``-on-the-instance convention as ``max_result_chars``
        and ``poll_class``: core cannot import the graph layer, so the structure
        IS the contract and a tool declaring nothing returns ``None`` here, so
        its event carries no headline.

        Every emit drains, including the error and timeout paths that use the
        headline for nothing. That is what keeps the state honest: a headline
        can never outlive the call that recorded it and surface as the title of
        the next one.

        Best-effort by construction — a log title must never be able to fail a
        tool call, so a raising hook is one warning and today's summary.
        """
        tool = self._session_tool(tool_id)
        hook = getattr(tool, "result_headline", None) if tool is not None else None
        if not callable(hook):
            return None
        try:
            declared = hook()
        except Exception as exc:  # noqa: BLE001 — a title must not break the call
            logging.warning("session tool {} result_headline failed: {}", tool_id, exc)
            return None
        if not isinstance(declared, str):
            return None
        # One line, bounded, markup-free. This rides the event into every client
        # and is replayed on every history read, so ``run_error.py``'s law holds
        # here too: a title is derived from structured facts, and anything that
        # starts to look like a response body (the first ``<``) is cut rather
        # than republished.
        one_line = " ".join(declared.split()).split("<", 1)[0]
        return one_line.strip()[:SESSION_TOOL_HEADLINE_MAX_CHARS].strip() or None

    def _tool_execution_timeout(self, tool_name: str, tool_call: Any) -> float | None:
        """The outer ``asyncio.wait_for`` ceiling for one tool call.

        A tool that bounds its OWN wait must not be clipped by this ceiling, or
        a legitimate long wait aborts with "timed out after 120.0s" the
        requested 300s never reached (production: three such clips on
        ``check_agents``). Worse, the flat ceiling silently applied to a tool
        whose whole contract is to block on a human: a SessionTool has no
        ``ToolSpec``, so ``ask_user_question`` inherited the 120s fallback and
        killed every question a user took two minutes to answer.

        So the exemption is DECLARED, per call, not a hardcoded name set — the
        same move ``DoomLoopGuard.poll_rules`` made for poll-class tools. A
        session tool may expose ``execution_timeout(tool_input)`` returning
        seconds, or ``None`` for NO ceiling; it is read defensively (a
        convention, not a Protocol member) and a raising hook falls back to the
        flat ceiling rather than failing the call. ``check_agents`` keeps its
        arm below because it is loop-injected and owns no class to declare on.
        """
        ceiling = self._get_tool_timeout(tool_name)
        args = tool_call.get("args") if isinstance(tool_call, dict) else None
        did_declare, declared = self._declared_execution_timeout(tool_name, args)
        if did_declare:
            return declared
        if tool_name not in DOOM_LOOP_EXEMPT_TOOLS:
            return ceiling
        if not isinstance(args, dict) or not args.get("wait"):
            return ceiling
        requested = args.get("timeout")
        if isinstance(requested, bool) or not isinstance(requested, (int, float)):
            return ceiling
        return max(ceiling, float(requested) + self._WAIT_TOOL_RENDER_MARGIN_S)

    def _declared_execution_timeout(
        self, tool_id: str, args: Any
    ) -> tuple[bool, float | None]:
        """A session tool's own ceiling for this call, as (declared?, seconds).

        The boolean is what lets ``None`` mean "no ceiling at all" — the whole
        point of the hook — while still distinguishing a tool that declined to
        declare anything from one that declared unboundedness.
        """
        tool = self._session_tool(tool_id)
        hook = getattr(tool, "execution_timeout", None) if tool is not None else None
        if not callable(hook):
            return False, None
        try:
            declared = hook(args)
        except Exception as exc:  # noqa: BLE001 — a ceiling must not break the call
            logging.warning("session tool {} execution_timeout failed: {}", tool_id, exc)
            return False, None
        if declared is None:
            return True, None
        if isinstance(declared, bool) or not isinstance(declared, (int, float)):
            return False, None
        return True, float(declared)

    def _partition_tool_calls(
        self, tool_calls: list[Any], specs_map: dict[str, ToolSpec]
    ) -> list[ToolBatch]:
        """Group consecutive concurrent-safe tools; isolate exclusive tools."""
        batches: list[ToolBatch] = []
        current_concurrent: list[Any] = []
        for tc in tool_calls:
            spec = specs_map.get(tc.get("name", ""))
            is_safe = spec.concurrency_safe if spec else True
            if is_safe:
                current_concurrent.append(tc)
            else:
                if current_concurrent:
                    batches.append(ToolBatch(calls=list(current_concurrent), concurrent=True))
                    current_concurrent = []
                batches.append(ToolBatch(calls=[tc], concurrent=False))
        if current_concurrent:
            batches.append(ToolBatch(calls=list(current_concurrent), concurrent=True))
        return batches

    async def _agent_def_name(self) -> str:
        """The bounded AgentDef name for this agent, for its trace span.

        Two stable literals stand in where no def name exists: ``root`` for the
        top agent, which registers no handle at all, and ``subagent`` for an
        ad-hoc spawn that named no ``agent_type``. Both are deliberate — the
        alternative is the runtime handle id, whose cardinality is one value per
        run and which the OTel agent conventions forbid recording for exactly
        that reason. `O(1)`.
        """
        if self._ctx.depth == 0:
            return "root"
        try:
            handle = await self._ctx.registry.get(self._ctx.agent_id)
        except Exception as lookup_exc:  # noqa: BLE001 — telemetry never fails a run
            logging.debug("agent def name lookup failed: {}", lookup_exc)
            return "subagent"
        agent_type = getattr(handle, "agent_type", None)
        return str(agent_type) if agent_type else "subagent"

    async def _safe_execute(
        self,
        tool_call: Any,
        tool_specs: list[ToolSpec],
    ) -> ToolCallResult:
        """Execute a tool call with timeout.

        Catches exceptions so gather does not cancel siblings.
        """
        tool_name = tool_call.get("name", "")
        timeout = self._tool_execution_timeout(tool_name, tool_call)
        # Stamp the in-flight tool BEFORE awaiting it so
        # the watchdog can attribute a stall to the tool actually running, not
        # the last one that completed (``update_step`` only fires after this
        # call returns). Cleared in ``finally`` — including on a concurrent
        # batch, where a sibling's clear can race this one; that only degrades
        # attribution to "unknown", never to a wrong tool name.
        await self._ctx.registry.mark_tool_start(self._ctx.agent_id, tool_name)
        self._emit_tool_call_event(tool_call)
        # The ONE seam every tool call passes through, so it is also the only
        # place a per-tool span covers every terminal — success, error,
        # rejection, timeout and cancellation alike. The span degrades to
        # ``None`` when Langfuse is disabled, so a deployment without it pays
        # a context-manager enter and nothing else.
        span_name, span_attributes = self._tool_span_identity(tool_name, tool_call, tool_specs)
        with langfuse_trace_span(
            span_name,
            as_type="tool",
            attributes=span_attributes,
            input_data=self._tool_span_input(tool_call),
        ) as tool_span:
            try:
                result = await asyncio.wait_for(
                    self._execute_tool_call(tool_call, tool_specs),
                    timeout=timeout,
                )
            except asyncio.TimeoutError:
                error_msg = f"Tool '{tool_name}' timed out after {timeout}s"
                logging.error(error_msg)
                # A timeout kills ``_execute_tool_call`` mid-flight, so the emit at
                # its tail never runs and the step left NO ``tool_result`` event
                # behind at all. The durable record then showed 100% tool success
                # for runs with known live timeouts — the failure was not
                # under-reported but structurally absent, and no amount of reading
                # the store could have found it. Every exit of this method emits.
                self._emit_timeout_or_crash_result(tool_call, error_msg)
                result = ToolCallResult(
                    tool_call_id=tool_call.get("id", ""),
                    tool_id=tool_name,
                    content=f"ERROR: {error_msg}",
                    success=False,
                )
            except asyncio.CancelledError:
                raise  # Must propagate for TaskGroup cancellation.
            except Exception as exc:
                # Same contract as the timeout branch: an exception that escaped
                # every inner handler would otherwise return a failed result the
                # event log has no record of.
                self._emit_timeout_or_crash_result(tool_call, str(exc))
                result = ToolCallResult(
                    tool_call_id=tool_call.get("id", ""),
                    tool_id=tool_name,
                    content=f"ERROR: {exc}",
                    success=False,
                )
            finally:
                await self._ctx.registry.mark_tool_start(self._ctx.agent_id, None)
            self._mark_tool_span_outcome(tool_span, result)
            return result

    def _tool_span_identity(
        self,
        tool_name: str,
        tool_call: Any,
        tool_specs: list[ToolSpec],
    ) -> tuple[str, dict[str, str]]:
        """Span name and OTel attributes for one tool execution.

        **The name carries the tool id and nothing else.** A tool's arguments
        are a file path, a shell command or a query, so a name built from them
        is unbounded cardinality — every call its own group — and republishes
        the argument text into every aggregation that reads the name. They ride
        the span INPUT instead, where a reader can still see them.

        For an MCP-backed tool the server NAME is what tells two servers
        exposing a same-named tool apart, and it is the only server identity
        reachable from here: ``server.address`` lives in the merged MCP config,
        which is a filesystem read per call. `O(len(tool_specs))`.
        """
        spec = next((s for s in tool_specs if s.tool_id == tool_name), None)
        attributes: dict[str, str] = {
            "gen_ai.operation.name": "execute_tool",
            "gen_ai.tool.name": tool_name,
        }
        call_id = self._tool_call_id_of(tool_call)
        if call_id:
            attributes["gen_ai.tool.call.id"] = call_id
        if spec is None or spec.kind != "mcp":
            return f"execute_tool {tool_name}", attributes
        attributes["mcp.method.name"] = "tools/call"
        server = str(spec.metadata.get("server", "") or "")
        if server:
            attributes["mcp.server.name"] = server
        remote_tool = str(spec.metadata.get("tool", "") or "")
        if remote_tool:
            attributes["mcp.tool.name"] = remote_tool
        return f"tools/call {tool_name}", attributes

    def _tool_span_input(self, tool_call: Any) -> dict[str, str]:
        """The bounded argument payload for a tool span's input.

        Windowed rather than sent whole: an edit's arguments carry a file's
        content, and exporting that per call costs the trace pipeline far more
        than the answer it buys. 2000 chars keeps both ends of the argument
        blob, which is where a path and a trailing flag live.
        """
        try:
            return {"args": self._windowed(str(tool_call.get("args", "")), 2000)}
        except Exception:  # noqa: BLE001 — telemetry never fails a tool call
            return {}

    @staticmethod
    def _mark_tool_span_outcome(span: Any, result: ToolCallResult) -> None:
        """Record a finished call's outcome on its tool span.

        Every failure terminal at this seam — a timeout, a crash, a permission
        denial, a rejected argument, an MCP tool reporting ``isError`` — arrives
        as a RESULT rather than as an exception, so nothing else would mark the
        span and a failed call would close indistinguishable from a clean one.
        Best-effort: the span is telemetry and must never be what fails a call.
        """
        if span is None or result.success:
            return
        try:
            span.update(
                level="ERROR",
                status_message=result.content[:500],
                metadata={"error.type": "tool_error"},
            )
        except Exception as span_exc:  # noqa: BLE001 — telemetry never fails a call
            logging.warning("failed to mark tool span outcome: {}", span_exc)

    def _emit_tool_call_event(self, tool_call: Any) -> None:
        """Emit one ``tool_call`` event for a call about to be dispatched.

        The transcript recorded a tool call only once it FINISHED, so every client
        showed a step's work retroactively: a call that takes a minute was
        indistinguishable from no call at all for that minute, and a call that never
        returned left nothing at all behind. This is the initiation half, emitted
        before the await rather than after it, and paired to its ``tool_result`` by
        ``tool_call_id``.

        Best-effort by construction: this runs immediately before real work, so a
        fault building the event must never prevent the call itself. `O(1)`.
        """
        try:
            action_step = self._tool_call_to_action_step(tool_call)
            self._emit_event(
                {
                    "type": "tool_call",
                    "payload": {
                        # The provider's call id, and the ONLY correlation key
                        # between this event and its result. A provider that omits
                        # one yields "", which a consumer must read as
                        # "not correlatable" rather than as a shared key.
                        "tool_call_id": self._tool_call_id_of(tool_call),
                        "tool_id": action_step.tool_id,
                        "operation": action_step.operation,
                        "tool_input": action_step.tool_input,
                        "agent_id": self._ctx.agent_id,
                        "depth": self._ctx.depth,
                        "model": self._ctx.model_name,
                    },
                }
            )
        except Exception as emit_exc:  # noqa: BLE001 — never block the call itself
            logging.warning("failed to emit tool_call event: {}", emit_exc)

    @staticmethod
    def _tool_call_id_of(tool_call: Any) -> str:
        """The provider call id of *tool_call*, or ``""`` when it carries none."""
        if isinstance(tool_call, dict):
            return str(tool_call.get("id") or "")
        return str(getattr(tool_call, "id", "") or "")

    def _emit_timeout_or_crash_result(self, tool_call: Any, error: str) -> None:
        """Emit the ``tool_result`` event for a call that died inside the wrapper.

        Best-effort by construction: this runs on a path that is ALREADY
        failing, so a fault while building the ActionStep must not replace the
        tool's real error with an event-emission error.
        """
        try:
            self._emit_tool_result_event(
                self._tool_call_to_action_step(tool_call),
                None,
                tool_call_id=self._tool_call_id_of(tool_call),
                error=error,
            )
        except Exception as emit_exc:  # noqa: BLE001 — never mask the real failure
            logging.warning("failed to emit tool_result for a failed call: {}", emit_exc)

    # ------------------------------------------------------------------
    # Message construction
    # ------------------------------------------------------------------

    def _build_messages(
        self,
        user_query: str,
        context: ContextSnapshot | None,
        plan: Plan | None,
        agent_tree: str = "",
    ) -> list[BaseMessage]:
        """Build the initial message list for the conversation."""
        system_prompt = self._render_system_prompt(context, plan, agent_tree)
        return self._messages_from_system_prompt(system_prompt, user_query, context)

    def _render_system_prompt(
        self,
        context: ContextSnapshot | None,
        plan: Plan | None,
        agent_tree: str = "",
    ) -> str:
        """Assemble the system-prompt text for the ACTIVE model.

        Extracted from ``_build_messages`` so the fallback ladder can
        re-render it against the ESCALATED model mid-run (its per-model prompt
        overrides) without rebuilding the whole transcript — see
        ``_apply_model_escalation``. Reads ``self._active_model`` (the escalated
        model after a sticky switch, else the configured primary).
        """
        registry = get_prompt_registry()
        model = self._active_model
        system_parts: list[str] = [get_system_prompt("system")]

        # Environment context.
        work_dir = self._cwd or str(Path.cwd())
        system_parts.append(
            registry.render(
                "loop.section.environment",
                model=model,
                work_dir=work_dir,
                platform=_platform.system().lower(),
                date=_date.today().isoformat(),
                version=get_version(),
            )
        )

        # Project instructions (CLAUDE.md / AGENTS.md).
        if self._project_instructions:
            system_parts.append(
                registry.render(
                    "loop.section.project_instructions",
                    model=model,
                    project_instructions=self._project_instructions,
                )
            )

        # Operator-authored custom instructions (rendered upstream, in the
        # orchestrator). Sits right after the project instructions so the
        # operator's text reads as an extension of them.
        if self._user_instructions:
            system_parts.append(
                registry.render(
                    "loop.section.user_instructions",
                    model=model,
                    user_instructions=self._user_instructions,
                )
            )

        # Root agent's global eye — live agent tree
        if agent_tree:
            system_parts.append(
                registry.render("loop.section.agent_tree", model=model, agent_tree=agent_tree)
            )

        # Git context (injected after project instructions).
        git_ctx = get_git_context(self._cwd)
        if git_ctx:
            system_parts.append(
                registry.render("loop.section.git_context", model=model, git_ctx=git_ctx)
            )

        # Active skill instructions (from user /skill invocation).
        if self._skill_instructions:
            system_parts.append(
                registry.render(
                    "loop.section.skill_instructions",
                    model=model,
                    skill_instructions=self._skill_instructions,
                )
            )

        # Resilience note (error-visibility seam): the harness's own retry /
        # fallback activity, fed back as grounded fact so the model is not blind
        # to its own retries. Its OWN slot — never overloaded onto
        # ``skill_instructions`` — and empty (skipped) on a clean run.
        resilience_note = self._resilience_note.render()
        if resilience_note:
            system_parts.append(resilience_note)

        # Auto-invocable skills catalog (for LLM-driven activation).
        if self._skill_registry is not None:
            catalog = self._skill_registry.render_catalog(self._session_capabilities)
            if catalog:
                system_parts.append(catalog)

        # Registered agent types catalog (for spawn_agent agent_type selection).
        if self._agent_registry is not None:
            agent_catalog = self._agent_registry.render_catalog(self._session_capabilities)
            if agent_catalog:
                system_parts.append(agent_catalog)

        # Session context.
        if context and context.summary:
            system_parts.append(
                registry.render(
                    "loop.section.session_summary", model=model, summary=context.summary
                )
            )
        if context and context.recent_events:
            rendered = render_event_lines(context.recent_events)
            if rendered:
                system_parts.append(
                    registry.render(
                        "loop.section.recent_conversation", model=model, rendered=rendered
                    )
                )

        # Attached file contents.
        if context and context.attachment_texts:
            system_parts.append(
                registry.render(
                    "loop.section.attached_files",
                    model=model,
                    joined="\n---\n".join(context.attachment_texts),
                )
            )

        # Tool-specific guidance from prompt files.
        tool_guidance = self._render_tool_guidance()
        if tool_guidance:
            system_parts.append(
                registry.render(
                    "loop.section.tool_guidance", model=model, tool_guidance=tool_guidance
                )
            )

        # Deferred-tool catalog (names only). Schemas are fetched on demand
        # by the model via the ``tool_search`` tool; the per-turn re-bind in
        # ``run()`` then makes the matched tools invocable.
        deferred_block = self._render_deferred_tool_block()
        if deferred_block:
            system_parts.append(deferred_block)

        # Plan context.
        if plan and plan.steps:
            plan_lines = "\n".join(
                f"{i + 1}. {s.title} — {s.description}" for i, s in enumerate(plan.steps)
            )
            system_parts.append(
                registry.render(
                    "loop.section.plan_execution", model=model, plan_lines=plan_lines
                )
            )

        # Depth-aware sub-agent guidance.
        system_parts.append(self._build_depth_guidance())

        # Plan-mode reminder — injected when the loop was started in plan
        # mode. Rendered with the session-scoped plan path and the shell
        # command allowlist so the model knows exactly what it may write
        # and which shell commands it is permitted to run.
        if self._current_mode == "plan" and self._session_id is not None:
            if self._ctx.depth == 0:
                # Root hypervisor: automata prompt + plan path for review.
                try:
                    hyper_template = registry.render("file.plan_hypervisor").strip()
                except OSError:
                    hyper_template = ""
                if hyper_template:
                    plan_path = plan_file_for(self._session_id)
                    system_parts.append(
                        hyper_template
                        + registry.render("loop.plan_file_suffix", plan_path=plan_path)
                    )
            else:
                # Plan sub-agent: full plan-mode prompt with placeholders.
                plan_path = plan_file_for(self._session_id)
                shell_allowlist = self._plan_mode_shell_allowlist()
                if shell_allowlist:
                    bullets = "\n".join(f"    - `{entry}`" for entry in shell_allowlist)
                else:
                    bullets = "    - (none — shell is disabled in plan mode)"
                system_parts.append(
                    registry.render(
                        "loop.plan_mode_reminder",
                        plan_path=plan_path,
                        shell_allowlist_bullets=bullets,
                    )
                )

        return "\n\n".join(p for p in system_parts if p)

    def _messages_from_system_prompt(
        self,
        system_prompt: str,
        user_query: str,
        context: ContextSnapshot | None,
    ) -> list[BaseMessage]:
        """Wrap the assembled system prompt + user turn into the message list."""
        # If the active context carries images for a vision-capable model,
        # build a multipart HumanMessage that interleaves the user's text
        # with ``image_url`` parts (LiteLLM/OpenAI Chat Completions format).
        # Otherwise stick with plain-string content to keep the cache
        # prefix friendly.
        image_parts = list(getattr(context, "attachment_images", []) or []) if context else []
        if image_parts:
            # langchain's HumanMessage expects ``list[str | dict]`` (invariant);
            # widen the element type so mypy accepts mixed text/image parts.
            human_content: list[str | dict] = [
                {"type": "text", "text": user_query},
                *image_parts,
            ]
            return [SystemMessage(content=system_prompt), HumanMessage(content=human_content)]
        return [SystemMessage(content=system_prompt), HumanMessage(content=user_query)]

    def _build_depth_guidance(self) -> str:
        """Build delegation-lifecycle-aware prompt guidance.

        Contract-first decomposition — root
        agents define verifiable acceptance criteria for sub-tasks.
        Manager/worker role separation — root synthesizes,
        sub-agents execute.
        Verification by same model in different role
        prevents confirmation bias.
        Liability firebreaks at chain boundaries.
        """
        depth = self._ctx.depth
        max_depth = self._ctx.max_depth
        remaining = self._ctx.remaining_depth
        is_root = depth == 0
        is_leaf = not self._ctx.can_spawn
        registry = get_prompt_registry()
        model = self._active_model

        if is_root:
            # Root = manager/hypervisor.
            # Non-blocking delegation protocol.
            # The base template carries both openings (plan vs direct execution)
            # plus the shared delegation/safety/synthesize/awareness/stop tail.
            return registry.render(
                "loop.depth.root",
                model=model,
                plan_mode=self._current_mode == "plan",
                depth=depth,
                max_depth=max_depth,
            )
        if is_leaf:
            # Liability firebreak at leaf
            return registry.render(
                "loop.depth.leaf", model=model, depth=depth, max_depth=max_depth
            )
        # Sub-orchestrator: can delegate but has bounded scope.
        guidance = registry.render(
            "loop.depth.suborchestrator",
            model=model,
            depth=depth,
            max_depth=max_depth,
            remaining=remaining,
        )
        if remaining <= 2:
            # Approaching delegation boundary
            guidance += registry.render("loop.depth.boundary", model=model)
        return guidance

    # ------------------------------------------------------------------
    # Background watchdog
    # ------------------------------------------------------------------

    async def _watchdog(self) -> None:
        """Code-level safety monitor — zero token cost.

        Periodically checks for stalled agents and injects NL warnings.
        Runs only for the root agent as a background asyncio task.

        Graduated enforcement via NL injection.
        Internal trigger: delegatee
        unresponsive → diagnose → intervene.
        """
        stall_threshold = float(get_config_value("agent", "stall_threshold_s", default=120.0))
        check_interval = float(get_config_value("agent", "stall_check_interval_s", default=30.0))
        # Deliberately well past the retry strategy's per-attempt timeout AND
        # its turn deadline: reaching this age means the bound that should have
        # cancelled the call did not, which is the only case worth an alarm.
        llm_liveness_s = float(
            get_config_value(
                "agent", "llm_call_liveness_s", default=DEFAULT_LLM_CALL_LIVENESS_S
            )
        )
        # Per-agent ids already warned of their OWN wall
        # deadline, so the 80% warning fires exactly once per agent across the
        # whole watchdog lifetime (never re-sent on every sweep).
        warned_wall_deadline: set[str] = set()
        # Steps already reported as wedged, so one wedge yields one event rather
        # than one per sweep. Re-arms naturally: the next call is a new step.
        reported_wedged_steps: set[int] = set()
        try:
            while True:
                await self._watchdog_sleep(check_interval)

                # LLM-call liveness. The stall sweep below cannot see this: it
                # keys on TOOL activity, and a wedged model call has no tool in
                # flight, so the agent looks merely quiet. Emitting makes the
                # wedge visible while it is happening instead of leaving an
                # ``llm_call_start`` with no end for a boot-time sweeper to find.
                age = self.llm_call_stall_age(_time.monotonic())
                if (
                    age is not None
                    and llm_liveness_s > 0
                    and age >= llm_liveness_s
                    and self._llm_call_step not in reported_wedged_steps
                ):
                    reported_wedged_steps.add(self._llm_call_step)
                    self._emit_event(
                        {
                            "type": "llm_call_stalled",
                            "payload": {
                                "agent_id": self._ctx.agent_id,
                                "depth": self._ctx.depth,
                                "step": self._llm_call_step,
                                "model": self._active_model,
                                "age_s": round(age, 1),
                                "threshold_s": llm_liveness_s,
                            },
                        }
                    )

                stalled = await self._ctx.registry.stalled_agents(threshold=stall_threshold)
                for h in stalled:
                    await self._ctx.registry.send_message(
                        h.agent_id,
                        get_prompt_registry().render("loop.stall_warning"),
                    )
                    # Root self-stall: this loop IS the one synchronously
                    # awaiting the stalled tool call, so a line queued here can
                    # only be drained AFTER that call resolves — stale and
                    # misleading by the time it arrives. Never enqueue it for
                    # self; the send_message warning above still reaches
                    # genuinely separate sub-agents (their own message_queue is
                    # drained mid-run by their own loop, not this one).
                    if h.agent_id == self._ctx.agent_id:
                        continue
                    if self._ctx.message_queue is not None:
                        # active_tool_id is the tool actually IN FLIGHT at
                        # detection time; last_tool_id (only updated on
                        # completion) would name the wrong, already-finished
                        # tool for a call still running.
                        detail = (
                            f"stalled on {h.active_tool_id}"
                            if h.active_tool_id
                            else f"stalled — no tool activity for over {stall_threshold:.0f}s"
                        )
                        self._ctx.message_queue.put_nowait(
                            f"[Watchdog: Agent {h.agent_id[:8]} {detail}]",
                        )

                # Per-contract wall-clock deadline sweep.
                # Mirrors the stall sweep above but keyed on an agent's OWN
                # declared bound, never global inactivity. Two-strike by
                # design: the 80% warn ALWAYS precedes the 100% cancel, so
                # cancel_agent here is the ONE contract-driven kill switch in
                # the codebase — every other agent termination is either
                # graceful (natural completion) or parent-requested
                # (steer_agent cancel).
                for h, wall_state in await self._ctx.registry.over_wall_deadline_agents():
                    if wall_state == "over":
                        await self._ctx.registry.cancel_agent(h.agent_id)
                        continue
                    if h.agent_id in warned_wall_deadline:
                        continue
                    warned_wall_deadline.add(h.agent_id)
                    await self._ctx.registry.send_message(
                        h.agent_id,
                        get_prompt_registry().render("loop.wall_deadline_warning"),
                    )
        except asyncio.CancelledError:
            pass  # Normal shutdown path.

    def _render_tool_guidance(self) -> str:
        """Render tool-specific prompt guidance for local tools.

        Routes through the prompt registry WITH the active model so a per-model
        override of a tool-guidance entry (e.g. the structured-patch discipline
        nudge on ``file.tools.file-edit``, paired with the edit-tool variant via
        the shared model prefix) actually applies. ``get_system_prompt``
        alone can't: a tool's ``prompt_path`` (``tools/file-edit``) maps to the
        dotted registry id ``file.tools.file-edit``, so the slash form misses the
        registry and would silently read the raw ``.txt``, dropping ``model=``.
        Falls back to a raw file read for any prompt the registry doesn't
        inventory.
        """
        registry = get_prompt_registry()
        prompts: list[str] = []
        for spec in self._tool_registry.list_specs():
            if spec.kind != "local" or not spec.prompt_path:
                continue
            try:
                reg_id = "file." + spec.prompt_path.replace("/", ".")
                if registry.has(reg_id):
                    prompt = registry.render(reg_id, model=self._active_model).strip()
                else:
                    prompt = get_system_prompt(spec.prompt_path)
            except OSError:
                continue
            if prompt:
                prompts.append(prompt)
        return "\n\n".join(prompts)

    # ------------------------------------------------------------------
    # Skill activation (internal tool handler)
    # ------------------------------------------------------------------

    def _handle_activate_skill(self, action_step: ActionStep) -> Any:
        """Handle an ``activate_skill`` tool call from the LLM.

        Returns the skill body as a mock result so it arrives as a
        ``ToolMessage`` — the LLM reads the instructions and follows them.
        """
        from mewbo_core.tooling.skills import activate_skill

        args = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
        skill_name = str(args.get("skill_name", ""))
        skill_args = str(args.get("args", ""))

        registry = self._skill_registry
        skill = (
            registry.get(skill_name, self._session_capabilities) if registry else None
        )
        if skill is None:
            msg = f"ERROR: Unknown skill '{skill_name}'"
            return type("R", (), {"content": msg})()
        if skill.disable_model_invocation:
            msg = f"ERROR: Skill '{skill_name}' is user-invocable only"
            return type("R", (), {"content": msg})()

        instructions, _ = activate_skill(skill, skill_args, cwd=self._cwd)
        logging.info("LLM auto-activated skill '{}'", skill_name)
        body = f"## Skill: {skill_name}\n\n{instructions}\n\nFollow these instructions now."
        return type("R", (), {"content": body})()

    # ------------------------------------------------------------------
    # Model binding
    # ------------------------------------------------------------------

    def _configured_edit_tool_id(self) -> str:
        """Return the tool_id for the edit tool appropriate for the active model.

        Prefers model-derived capability detection (via
        ``llm.model_prefers_structured_patch``).  The ``agent.edit_tool``
        config value is honoured as an explicit override when non-empty.
        """
        from mewbo_core.llm.llm import model_prefers_structured_patch

        # Explicit user override always wins
        override = get_config_value("agent", "edit_tool", default="")
        if override == "structured_patch":
            return "file_edit_tool"
        if override == "search_replace_block":
            return "aider_edit_block_tool"

        # Derive from the ACTIVE model identity (escalated model after a
        # sticky switch, else the configured primary) so the tool VARIANT adapts
        # in lockstep with the per-model prompt on escalation.
        model_name: str | None = getattr(self, "_active_model", None) or getattr(
            self._ctx, "model_name", None
        )
        if model_prefers_structured_patch(model_name):
            return "file_edit_tool"
        return "aider_edit_block_tool"

    def _plan_mode_shell_allowlist(self) -> list[str]:
        """Return the configured shell command prefix allowlist for plan mode.

        Read from ``agent.plan_mode_shell_allowlist``. An empty list means
        the shell tool is disabled in plan mode.
        """
        raw = get_config_value("agent", "plan_mode_shell_allowlist", default=[])
        if isinstance(raw, list):
            return [str(item).strip() for item in raw if str(item).strip()]
        if isinstance(raw, str):
            return [entry.strip() for entry in raw.split(",") if entry.strip()]
        return []

    # ------------------------------------------------------------------
    # LLM call resilience (retry / fallback / circuit-break / budget)
    # ------------------------------------------------------------------

    @staticmethod
    def _chunk_delta_text(chunk: Any) -> str:
        """Incremental text carried by one streamed chunk (whitespace-preserving).

        Unlike ``_extract_text_content`` this never strips — token deltas must
        keep their spacing — and concatenates list-form text blocks without the
        newline separator (a chunk is a fragment, not a finished message).
        """
        content = getattr(chunk, "content", "")
        if isinstance(content, str):
            return content
        if isinstance(content, list):
            return "".join(
                block.get("text", "")
                for block in content
                if isinstance(block, dict) and block.get("type") == "text"
            )
        return ""

    async def _acall_model(
        self,
        bound: Any,
        messages: list[BaseMessage],
        *,
        config: dict[str, Any] | None,
        step: int,
    ) -> AIMessage:
        """Invoke the model, streaming token deltas for true time-to-first-token.

        Consumes ``bound.astream`` and emits one ``agent_message_delta`` event per
        text chunk so clients render tokens as the model produces them, then
        returns the fully-accumulated message — content, ``tool_calls`` and
        ``usage_metadata`` are identical in shape to ``ainvoke`` (LangChain's
        ``AIMessageChunk.__add__`` aggregates tool-call chunks and sums usage;
        ``ChatLiteLLM._astream`` already requests ``stream_options.include_usage``
        so the usage chunk arrives). Falls back to ``ainvoke`` for a model that
        exposes no usable stream at all (a stubbed model), so a non-streaming
        caller is unaffected — but NEVER for a real stream that simply produced
        nothing, which is a failed rung the ladder must see rather than a second
        billed call. Real provider errors mid-stream propagate to the resilience
        strategy; only the structural "no usable stream" signals are swallowed.
        """
        cfg = config or None
        accumulated: AIMessageChunk | None = None
        # "This object has no async stream" and "a real stream produced nothing"
        # are DIFFERENT facts, and reading both off ``accumulated is None``
        # conflated them: a stream that yielded zero chunks silently issued a
        # second full call — billed again, counted as no attempt, given no
        # backoff, recorded in no event. Two independent signals separate them,
        # because neither alone is sufficient.
        #
        # ``streams`` is the CAPABILITY probe, and it decides what zero chunks
        # MEAN. Every real client answers True (``BaseChatModel.astream`` and
        # the ``RunnableBinding`` returned by ``bind_tools`` are both async
        # generator functions), while a ``MagicMock`` — whose ``__aiter__``
        # yields an empty sequence rather than raising — answers False. Without
        # it a stubbed model is indistinguishable from a provider that returned
        # an empty stream, and the safe reading of that ambiguity is the
        # buffered call, never a failed rung invented from a test double.
        streams = inspect.isasyncgenfunction(getattr(bound, "astream", None))
        stream_unusable = False
        try:
            async for chunk in bound.astream(messages, config=cfg):
                accumulated = chunk if accumulated is None else accumulated + chunk
                delta = self._chunk_delta_text(chunk)
                if delta:
                    self._emit_event(
                        {
                            "type": "agent_message_delta",
                            "payload": {
                                "text": delta,
                                "agent_id": self._ctx.agent_id,
                                "depth": self._ctx.depth,
                                "step": step,
                            },
                        }
                    )
        except (TypeError, AttributeError, NotImplementedError):
            # ``bound`` exposes no usable async stream (e.g. a stubbed model) —
            # fall through to the buffered path. Provider/runtime errors are NOT
            # caught here; they belong to the resilience strategy.
            if accumulated is not None:
                # Chunks had already arrived, so this is a fault PARTWAY through
                # a working stream, not a model without one. Re-issuing buffered
                # here would re-send a request whose partial stream may already
                # carry a tool call — and the buffered reply would carry it
                # again, executing the same tool twice.
                raise
            stream_unusable = True
        if accumulated is None:
            if stream_unusable or not streams:
                return await bound.ainvoke(messages, config=cfg)
            # A real stream opened and closed without one chunk. It is reported
            # with an explicit zero-token usage rather than a bare empty message,
            # because that is literally what happened — nothing was generated —
            # and because it is the ONE signal the result contract convicts on.
            # A bare empty message would carry no usage at all, which that
            # contract deliberately reads as "not proven" and would let through.
            # The ladder decides from there, rather than this seam quietly buying
            # a second call the ladder never learns about.
            return AIMessage(
                content="",
                usage_metadata=UsageMetadata(
                    input_tokens=0, output_tokens=0, total_tokens=0
                ),
            )
        return AIMessage(
            content=accumulated.content,
            tool_calls=list(getattr(accumulated, "tool_calls", []) or []),
            additional_kwargs=getattr(accumulated, "additional_kwargs", {}) or {},
            response_metadata=getattr(accumulated, "response_metadata", {}) or {},
            usage_metadata=getattr(accumulated, "usage_metadata", None),
            id=getattr(accumulated, "id", None),
        )

    async def _capture_usage(self, response: AIMessage) -> None:
        """Accumulate token usage + the compaction anchor from a response."""
        usage = getattr(response, "usage_metadata", None)
        if not usage:
            return
        handle = await self._ctx.registry.get(self._ctx.agent_id)
        if handle:
            handle.input_tokens += usage.get("input_tokens", 0)
            handle.output_tokens += usage.get("output_tokens", 0)
        # Authoritative signal for compaction: what the API said this consumed.
        self._last_input_tokens = int(usage.get("input_tokens", 0) or 0)

    async def _invoke_with_resilience(
        self,
        *,
        primary_model: Any,
        messages: list[BaseMessage],
        tool_schemas: Any,
        turns: int,
        invoke_config: dict[str, Any] | None,
        strategy: RetryStrategy,
    ) -> tuple[AIMessage, str]:
        """Drive the resilience strategy for one turn with loop-local I/O.

        The model call, event emission and reactive compaction are injected so
        the strategy stays loop-agnostic and unit-testable. ``messages.append``
        is the caller's job and only happens after this returns, so a retry
        never duplicates a tool call or bloats context with a partial output.
        """

        async def _invoke(model_name: str, is_fallback: bool) -> AIMessage:
            # Re-arm the liveness leg per ATTEMPT, not per turn. The detector asks
            # "is THIS provider read wedged below the event loop", so measuring
            # from the turn's first attempt made a healthy retry sequence — two
            # capped attempts that each timed out cleanly — read as one stalled
            # call and drown the genuine wedge it exists to catch.
            self._llm_call_started_at = _time.monotonic()
            # ``primary_model`` is pre-bound to the ACTIVE model (the configured
            # primary, or — after a sticky escalation — the pinned escalated
            # model, since ``_apply_model_escalation`` rebinds it). Key the reuse
            # off the model NAME, not ``is_fallback``: sticky escalation reorders
            # a pinned rescue model to idx 0 (so ``is_fallback`` is False) —
            # reusing a stale binding there would silently call the wrong model.
            # Rebind through ``_bind_model`` — the only binder that appends the
            # directly-bound legs. Binding ``tool_schemas`` here directly would
            # bind the caller's REGISTRY-ONLY list, dropping the spawn family,
            # the session tools and ``activate_skill`` for exactly the rung a
            # struggling run most needs them on.
            bound = (
                primary_model
                if not is_fallback and model_name == self._active_model
                else self._bind_model(tool_schemas, model_name=model_name)
            )
            return await self._acall_model(bound, messages, config=invoke_config, step=turns)

        async def _compact() -> bool:
            info = await self._compact_messages(messages)
            if not info:
                return False
            await self._ctx.registry.record_compaction(self._ctx.agent_id)
            self._emit_event(
                {
                    "type": "context_compacted",
                    "payload": {
                        **info,
                        "agent_id": self._ctx.agent_id,
                        "depth": self._ctx.depth,
                        "mode": "reactive",
                        "turn": turns,
                    },
                }
            )
            return True

        # Chain head is the ACTIVE model, not the frozen configured one. Once a
        # sticky escalation has promoted ``_active_model``, re-offering
        # ``_ctx.model_name`` puts the model that just died back at the head of
        # every subsequent turn's chain; escalation survived that only because
        # ``RetryStrategy._order_models`` reorders on its own pin, i.e. by
        # accident of a second mechanism rather than by this call being right.
        # Before any escalation the two are identical, so the ordinary path is
        # unchanged.
        response, model_name = await strategy.run(
            models=[self._active_model, *self._ctx.fallback_models],
            invoke=_invoke,
            emit=self._emit_event,
            compact=_compact,
            agent_id=self._ctx.agent_id,
            depth=self._ctx.depth,
            step=turns,
        )
        await self._capture_usage(response)
        return response, model_name

    def _request_model_switch(self, target: str) -> None:
        """Record a deliberate model switch the ``model_control`` tool sanctioned.

        Deferred, never applied inline: the loop promotes the model at the next
        turn boundary through :meth:`_apply_model_escalation`, so the switch
        lands exactly like a sticky fallback (transcript tail untouched) — the
        tool-loop-continuity contract the tool already enforced at request time.
        """
        self._requested_switch = target

    def _has_dangling_tool_use(self) -> bool:
        """True when a PRIOR turn's ``tool_use`` is still unanswered.

        The continuity-lock guardrail for ``model_control``: switching models
        while a tool call sits unpaired would carry the dangling pair into the
        first request on the new model. The in-flight batch (the most recent
        ``AIMessage``'s tool_calls, whose results are being produced right now)
        is excluded — counting it would make every switch look unsafe; only an
        EARLIER unpaired call, the kind a compaction slice can strand, blocks a
        switch.
        """
        messages = self._live_messages or []
        answered = {
            m.tool_call_id for m in messages if isinstance(m, ToolMessage)
        }
        ai_batches = [m for m in messages if isinstance(m, AIMessage) and m.tool_calls]
        for m in ai_batches[:-1]:  # exclude the in-flight (last) batch
            for tc in m.tool_calls:
                tcid = tc.get("id") if isinstance(tc, dict) else getattr(tc, "id", None)
                if tcid and tcid not in answered:
                    return True
        return False

    def _apply_model_escalation(
        self,
        final_model: str,
        messages: list[BaseMessage],
        *,
        context: ContextSnapshot | None,
        plan: Plan | None,
        agent_tree: str,
        tool_schemas: list[dict[str, Any]],
        model: Any,
    ) -> tuple[list[dict[str, Any]], Any]:
        """Promote the active model to an escalated one and re-derive prompts.

        No-op (returns the inputs unchanged) unless the resilience strategy ended
        the turn on a DIFFERENT model than the one currently active — i.e. a
        sticky escalation down the fallback ladder. On a real switch it:

        - promotes ``self._active_model`` so every subsequent per-step render +
          the edit-tool variant selection resolve the ESCALATED model's overrides;
        - re-renders ``messages[0]`` (the system prompt) in place against it, so
          the next turn carries that model's compatibility-adjusted prompt;
        - recomputes the bound tool schemas (the edit-tool VARIANT can differ
          per model) + rebinds, mirroring the deferred-tool re-bind path.

        The transcript tail is untouched; the change takes effect next turn.
        """
        if not final_model or final_model == self._active_model:
            return tool_schemas, model
        self._active_model = final_model
        # Re-seat the spawn seam on the healed model. An un-overridden child
        # resolves its model from the spawn tool's captured context — the model
        # this loop just escalated AWAY from — so a self-healed parent would
        # otherwise fan its children onto the dead one. The re-seat goes through
        # the tool's own ``rebind_active_model`` seam: that ``AgentContext`` is
        # frozen is the tool's business, not the loop's.
        if self._spawn_agent_tool is not None:
            self._spawn_agent_tool.rebind_active_model(final_model)
        messages[0] = SystemMessage(
            content=self._render_system_prompt(context, plan, agent_tree)
        )
        # This re-render already baked in the current resilience-note slot, so
        # keep the tracker honest — else the top-of-loop refresh re-renders once
        # more for a note that is already present.
        self._active_resilience_note = self._resilience_note.render()
        discovered = self._discovered_from_messages(messages)
        active_specs = self._select_active_specs(self._tool_specs_full, discovered=discovered)
        tool_schemas = self._build_tool_schemas_for_mode(active_specs, self._current_mode)
        model = self._bind_model(tool_schemas)
        self._last_active_ids = {s.tool_id for s in active_specs}
        self._emit_event(
            {
                "type": "llm_prompt_revariant",
                "payload": {
                    "agent_id": self._ctx.agent_id,
                    "depth": self._ctx.depth,
                    "model": final_model,
                    "edit_tool": self._configured_edit_tool_id(),
                },
            }
        )
        return tool_schemas, model

    # ------------------------------------------------------------------
    # Graduated exhaustion — forced wrap-up turns
    # ------------------------------------------------------------------

    async def _unbound_wrapup_invoke(
        self,
        *,
        prompt_key: str,
        model_name: str,
        messages: list[BaseMessage],
        tool_outputs: list[str],
        invoke_config: dict[str, Any],
    ) -> str:
        """One text-only LLM call with tools UNBOUND — never raises.

        Falls back to recent tool output on failure. Shared by the
        (currently unreachable, defensive) final-synthesis safety net and
        :meth:`_budget_wrapup_turn`; the only difference between callers is
        which directive gets injected and which model answers it.
        """
        final_response = ""
        try:
            messages.append(SystemMessage(content=get_prompt_registry().render(prompt_key)))
            unbound = build_chat_model(model_name=model_name)
            synthesis_timeout = float(get_config_value("agent", "llm_call_timeout", default=60.0))
            synthesis: AIMessage = await asyncio.wait_for(
                unbound.ainvoke(messages, config=invoke_config or None),
                timeout=synthesis_timeout,
            )
            _usage = getattr(synthesis, "usage_metadata", None)
            if _usage:
                _h = await self._ctx.registry.get(self._ctx.agent_id)
                if _h:
                    _h.input_tokens += _usage.get("input_tokens", 0)
                    _h.output_tokens += _usage.get("output_tokens", 0)
            final_response = self._extract_text_content(getattr(synthesis, "content", ""))
        except Exception:
            logging.warning("Unbound wrap-up turn failed.", exc_info=True)

        # Fallback: if the call produced nothing, build a minimal response
        # from successful tool outputs so the caller isn't left empty-handed.
        if not final_response or not final_response.strip():
            successful = [o for o in tool_outputs if o and not o.startswith("ERROR")]
            if successful:
                final_response = "Partial results:\n\n" + "\n\n".join(successful[-5:])
        return final_response

    async def _budget_wrapup_turn(
        self,
        reason: str,
        *,
        state: OrchestrationState,
        messages: list[BaseMessage],
        tool_outputs: list[str],
        invoke_config: dict[str, Any],
    ) -> str:
        """Force one text-only wrap-up call in place of a bare halt.

        Shared by every graduated-exhaustion caller: the session step budget
        passes ``reason="budget_exhausted"``, the per-agent contract passes
        ``"halted_agent_budget"``. Runs against ``self._active_model`` (the
        escalated model if a sticky switch pinned one) — unlike the
        final-synthesis safety net, which answers on ``_ctx.model_name``.
        Deliberately skips
        ``registry.update_step`` — this closing call must not re-trip the very
        budget that triggered it. Sets ``state.done``/``state.done_reason``;
        the caller's ``break`` is the only remaining step, and the (already
        done-guarded) final-synthesis safety net is skipped as a result.
        """
        final_response = await self._unbound_wrapup_invoke(
            prompt_key="loop.budget_exhausted_wrapup",
            model_name=self._active_model,
            messages=messages,
            tool_outputs=tool_outputs,
            invoke_config=invoke_config,
        )
        state.done = True
        state.done_reason = reason
        return final_response

    async def _run_verifier(self, *, step: int, attempt: int) -> VerifierOutcome:
        """Run the ground-truth completion verifier once and emit its verdict.

        The model does no I/O: the injected runner performs the subprocess
        (never raising — a timeout/OS error comes back as a failed result), the
        spec interprets that raw result into a pass/fail ``VerifierOutcome``, and
        this method emits the BOUNDED telemetry (``verification`` event, scalars
        only — the grounded stdout/stderr rides only the outcome's feedback into
        the model's own context, never onto the wire). Guarded by
        ``self._verification_active`` at the sole call site, so the spec is
        non-None here.
        """
        assert self._verification is not None  # _verification_active guards this
        result = await self._verifier_runner.run(self._verification, cwd=self._cwd)
        outcome = self._verification.interpret(result)
        payload: VerificationPayload = {
            "agent_id": self._ctx.agent_id,
            "depth": self._ctx.depth,
            "step": step,
            "passed": outcome.passed,
            "attempt": attempt,
            "exit_code": outcome.exit_code,
            "timed_out": outcome.timed_out,
        }
        self._emit_event({"type": "verification", "payload": payload})
        return outcome

    # ------------------------------------------------------------------
    # Mid-loop context compaction
    # ------------------------------------------------------------------

    def _should_compact_messages(self, messages: list[BaseMessage]) -> bool:
        """Token-reality check anchored on the last response's usage_metadata.

        ``messages`` is intentionally unused — the LLM's ``usage_metadata``
        already reflects everything it saw (system prompt, tool schemas,
        all messages). No char-count estimation, no manual overhead.
        """
        del messages  # Signature preserved for callers; data comes from the API.
        if self._last_input_tokens <= 0:
            return False  # No LLM call yet; nothing authoritative to check.
        from mewbo_core.session.token_budget import get_model_max_input_tokens

        # Size the window against the ACTIVE model. Reading the frozen
        # configured one keeps a large window after escalating DOWN to a small
        # model, so the threshold sits above that model's real ceiling,
        # auto-compaction never fires, and the run dies on
        # ContextWindowExceededError instead of compacting.
        max_input = get_model_max_input_tokens(self._active_model)
        threshold = float(get_config_value("token_budget", "auto_compact_threshold", default=0.8))
        return self._last_input_tokens >= max_input * threshold

    async def _compact_messages(
        self,
        messages: list[BaseMessage],
    ) -> dict[str, Any] | None:
        """Compact the in-flight message list by summarizing older messages.

        Keeps messages[0] (system prompt) and the last ``recent_keep``
        messages, summarizing everything in between via the structured
        compaction prompt.  Returns a dict with ``summary`` and
        ``events_summarized`` on success, or ``None`` if skipped.
        """
        recent_keep = 6  # ~3 turn pairs (AI + Tool)
        if len(messages) <= recent_keep + 2:
            return None  # Not enough to compact

        to_summarize = messages[1:-recent_keep]
        kept_tail = messages[-recent_keep:]

        # Take stale images out of the tail we are KEEPING. Compaction is the
        # one moment this is free: the message list is being rewritten anyway,
        # so the prompt-cache prefix is already void and stripping costs no
        # additional invalidation. The newest image survives — a run driving a
        # phone that loses sight of the current screen has to spend a turn
        # re-observing it. The rest become a placeholder that says the image is
        # re-requestable, which for a screenshot is the only honest recovery:
        # the screen has moved on, so a retained copy would answer a question
        # about the CURRENT screen with a stale picture.
        # (The summarized half needs no strip — it is replaced by text.)
        images_stripped = self._image_history.strip(kept_tail)

        # Build text representation for the summarizer.
        lines: list[str] = []
        for m in to_summarize:
            role = getattr(m, "type", "unknown")
            # A multimodal message's content is a LIST of parts, and ``str()``
            # on it would inline a whole base64 data URI into the summarizer's
            # prompt — paying for an image the summarizer cannot see, then
            # cutting it at 2000 chars. Keep the text parts only.
            text = m.content if isinstance(m.content, str) else _text_of_parts(m.content)
            lines.append(f"[{role}] {text[:2000]}")
        summary_input = "\n".join(lines)

        # If root agent, include agent tree so delegation context survives.
        if self._ctx.depth == 0:
            tree = await self._ctx.registry.render_agent_tree(
                exclude_agent_id=self._ctx.agent_id,
            )
            if tree:
                summary_input += f"\n\n# Active agent tree at compaction:\n{tree}"

        # Invoke the compaction LLM with priority-ordered model fallback.
        from mewbo_core.session.compact import (
            _extract_summary,
            get_compact_prompt,
            resolve_compact_models,
        )

        _compact_models = resolve_compact_models(self._active_model)
        _compact_model = _compact_models[0]
        _msgs = [
            SystemMessage(content=get_compact_prompt(model=_compact_model)),
            HumanMessage(
                content=get_prompt_registry().render(
                    "loop.compaction_drive", summary_input=summary_input
                )
            ),
        ]
        response = None
        for _i, _candidate in enumerate(_compact_models):
            _compact_model = _candidate
            try:
                llm = build_chat_model(model_name=_compact_model)
                response = await llm.ainvoke(_msgs)
                break
            except Exception:
                if _i < len(_compact_models) - 1:
                    logging.warning(
                        "Mid-loop compact model %s failed, trying next: %s",
                        _compact_model,
                        _compact_models[_i + 1],
                        exc_info=True,
                    )
                else:
                    logging.warning("Mid-loop compaction LLM call failed", exc_info=True)
                    return None

        if response is None:
            return None  # all models failed

        # Capture compaction LLM tokens on the agent handle.
        _usage = getattr(response, "usage_metadata", None)
        if _usage:
            _h = await self._ctx.registry.get(self._ctx.agent_id)
            if _h:
                _h.input_tokens += _usage.get("input_tokens", 0)
                _h.output_tokens += _usage.get("output_tokens", 0)

        summary = _extract_summary(response_text(response))
        events_summarized = len(to_summarize)

        # Rebuild messages in-place.
        system_msg = messages[0]
        messages.clear()
        messages.append(system_msg)
        messages.append(
            SystemMessage(
                content=get_prompt_registry().render("loop.compacted_marker", summary=summary)
            )
        )
        messages.extend(kept_tail)
        # The recent-tail slice can orphan a tool_use/tool_result pair, which
        # Anthropic rejects with a 400. Rebalance before the list is replayed.
        repair_tool_pairing(messages)
        info: dict[str, Any] = {
            "summary": summary,
            "events_summarized": events_summarized,
            "model": _compact_model,
        }
        # Reported only when it happened, so every existing compaction event
        # stays byte-identical and no consumer has to learn a new always-zero
        # field. Stripping images silently would make a run that lost its
        # screenshots indistinguishable from one that never took any.
        if images_stripped:
            info["images_stripped"] = images_stripped
        return info

    def _is_tool_search_enabled(self, tool_specs: list[ToolSpec] | None = None) -> bool:
        """Return True if the deferred-tool / on-demand-schema feature is on.

        Read fresh from config so the field can be flipped without a
        process restart. Sub-orchestrators inherit by reading the same
        ``agent.tool_search.mode`` value.

        - ``on`` (the default) → always defer.
        - ``off`` → never defer.
        - ``auto`` → defer only when the number of deferrable specs in
          ``tool_specs`` exceeds ``agent.tool_search.auto_threshold``.
          When called without ``tool_specs`` (no set to measure) ``auto``
          conservatively stays off.

        The ``default=`` below must track ``ToolSearchConfig.mode``'s own
        default. It is unreachable while the field exists — ``get_config``
        returns the validated ``AppConfig``, so the ``getattr`` walk yields
        the model default and never falls through — but a stale literal here
        reads as the real default to anyone tracing this path, and it said
        ``off`` for the whole life of the feature.
        """
        mode = str(get_config_value("agent", "tool_search", "mode", default="on")).lower()
        if mode == "on":
            return True
        if mode == "auto":
            if not tool_specs:
                return False
            raw_threshold = get_config_value("agent", "tool_search", "auto_threshold", default=25)
            try:
                threshold = int(raw_threshold)
            except (TypeError, ValueError):
                threshold = 25
            deferrable = sum(1 for s in tool_specs if is_deferred(s))
            return deferrable > threshold
        return False

    def _select_active_specs(
        self,
        specs: list[ToolSpec],
        *,
        discovered: set[str],
    ) -> list[ToolSpec]:
        """Return the spec subset to bind on the model this turn.

        ``non_deferred ∪ {tool_search} ∪ (deferred ∩ discovered)``. When
        deferred-loading is off, returns ``specs`` unchanged. The same
        function drives both the run-start bind and the per-turn re-bind
        so there is exactly one source of truth for what is bound.

        ``tool_search`` reaches here having been EXEMPTED from the allowlist by
        ``filter_specs`` (it is ``always_load``). That exemption is load-bearing
        and stays: strip it while tools are deferred and a scoped sub-agent
        loses its MCP schemas AND the only means to fetch them. It is also
        justified ONLY while something is deferred — with nothing to fetch, the
        tool is pure surface area, and it was among the top repeat callees in
        the runs that burned a session's whole budget without answering. So a
        STRICT scope that did not name it caps it in exactly that case.
        """
        deferral_active = self._tool_search_enabled and bool(self._deferred_ids)
        if not deferral_active and not self._loop_injected_admitted(TOOL_SEARCH_TOOL_ID):
            specs = [s for s in specs if s.tool_id != TOOL_SEARCH_TOOL_ID]
        if not deferral_active:
            return list(specs)
        keep: list[ToolSpec] = []
        for spec in specs:
            if spec.tool_id in self._deferred_ids and spec.tool_id not in discovered:
                continue
            keep.append(spec)
        return keep

    # Match each ``<function>{...}</function>`` line emitted by ToolSearchRunner.
    # The runner serialises ``{"name": ..., "description": ..., "parameters": {...}}``
    # with ``json.dumps`` so ``"name"`` is always the first field; anchoring there
    # avoids brace-counting through the nested ``parameters`` schema.
    _DISCOVERED_FUNC_RE = re.compile(r'<function>\s*\{\s*"name"\s*:\s*"([^"]+)"')

    def _discovered_from_messages(self, messages: list[BaseMessage]) -> set[str]:
        """Scan message history for tool names exposed by past tool_search calls.

        Discovery is derived from messages — no separate state — so the
        set survives compaction unchanged: whatever messages remain after
        compaction still parse the same way. Only ``ToolMessage`` content
        is scanned; the regex matches the ``<function>...</function>``
        line format produced by ``ToolSearchRunner``.
        """
        if not self._deferred_ids:
            return set()
        names: set[str] = set()
        for msg in messages:
            if not isinstance(msg, ToolMessage):
                continue
            content = msg.content
            if isinstance(content, str):
                for match in self._DISCOVERED_FUNC_RE.finditer(content):
                    name = match.group(1)
                    if name in self._deferred_ids:
                        names.add(name)
        return names

    # Budget for the tool-id detail in the deferred catalog, in characters.
    # An id is ~5 tokens against a schema's ~240, so naming every deferred tool
    # costs a fraction of binding one — but a pathological fleet still has to be
    # bounded somewhere. When the budget runs out the remaining servers degrade
    # to the name+count summary: they stay VISIBLE and the block SAYS their ids
    # were dropped, because a silently truncated catalog reads to the model as a
    # complete one and it will conclude a tool does not exist.
    _DEFERRED_ID_CHAR_BUDGET = 6000

    def _render_deferred_tool_block(self) -> str:
        """Render the ``<available-mcp-servers>`` system-prompt catalog.

        Lists each MCP server with its tool COUNT **and its tool ids**, plus any
        non-MCP deferred tool ids. Naming the ids is the point: the model can
        then issue ``tool_search`` with ``select:<tool_id>`` and get the schema
        in one call. Listing servers alone forced a discovery round-trip whose
        only product was a list of names we already held — a wasted turn on
        every session that touched an MCP tool, and a turn the model sometimes
        spent guessing ids instead.

        Bounded by ``_DEFERRED_ID_CHAR_BUDGET``; overflow degrades to
        counts-only for the remaining servers and says so.
        """
        if not getattr(self, "_tool_search_enabled", False):
            return ""
        deferred_ids: set[str] = getattr(self, "_deferred_ids", set())
        if not deferred_ids:
            return ""
        deferred_specs = [s for s in self._tool_specs_full if s.tool_id in deferred_ids]

        servers: dict[str, list[str]] = {}
        other: list[str] = []
        for spec in deferred_specs:
            if spec.kind == "mcp":
                server = str(spec.metadata.get("server") or "unknown")
                servers.setdefault(server, []).append(spec.tool_id)
            else:
                other.append(spec.tool_id)

        parts: list[str] = []
        if servers:
            lines: list[str] = []
            elided: list[str] = []
            spent = 0
            for name, ids in sorted(servers.items()):
                body = ", ".join(sorted(ids))
                # Always detail the first server even if it alone overruns the
                # budget — a catalog naming no ids at all is worse than a long one.
                if lines and spent + len(body) > self._DEFERRED_ID_CHAR_BUDGET:
                    lines.append(f"{name} ({len(ids)})")
                    elided.append(name)
                    continue
                spent += len(body)
                lines.append(f"{name} ({len(ids)}): {body}")
            summary = "\n".join(lines)
            parts.append(f"<available-mcp-servers>\n{summary}\n</available-mcp-servers>")
            if elided:
                parts.append(
                    f"Tool ids under {', '.join(elided)} are omitted for length — "
                    "search those by server name or keyword."
                )
        if other:
            parts.append(f"Other deferred tools: {', '.join(sorted(other))}.")
        parts.append(get_prompt_registry().render("loop.section.deferred_tools"))
        return "\n".join(parts)

    def _build_tool_schemas_for_mode(
        self,
        specs: list[ToolSpec],
        mode: str,
    ) -> list[dict[str, Any]]:
        """Return langchain tool schemas appropriate for ``mode``.

        In plan mode the schema is filtered to: read-only tools + the
        configured edit tool (path-scoped at the permission layer) + the
        shell tool (command-allowlisted at the permission layer, iff the
        allowlist is non-empty) + all MCP tools. MCP specs carry no
        read-only signal (the wire protocol exposes no ``readOnlyHint``),
        so Mewbo cannot classify a third-party MCP tool's effect — a mode
        filter over user-land tools would be guesswork, not gating. In act
        mode all specs pass through.
        """
        if mode != "plan":
            return specs_to_langchain_tools(specs)
        edit_tool_id = self._configured_edit_tool_id()
        shell_enabled = bool(self._plan_mode_shell_allowlist())
        filtered = [
            spec
            for spec in specs
            if spec.read_only
            or spec.kind == "mcp"
            or (spec.tool_id == edit_tool_id and self._ctx.depth > 0)
            or (shell_enabled and spec.tool_id in SHELL_TOOL_IDS)
        ]
        return specs_to_langchain_tools(filtered)

    def _directly_bound_tool_schemas(self, *, plan_mode: bool) -> list[dict[str, Any]]:
        """The schemas the loop binds DIRECTLY, outside the ``ToolRegistry``.

        Three populations ``_bind_model`` appends beyond the registry specs: the
        spawn family (gated on ``self._spawn_agent_tool`` + the plan-mode rule),
        ``activate_skill`` (when auto-invocable skills exist and skills are on),
        and the per-agent SESSION TOOLS (mode-filtered). This is the ONE source
        of truth, shared by ``_bind_model`` (what gets BOUND) and the
        ``tool_search`` supplement (what search can FIND) — so the searchable set
        can never drift from the bound set, and a strictly-scoped agent can never
        widen its surface via search.
        """
        extra: list[dict[str, Any]] = []
        # In plan mode, the root (depth=0) gets agent management tools so it
        # can spawn and monitor the plan sub-agent. Non-root plan agents get
        # no agent tools — they explore and draft only.
        plan_root = plan_mode and self._ctx.depth == 0
        if (not plan_mode or plan_root) and self._spawn_agent_tool is not None:
            from mewbo_core.agents.spawn_agent import SPAWN_AGENT_SCHEMA, SPAWN_AGENTS_SCHEMA

            extra.extend([SPAWN_AGENT_SCHEMA, SPAWN_AGENTS_SCHEMA])
            # Root-only management tools
            # for non-blocking agent monitoring and steering. Each is ceiling-
            # checked on its OWN id: a delegating AgentDef that means to monitor
            # names ``check_agents`` (every strict def in the tree does), and one
            # that never steers should not be handed ``steer_agent`` merely for
            # having named a spawn tool.
            if self._ctx.depth == 0:
                from mewbo_core.agents.spawn_agent import (
                    CHECK_AGENTS_SCHEMA,
                    STEER_AGENT_SCHEMA,
                )

                extra.extend(
                    schema
                    for tool_id, schema in (
                        ("check_agents", CHECK_AGENTS_SCHEMA),
                        ("steer_agent", STEER_AGENT_SCHEMA),
                    )
                    if self._loop_injected_admitted(tool_id)
                )
        # Inject activate_skill schema when auto-invocable skills exist — unless
        # the drive opted out (``enable_skills=False``), so a headless product
        # run never burns a step activating a host ``~/.claude`` skill.
        if (
            self._enable_skills
            and self._loop_injected_admitted("activate_skill")
            and not plan_mode
            and self._skill_registry is not None
            and self._skill_registry.list_auto_invocable(self._session_capabilities)
        ):
            from mewbo_core.tooling.skills import ACTIVATE_SKILL_SCHEMA

            extra.append(ACTIVATE_SKILL_SCHEMA)
        # Session-tool schemas only for tools whose ``modes`` include the current
        # orchestration mode. Data-driven — no tool_id string match. Plugin tools
        # missing the attribute default to act-mode.
        current_mode = "plan" if plan_mode else "act"
        for session_tool in self._session_tools:
            tool_modes = getattr(session_tool, "modes", None) or DEFAULT_SESSION_TOOL_MODES
            if current_mode in tool_modes:
                extra.append(session_tool.schema)
        return extra

    def _bind_model(
        self,
        tool_schemas: list[dict[str, Any]],
        *,
        model_name: str | None = None,
    ) -> Any:
        """Build a chat model and bind the COMPLETE tool surface to it.

        The only place the full bind list is assembled: the caller's registry
        schemas plus every directly-bound leg (the spawn family,
        ``activate_skill``, the per-agent SessionTools) from the shared
        ``_directly_bound_tool_schemas``. So every caller that needs a bound
        model comes through here — ``model_name`` defaults to the ACTIVE model,
        and a fallback rung passes the model it is escalating to. Binding a rung
        from the caller's registry-only list instead is how an escalated call
        silently lost delegation and its session tools for one generation.
        """
        model = build_chat_model(model_name=model_name or self._active_model)
        plan_mode = self._current_mode == "plan"
        tool_schemas = [
            *tool_schemas,
            *self._directly_bound_tool_schemas(plan_mode=plan_mode),
        ]
        self._bound_tool_count = len(tool_schemas)
        if tool_schemas:
            return model.bind_tools(tool_schemas)
        return model

    # ------------------------------------------------------------------
    # Tool execution (async)
    # ------------------------------------------------------------------

    async def _execute_tool_call(
        self,
        tool_call: Any,
        tool_specs: list[ToolSpec],
    ) -> ToolCallResult:
        """Execute a single LLM tool_call: permission → hooks → run → emit event."""
        tool_call_id: str = tool_call.get("id") or ""
        tool_id: str = tool_call.get("name") or ""

        action_step = self._tool_call_to_action_step(tool_call)

        # A session tool that RETURNS a structured-error envelope (wiki/scg
        # ``err_result``) is a handled failure, not a raise — captured here and
        # applied at the common emit/return path so the step records
        # ``success=False`` and the loop's failure nudge fires (the model still
        # gets the envelope text). ``None`` for every normal result.
        session_tool_error: _SessionToolError | None = None

        # MCP input coercion.
        spec = self._tool_registry.get_spec(tool_id)
        if spec is not None:
            coercion_error = _coerce_mcp_tool_input(action_step, spec)
            if coercion_error:
                self._emit_tool_result_event(
                    action_step, None, error=coercion_error, tool_call_id=tool_call_id
                )
                return ToolCallResult(
                    tool_call_id=tool_call_id,
                    tool_id=tool_id,
                    content=f"ERROR: {coercion_error}",
                    success=False,
                )

        # Safety-plane gate. Runs BEFORE the normal permission check — an
        # operator-owned guardrail outranks the session's own permission
        # policy, and a call the plane denies has no side effect at all.
        # ``spec.capability`` (undeclared for a plugin/MCP tool) is the same
        # tier the capability-mode ceiling already reads, so a rule scoped to
        # "write" behaves identically for a built-in and a runtime-resolved
        # tool. Counts feed the observer's turn-boundary checks even for a
        # call no rule judges directly.
        if self._safety_plane is not None:
            self._safety_tool_calls += 1
            if spec is not None and spec.capability == "write":
                self._safety_write_calls += 1
            call = ToolCallObservation(
                tool_id=tool_id,
                operation=action_step.operation,
                capability=spec.capability if spec is not None else None,
                arguments=(
                    action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
                ),
            )
            verdict = self._safety_plane.evaluate_tool_call(call)
            if verdict is not None and verdict.blocks:
                self._emit_safety_deny(verdict, tool_id)
                denial_content = verdict.reason or f"Blocked by safety policy {verdict.rule!r}."
                self._emit_tool_result_event(
                    action_step, None, error=denial_content, tool_call_id=tool_call_id
                )
                return ToolCallResult(
                    tool_call_id=tool_call_id,
                    tool_id=tool_id,
                    content=f"ERROR: {denial_content}",
                    success=False,
                )

        if not self._check_permission(action_step):
            # If the permission branch wrote a detailed error to
            # ``action_step.result`` (e.g., plan-mode path scoping), use
            # that as the tool-result content so the model can self-correct.
            denial_content = getattr(action_step.result, "content", None) or "Permission denied"
            self._emit_tool_result_event(
                action_step, None, error=str(denial_content), tool_call_id=tool_call_id
            )
            return ToolCallResult(
                tool_call_id=tool_call_id,
                tool_id=tool_id,
                content=str(denial_content),
                success=False,
            )

        action_step = self._hook_manager.run_pre_tool_use(action_step)

        # File read dedup: return a stub if this file was already read
        # with the same params and hasn't changed on disk.
        if tool_id == "read_file":
            dedup_stub = self._check_file_read_cache(action_step)
            if dedup_stub is not None:
                self._emit_tool_result_event(action_step, dedup_stub, tool_call_id=tool_call_id)
                return ToolCallResult(
                    tool_call_id=tool_call_id,
                    tool_id=tool_id,
                    content=dedup_stub,
                    success=True,
                )

        # Execute — internal tools (spawn_agent, session tools, activate_skill)
        # first, then the registry.
        session_tool = next(
            (t for t in self._session_tools if t.tool_id == tool_id), None
        )
        # Publish the active containment for the duration of
        # THIS tool's run — around the WHOLE dispatch chain, not just the
        # registry branch, because a session tool, a skill activation and a
        # spawn each touch the filesystem too, and a branch outside this
        # block sees ``get_active_project_root() is None`` and falls back to
        # the unscoped tenant union. The scope must not depend on which
        # dispatch arm a tool happens to live on.
        # It matters so ``resolve_safe_path`` (called deep inside the
        # aider file/edit/shell tools, and the LSP tool) enforces the
        # firebreak without the loop threading a live object through JSON
        # args. A no-op when ``self._containment`` is None/inactive, so a
        # full_access / enforcement-off run pays nothing. ``contextvars``
        # propagate into ``asyncio.to_thread``, so the sync-tool thread
        # sees the same active containment as this awaiting frame.
        # ``active_project_root`` is published UNCONDITIONALLY, unlike the
        # containment above: the shell sandbox scopes which project's
        # DATA is reachable, which is a different axis from the privilege
        # tier. A root agent is always ``full_access`` and therefore never
        # contained, so a shell scope hung off containment would never
        # apply to the session an operator actually drives.
        with active_containment(self._containment), active_project_root(self._cwd):
            if tool_id == "spawn_agent" and self._spawn_agent_tool is not None:
                try:
                    result = await self._spawn_agent_tool.run_async(action_step)
                except Exception as exc:
                    logging.error("spawn_agent failed: {}", exc)
                    self._emit_tool_result_event(
                        action_step, None, error=str(exc), tool_call_id=tool_call_id
                    )
                    return ToolCallResult(
                        tool_call_id=tool_call_id,
                        tool_id=tool_id,
                        content=f"ERROR: {exc}",
                        success=False,
                    )
                # A REFUSED spawn returns normally — the same "handled failure as a
                # successful return" shape a session tool has, and it fell through
                # to ``success=True`` for exactly the same reason. So it now rides
                # exactly the same envelope and the same parse: no delegation
                # happened, so the step must record as failed (failure nudge,
                # doom-loop tracking, ``permanence``) rather than as "✓ ok".
                session_tool_error = _SessionToolError.parse(result)
            elif tool_id == "spawn_agents" and self._spawn_agent_tool is not None:
                try:
                    result = await self._spawn_agent_tool.run_batch_async(action_step)
                except Exception as exc:
                    logging.error("spawn_agents failed: {}", exc)
                    self._emit_tool_result_event(
                        action_step, None, error=str(exc), tool_call_id=tool_call_id
                    )
                    return ToolCallResult(
                        tool_call_id=tool_call_id,
                        tool_id=tool_id,
                        content=f"ERROR: {exc}",
                        success=False,
                    )
            elif tool_id == "check_agents" and self._spawn_agent_tool is not None:
                try:
                    result = await self._spawn_agent_tool.handle_check_agents(action_step)
                except Exception as exc:
                    logging.error("check_agents failed: {}", exc)
                    self._emit_tool_result_event(
                        action_step, None, error=str(exc), tool_call_id=tool_call_id
                    )
                    return ToolCallResult(
                        tool_call_id=tool_call_id,
                        tool_id=tool_id,
                        content=f"ERROR: {exc}",
                        success=False,
                    )
            elif tool_id == "steer_agent" and self._spawn_agent_tool is not None:
                try:
                    result = await self._spawn_agent_tool.handle_steer_agent(action_step)
                except Exception as exc:
                    logging.error("steer_agent failed: {}", exc)
                    self._emit_tool_result_event(
                        action_step, None, error=str(exc), tool_call_id=tool_call_id
                    )
                    return ToolCallResult(
                        tool_call_id=tool_call_id,
                        tool_id=tool_id,
                        content=f"ERROR: {exc}",
                        success=False,
                    )
            elif session_tool is not None:
                try:
                    result = await session_tool.handle(action_step)
                except Exception as exc:
                    logging.error("session tool {} failed: {}", tool_id, exc)
                    self._emit_tool_result_event(
                        action_step, None, error=str(exc), tool_call_id=tool_call_id
                    )
                    return ToolCallResult(
                        tool_call_id=tool_call_id,
                        tool_id=tool_id,
                        content=f"ERROR: {exc}",
                        success=False,
                    )
                # A handled error envelope (returned, not raised) → reclassify as a
                # FAILED step while keeping the envelope text as the model output.
                session_tool_error = _SessionToolError.parse(result)
            elif (
                tool_id == TOOL_SEARCH_TOOL_ID
                and (tool_search_runner := self._tool_registry.get(tool_id)) is not None
            ):
                # tool_search searches the ToolRegistry, but the loop ALSO binds
                # spawn-family / activate_skill / session-tool schemas directly —
                # invisible to the registry, and so unsearchable on their own.
                # Hand the runner the SAME directly-bound schemas as this turn's
                # bind so those tools are findable; by construction search can only
                # surface what is already bound, never widen a scoped agent's scope.
                supplement = self._directly_bound_tool_schemas(
                    plan_mode=self._current_mode == "plan"
                )
                try:
                    result = await asyncio.to_thread(
                        tool_search_runner.run, action_step, supplement=supplement
                    )
                except Exception as exc:
                    logging.error("tool_search failed: {}", exc)
                    self._emit_tool_result_event(
                        action_step, None, error=str(exc), tool_call_id=tool_call_id
                    )
                    return ToolCallResult(
                        tool_call_id=tool_call_id,
                        tool_id=tool_id,
                        content=f"ERROR: {exc}",
                        success=False,
                    )
            elif tool_id == "activate_skill" and self._skill_registry is not None:
                result = self._handle_activate_skill(action_step)
            else:
                tool = self._tool_registry.get(tool_id)
                if tool is None:
                    self._emit_tool_result_event(
                        action_step, None, error="Tool not available", tool_call_id=tool_call_id
                    )
                    return ToolCallResult(
                        tool_call_id=tool_call_id,
                        tool_id=tool_id,
                        content="ERROR: Tool not available",
                        success=False,
                    )
                try:
                    # Prefer async execution for tools that support it (MCP tools).
                    # Falls back to to_thread for sync-only tools (aider_*, etc.).
                    if hasattr(tool, "arun"):
                        result = await tool.arun(action_step)
                    else:
                        result = await asyncio.to_thread(tool.run, action_step)
                except Exception as exc:
                    logging.error("Tool execution failed: {}", exc)
                    self._emit_tool_result_event(
                        action_step, None, error=str(exc), tool_call_id=tool_call_id
                    )
                    return ToolCallResult(
                        tool_call_id=tool_call_id,
                        tool_id=tool_id,
                        content=f"ERROR: {exc}",
                        success=False,
                    )

        result = self._hook_manager.run_post_tool_use(action_step, result)

        content = getattr(result, "content", None)
        if content is None:
            content = "" if result is None else str(result)
        # A tool may carry image parts alongside its text (a device
        # screenshot). They are read off a SEPARATE attribute rather than
        # smuggled into ``content``, so everything below stays string-only:
        # ``str()``-ing a list of content parts would hand the model the Python
        # repr of that list, which reads as working and is unusable. The images
        # also never reach ``event_str``, so no base64 is persisted to the
        # session store or replayed on a transcript read.
        multimodal = ToolResultContent.parse(content)
        content = multimodal.text if multimodal.has_images else content
        result_images = multimodal.images or tuple(getattr(result, "images", ()) or ())
        content_str = str(content) if not isinstance(content, str) else content
        max_chars = self._result_char_cap(tool_id)
        if isinstance(content, dict):
            content_str = json.dumps(content, ensure_ascii=False, default=str)

        # Micro-compaction: strip ANSI escapes.
        content_str = _ANSI_ESCAPE_RE.sub("", content_str)

        # Snapshot for the event — preserves the JSON envelope so the
        # frontend can always parse structured fields (exit_code, duration_ms).
        # Use a generous event-side cap (decoupled from `max_chars`, which is
        # the LLM-context cap): the frontend can scroll the full output, and
        # `result_file` still backstops pathological results when an export dir
        # is configured. Without this, MCP tool responses get truncated to
        # 2000 chars in the UI even though the full content exists in memory.
        event_str = content_str
        # Floor the event cap at the model-facing cap. Without that, a tool
        # whose declared cap exceeds the event cap would record LESS than the
        # model read — inverting the very fidelity the paired
        # ``result``/``result_seen`` keys exist to establish.
        event_max_chars = max(_EVENT_SNAPSHOT_MAX_CHARS, max_chars)
        if len(event_str) > event_max_chars:
            event_str = event_str[:event_max_chars] + "\n[truncated — see result_file]"

        # Decided ONCE, and before the export branch below can shrink
        # ``content_str`` to a pointer: whether the model read less than the tool
        # produced has two consumers — the cut itself and the read cache — and
        # two derivations of it would be free to disagree.
        result_truncated = bool(max_chars) and len(content_str) > max_chars

        # Save large results to file when export dir is configured.
        result_file: str | None = None
        export_dir = str(get_config_value("runtime", "result_export_dir", default="") or "")
        if export_dir and result_truncated:
            try:
                export_path = Path(export_dir)
                export_path.mkdir(parents=True, exist_ok=True)
                fid = tool_call_id or f"{tool_id}-{int(_time.time() * 1000)}"
                safe_id = re.sub(r"[^\w\-]", "_", fid)
                result_path = export_path / f"{safe_id}.txt"
                result_path.write_text(content_str, encoding="utf-8")
                result_file = str(result_path)
                content_str = (
                    f"[Full output ({len(content_str)} chars) saved to {result_file}. "
                    f"Read the file for complete content.]"
                )
            except OSError:
                pass  # Fall through to normal truncation

        # Final truncation for the LLM. A dict fits ITSELF rather than being cut
        # as a string, because the cap applies to the serialized envelope and a
        # string-level cut lands mid-value — see ``_fitted_json``. The length is
        # re-tested because the export branch may already have replaced the
        # payload with a short pointer.
        #
        # A plain STRING result is windowed, for the reason ``_windowed`` states
        # at length: the verdict of a command lives at its END. This arm used to
        # be a head-only ``content_str[:max_chars]``, so ``_windowed`` reached
        # only the fields of a dict result and every string-returning tool was
        # cut head-first — while the ``mewbo-harness`` skill told the model the
        # opposite, in the engine's own voice. A model that trusts a documented
        # invariant and gets the other behaviour cannot tell a bounded read from
        # a complete one, which is the failure the marker exists to prevent.
        if result_truncated and len(content_str) > max_chars:
            content_str = (
                self._fitted_json(content, max_chars)
                if isinstance(content, dict)
                else self._windowed(content_str, max_chars)
            )

        # Populate file read cache after successful read. A truncated read did
        # not deliver the file, so caching it would let the dedup stub send the
        # model back to an incomplete payload AND refuse the re-read that would
        # have completed it — the cache exists to spare a redundant read, and a
        # read that did not deliver the file is not one.
        if tool_id == "read_file" and not result_truncated:
            self._populate_file_read_cache(action_step)

        # Invalidate file read cache when a file is edited.
        if tool_id in ("file_edit_tool", "aider_edit_block_tool"):
            edit_args = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
            edited_path = str(edit_args.get("file_path", "") or edit_args.get("path", ""))
            if edited_path:
                norm = os.path.normpath(edited_path)
                self._file_read_cache.pop(norm, None)
                # Passive LSP diagnostics: surface errors after edits.
                content_str = _append_lsp_feedback(
                    content_str,
                    norm,
                    self._cwd or "",
                )

        # A session tool's handled error envelope records as a FAILED step (so
        # the loop's per-step failure nudge fires) while the model still receives
        # the full envelope JSON as the tool output — error surfacing without
        # hiding the structured detail the tool chose to return.
        if session_tool_error is not None:
            self._emit_tool_result_event(
                action_step,
                event_str,
                tool_call_id=tool_call_id,
                error=session_tool_error.summary,
                result_file=result_file,
                seen=content_str,
                permanence=session_tool_error.permanence,
            )
            return ToolCallResult(
                tool_call_id=tool_call_id,
                tool_id=tool_id,
                content=content_str,
                success=False,
                blocked_code=(
                    session_tool_error.code
                    if session_tool_error.blocks_completion
                    else None
                ),
                permanence=session_tool_error.permanence,
            )

        self._emit_tool_result_event(
            action_step,
            event_str,
            tool_call_id=tool_call_id,
            result_file=result_file,
            seen=content_str,
        )
        return ToolCallResult(
            tool_call_id=tool_call_id,
            tool_id=tool_id,
            content=content_str,
            success=True,
            # Images ride ONLY on the success path. The provider rejects a
            # ``tool_result`` that carries a non-text block while marked as an
            # error, so a failed capture must degrade to text — the failure
            # arrives as a 400 on the whole request, not as a bad image.
            images=result_images,
        )

    # ------------------------------------------------------------------
    # ActionStep construction
    # ------------------------------------------------------------------

    def _tool_call_to_action_step(self, tool_call: Any) -> ActionStep:
        """Convert an LLM tool_call dict to an ActionStep."""
        tool_id: str = tool_call.get("name") or ""
        args: Any = tool_call.get("args") or {}
        # Inject session cwd as `root` for registered local tools (aider-style
        # file/shell tools consume it via `argument.get("root")`). Unregistered
        # tools — session tools, spawn_agent, activate_skill — use strict
        # schemas that reject stray keys, so they opt out by not being here.
        #
        # Two regimes for the `root` injection:
        #   * ACTIVE containment (`self._containment is not None`): the workspace
        #     root is AUTHORITATIVE — always forced, OVERRIDING any model-supplied
        #     `root`, so a model cannot widen its way out of the jail by naming a
        #     sibling root. (`resolve_safe_path` then also collapses the tenant
        #     union to the containment's allowed roots via the active-containment
        #     context set around execution below.)
        #   * No active containment (flag off / full_access / no cwd): ADVISORY
        #     only — inject just when the model did not supply its own `root`.
        if isinstance(args, dict):
            spec = self._tool_registry.get_spec(tool_id)
            if spec is not None and spec.kind != "mcp":
                if self._containment is not None:
                    args = {**args, "root": self._containment.root}
                elif self._cwd and "root" not in args:
                    args = {**args, "root": self._cwd}
        operation = _infer_operation(tool_id)
        return ActionStep(
            title=tool_id,
            tool_id=tool_id,
            operation=operation,
            tool_input=args,
        )

    # ------------------------------------------------------------------
    # File-read dedup cache
    # ------------------------------------------------------------------

    def _check_file_read_cache(self, action_step: ActionStep) -> str | None:
        """Return a stub if this file was already read with the same params, else ``None``."""
        args = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
        path = str(args.get("path", ""))
        if not path:
            return None
        root = str(args.get("root") or "")
        offset = int(args.get("offset", 0) or 0)
        limit = args.get("limit")
        if limit is not None:
            try:
                limit = int(limit)
            except (TypeError, ValueError):
                limit = None

        try:
            joined = os.path.join(root, path) if root else path
            full_path = os.path.normpath(joined)
        except (TypeError, ValueError):
            return None

        cached = self._file_read_cache.get(full_path)
        if cached is None:
            return None
        if cached.offset != offset or cached.limit != limit:
            return None

        # Check mtime — file may have been edited externally.
        try:
            current_mtime = os.path.getmtime(full_path)
        except OSError:
            return None
        if current_mtime != cached.mtime:
            del self._file_read_cache[full_path]
            return None

        return (
            "File unchanged since last read. The content from the earlier "
            "Read tool_result in this conversation is still current — "
            "refer to that instead of re-reading."
        )

    def _populate_file_read_cache(self, action_step: ActionStep) -> None:
        """Record a successful file read for future dedup."""
        args = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
        path = str(args.get("path", ""))
        if not path:
            return
        root = str(args.get("root") or "")
        offset = int(args.get("offset", 0) or 0)
        limit = args.get("limit")
        if limit is not None:
            try:
                limit = int(limit)
            except (TypeError, ValueError):
                limit = None
        try:
            joined = os.path.join(root, path) if root else path
            full_path = os.path.normpath(joined)
            mtime = os.path.getmtime(full_path)
        except (TypeError, ValueError, OSError):
            return
        self._file_read_cache[full_path] = _CachedFileRead(
            path=full_path,
            offset=offset,
            limit=limit,
            mtime=mtime,
        )

    # ------------------------------------------------------------------
    # Safety plane
    # ------------------------------------------------------------------

    def _emit_safety_disclosure(self) -> None:
        """Disclose the active plane to the user before it evaluates anything.

        Called once, near the top of ``run()``, before the first tool call of
        the run can execute. Silent inspection is exactly what disclosure
        exists to rule out, so this fires unconditionally when a plane is
        attached — there is no configuration that attaches a plane without
        disclosing it.
        """
        if self._safety_plane is None:
            return
        disclosure = self._safety_plane.disclosure()
        self._emit_event(
            {
                "type": "safety_plane",
                "payload": {
                    "phase": "disclosed",
                    "documents": list(disclosure.documents),
                    "rules": [
                        {
                            "name": r.name,
                            "kind": r.kind,
                            "decision": r.decision,
                            "inspects": r.inspects,
                        }
                        for r in disclosure.rules
                    ],
                    "inspects_tool_calls": disclosure.inspects_tool_calls,
                    "inspects_turns": disclosure.inspects_turns,
                },
            }
        )

    def _emit_safety_deny(self, verdict: SafetyVerdict, tool_id: str = "") -> None:
        """Disclose a safety-plane verdict that stopped a call or the run."""
        self._emit_event(
            {
                "type": "safety_plane",
                "payload": {
                    "phase": "deny",
                    "rule": verdict.rule,
                    "reason": verdict.reason,
                    "tool_id": tool_id,
                },
            }
        )

    def _check_safety_turn(self, turns: int) -> bool:
        """Evaluate the observer at the top of a turn.

        Returns ``True`` when the run must stop. Reads a plain scalar struct —
        no message history, no agent tree — so the cost of this check does not
        grow with how long the session has been running.
        """
        if self._safety_plane is None:
            return False
        if self._safety_started_at is None:
            self._safety_started_at = _time.monotonic()
        observation = TurnObservation(
            step=turns,
            elapsed_seconds=_time.monotonic() - self._safety_started_at,
            tool_calls=self._safety_tool_calls,
            write_calls=self._safety_write_calls,
        )
        verdict = self._safety_plane.evaluate_turn(observation)
        if verdict is None or not verdict.blocks:
            return False
        self._emit_safety_deny(verdict)
        return True

    # ------------------------------------------------------------------
    # Permission
    # ------------------------------------------------------------------

    def _check_permission(self, action_step: ActionStep) -> bool:
        """Check permission for an action step. Returns True if allowed."""
        # Plan-mode gating is authoritative: read-only tools, the scoped
        # edit tool, and exit_plan_mode are allowed; everything else is
        # denied. The normal approval policy is bypassed so that plan-mode
        # exploration does not get blocked by ASK rules.
        if self._current_mode == "plan":
            return self._plan_mode_permission(action_step)

        decision = self._permission_policy.decide(action_step)
        decision = self._hook_manager.run_permission_request(action_step, decision)
        if decision == PermissionDecision.ASK:
            approved = self._approval_callback(action_step) if self._approval_callback else False
            decision = PermissionDecision.ALLOW if approved else PermissionDecision.DENY
            self._emit_event(
                {
                    "type": "permission",
                    "payload": {
                        "tool_id": action_step.tool_id,
                        "operation": action_step.operation,
                        "tool_input": action_step.tool_input,
                        "decision": decision.value,
                    },
                }
            )
        if decision == PermissionDecision.DENY:
            mock = get_mock_speaker()
            action_step.result = mock(content=f"Permission denied for {action_step.tool_id}.")
            return False
        return True

    def _plan_mode_permission(self, action_step: ActionStep) -> bool:
        """Plan-mode permission branch: allow read-only + path-scoped edits.

        Returns True if the action is allowed in plan mode (and normal
        policy checks should also run), or False to deny outright. The
        denial path sets an actionable error on ``action_step.result`` and
        emits a permission event so the model and user see the refusal.
        """
        tool_id = action_step.tool_id
        # Always allow internal tools that signal loop termination.
        if tool_id == "exit_plan_mode":
            return True
        # Root (depth=0) can use agent management tools to spawn and
        # monitor the plan sub-agent.  Non-root plan agents cannot.
        _AGENT_MGMT_TOOLS = {"spawn_agent", "spawn_agents", "check_agents", "steer_agent"}
        if tool_id in _AGENT_MGMT_TOOLS and self._ctx.depth == 0:
            return True
        spec = self._tool_registry.get_spec(tool_id)
        # Read-only tools are unrestricted in plan mode.
        if spec is not None and spec.read_only:
            return True
        # The configured edit tool is allowed ONLY when its target path
        # resolves inside the session's plan directory.
        edit_tool_id = self._configured_edit_tool_id()
        if tool_id == edit_tool_id and self._session_id is not None:
            args = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
            candidate = str(args.get("file_path", "") or "")
            if candidate and is_inside_plan_dir(candidate, self._session_id):
                return True
            attempted = candidate or "<missing file_path>"
            plan_path = plan_file_for(self._session_id)
            msg = get_prompt_registry().render(
                "loop.plan_edit_restricted", plan_path=plan_path, attempted=attempted
            )
            mock = get_mock_speaker()
            action_step.result = mock(content=msg)
            self._emit_event(
                {
                    "type": "permission",
                    "payload": {
                        "tool_id": tool_id,
                        "operation": action_step.operation,
                        "tool_input": action_step.tool_input,
                        "decision": "deny",
                    },
                }
            )
            return False
        # Shell tool: permitted only when the command matches an allowlisted
        # prefix AND contains no shell metacharacters. The denial message
        # includes the allowlist so the model can self-correct in one turn.
        shell_allowlist = self._plan_mode_shell_allowlist()
        if tool_id in SHELL_TOOL_IDS and shell_allowlist:
            args = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
            command = str(args.get("command", "") or "").strip()
            if is_shell_command_plan_safe(command, shell_allowlist):
                return True
            allowed_preview = ", ".join(shell_allowlist)
            attempted = command or "<missing command>"
            plan_hint = ""
            if self._session_id is not None:
                if self._ctx.depth == 0:
                    plan_hint = (
                        " You cannot write the plan directly. Spawn a sub-agent to draft it."
                    )
                else:
                    plan_hint = (
                        f" To write the plan, use your edit tool on "
                        f"{plan_file_for(self._session_id)}."
                    )
            msg = (
                f"Plan mode: shell command blocked. You attempted: `{attempted}`. "
                f"Allowed prefixes: {allowed_preview}. No pipes, redirects, "
                f"`&&`/`;`, `$VAR` expansion, or backticks.{plan_hint}"
            )
            mock = get_mock_speaker()
            action_step.result = mock(content=msg)
            self._emit_event(
                {
                    "type": "permission",
                    "payload": {
                        "tool_id": tool_id,
                        "operation": action_step.operation,
                        "tool_input": action_step.tool_input,
                        "decision": "deny",
                    },
                }
            )
            return False
        # User-enabled MCP tools: unconditionally allowed. Mewbo cannot
        # classify a third-party MCP tool's effect, so plan mode trusts the
        # user's mcp.json rather than guessing at a mode filter.
        if spec is not None and spec.kind == "mcp":
            return True
        # Everything else (agent tools for non-root, shell when allowlist
        # empty, MCP when flag is False, hallucinated tool names) is denied.
        plan_hint = ""
        if self._session_id is not None:
            if self._ctx.depth == 0:
                plan_hint = " You cannot write the plan directly. Spawn a sub-agent to draft it."
            else:
                plan_hint = f" Plan file: {plan_file_for(self._session_id)}."
        mock = get_mock_speaker()
        action_step.result = mock(
            content=(
                f"Plan mode: tool '{tool_id}' is unavailable. Use read-only "
                "tools to explore and your edit tool to draft the plan, "
                f"then call exit_plan_mode.{plan_hint}"
            )
        )
        self._emit_event(
            {
                "type": "permission",
                "payload": {
                    "tool_id": tool_id,
                    "operation": action_step.operation,
                    "tool_input": action_step.tool_input,
                    "decision": "deny",
                },
            }
        )
        return False

    # ------------------------------------------------------------------
    # Event emission
    # ------------------------------------------------------------------

    def _emit_tool_result_event(
        self,
        action_step: ActionStep,
        result: str | None,
        *,
        tool_call_id: str = "",
        error: str | None = None,
        result_file: str | None = None,
        seen: str | None = None,
        permanence: str | None = None,
    ) -> None:
        """Emit one ``tool_result`` event for a finished (or failed) tool call.

        *tool_call_id* pairs this event with the ``tool_call`` emitted before the
        call was dispatched. It defaults to ``""`` — "not correlatable" — because a
        provider need not supply an id and because every transcript written before
        the initiation event existed has none; a consumer must therefore treat an
        empty value as absence rather than as a shared key.

        *result* is the RAW payload snapshot (event-side cap) and *seen* is the
        string actually handed to the model after truncation. Both are recorded,
        labeled, because they routinely differ by two orders of magnitude — a
        100K-character result for a call the model read 2,000 characters of.
        Keeping only the raw one makes every "the model had this and ignored
        it" reading of a transcript unfalsifiable. *seen* is omitted when it is
        identical to *result*.

        *permanence* carries a tool's own verdict on whether its failure can
        ever succeed on retry (see :class:`_SessionToolError`).
        """
        max_chars = self._result_char_cap(action_step.tool_id)
        # `summary` is a short preview for log titles / agent-tree rendering;
        # the full payload lives in `result`, which the frontend renders in a
        # scrollable container. Keep `summary` capped at the LLM-context size
        # so it stays human-skimmable.
        if max_chars and result and len(result) > max_chars:
            summary = error or result[:max_chars]
        else:
            summary = error or result or ""
        # Drained unconditionally so a recorded headline can never survive into
        # the next call's event — but only USED on the success path, where the
        # alternative is a prefix of the payload that can restate it and nothing
        # more. An error already has a title derived from structured facts.
        headline = self._declared_headline(action_step.tool_id)
        if headline and not error:
            summary = headline
        payload: dict[str, Any] = {
            "tool_call_id": tool_call_id,
            "tool_id": action_step.tool_id,
            "operation": action_step.operation,
            "tool_input": action_step.tool_input,
            "result": result,
            "success": error is None,
            "summary": f"ERROR: {error}" if error else summary,
        }
        if error:
            payload["error"] = error
        if result_file:
            payload["result_file"] = result_file
        # What the model actually read, recorded only when it differs from the
        # raw snapshot — so a consumer can tell a truncated read from a full one
        # instead of inferring it from a cap it would have to re-derive.
        if seen is not None and seen != result:
            payload["result_seen"] = seen
            payload["result_truncated"] = True
        if permanence:
            payload["permanence"] = permanence
        # Always tag with agent_id and model so the console can display
        # badges for all agents including the root. The ACTIVE model, not the
        # frozen configured one: after a sticky escalation the context still
        # names the model that died, so every tool badge for the rest of the run
        # credited a model that served none of the calls.
        payload["agent_id"] = self._ctx.agent_id
        payload["model"] = self._active_model
        self._emit_event({"type": "tool_result", "payload": payload})

    def _emit_event(self, event: Event) -> None:
        # Capture retry/fallback events for the resilience note before they
        # leave for the sink (no-op for every other event type). This is the ONE
        # place both the strategy's automatic switches and the model_control
        # tool's deliberate ones flow through, so the note sees them all.
        self._resilience_note.record(event)
        if self._ctx.event_logger is not None:
            self._ctx.event_logger(event)

    @staticmethod
    def _extract_text_content(content: object) -> str:
        """Extract plain text from an AIMessage content field.

        Delegates to the ONE definition of what a response's text is
        (``RetryStrategy.response_text``, which also drops the internal
        placeholder so it never surfaces in ``agent_message`` events). A second
        copy here is precisely how the loop came to discard, at this seam,
        evidence the ladder needed one layer below.
        """
        return RetryStrategy.response_text(content)

active_model: str property

The model this loop last generated with — the one that SERVED the run.

Read by the orchestrator, which otherwise builds its failure record from the model frozen at construction: a run that escalated A -> B -> C then died named A, and an operator reading that benches a model that had not been active for minutes. Equal to the configured model until a sticky escalation promotes it, so the ordinary path is unchanged.

__init__(*, agent_context: AgentContext, tool_registry: ToolRegistry, permission_policy: PermissionPolicy, approval_callback: Callable[[ActionStep], bool] | None = None, hook_manager: HookManager, safety_plane: SafetyPlane | None = None, project_instructions: str | None = None, user_instructions: str | None = None, skill_instructions: str | None = None, skill_registry: Any = None, agent_registry: Any = None, session_tool_registry: SessionToolRegistry | None = None, allowed_tools: list[str] | None = None, denied_tools: list[str] | None = None, strict_tool_scope: bool = False, cwd: str | None = None, session_id: str | None = None, session_capabilities: tuple[str, ...] = (), extra_session_tools: list[SessionTool] | None = None, enable_skills: bool = True, project_autoselect: bool = False, project_catalog: ProjectCatalog | None = None, session_context_reader: Callable[[], dict[str, object]] | None = None, contract: DelegationContract | None = None, verification: CommandVerification | None = None, verifier_runner: VerifierRunner | None = None, watchdog_sleeper: Callable[[float], Awaitable[None]] | None = None) -> None

Initialize the tool-use loop.

Parameters:

Name Type Description Default
agent_context AgentContext

Required — carries model, cancel, logger, registry.

required
tool_registry ToolRegistry

Registered tools available to this agent.

required
permission_policy PermissionPolicy

Permission rules for tool execution.

required
approval_callback Callable[[ActionStep], bool] | None

Optional callback for ASK decisions (None for sub-agents).

None
hook_manager HookManager

Lifecycle hooks.

required
safety_plane SafetyPlane | None

Operator-owned tool-call gate and session observer, built once by the orchestrator from config + .mewbo/ and forwarded by value to every child loop. None (the default) when the deployment-wide switch is off — every safety call site is a single is not None check away from a no-op.

None
project_instructions str | None

CLAUDE.md / AGENTS.md content discovered at session start.

None
user_instructions str | None

Operator-authored custom instructions, ALREADY RENDERED by Orchestrator from the stored template (system_instructions/). A plain string — the loop does no Jinja and no store work. Lives on the instance, not on a per-call arg, so it survives the in-place system-prompt re-render that a model escalation performs.

None
skill_instructions str | None

Pre-rendered skill body (from user /skill invocation).

None
skill_registry Any

SkillRegistry for auto-invocation catalog + activate_skill handling.

None
agent_registry Any

AgentRegistry for agent type catalog + spawn_agent type lookup.

None
session_tool_registry SessionToolRegistry | None

Registry of plugin-contributed session-tool factories. Each matching factory (filtered by allowed_tools) is instantiated for this agent and added to self._session_tools.

None
allowed_tools list[str] | None

The agent's allowlist used to filter which session tools the plugin registry should build for this agent. None means "no plugin session tools" (root agents get only the built-in ExitPlanModeTool).

None
denied_tools list[str] | None

Session-tool ids withheld regardless of which gate in SessionToolRegistry.build_for would otherwise admit them — the unconditional auto-surface and the capability auto-surface included, and it beats a NAMED allowed_tools entry too. Deliberately NOT three-state like allowed_tools: deny is purely subtractive, so None and [] are the same "nothing denied" set. None (the default) changes nothing for every existing session.

None
strict_tool_scope bool

Whether allowed_tools is AUTHORITATIVE for this agent. True (spawned leaf sub-agents, wiki-qa/search runs) — the allowlist is the whole tool scope, so it also gates spawn_agent. False (the FE default) — allowed_tools is only a PERMISSIVE ceiling over MCP tools (context.mcp_tools); built-ins and the internal spawn_agent are NOT scoped by it, mirroring the orchestrator's permissive filter_specs branch.

False
cwd str | None

Working directory for this agent (project root).

None
session_id str | None

Session identifier — used for plan-mode path scoping.

None
session_capabilities tuple[str, ...]

Client-advertised capability tuple from the X-Mewbo-Capabilities header (persisted on the session context event). Used to filter capability-gated agents and skills out of the system-prompt catalogs and activate_skill / spawn_agent lookups.

()
extra_session_tools list[SessionTool] | None

Caller-injected SessionTool instances appended to self._session_tools without a plugin manifest (e.g. the structured-response emit_result tool).

None
enable_skills bool

When False, the auto-invocable activate_skill schema is never injected even if the registry holds skills — a headless product drive (search/wiki) can opt out so it doesn't burn its first step activating a host ~/.claude skill it never intended to expose. Default True (unchanged behavior).

True
project_autoselect bool

When True (and this is the ROOT agent), bind list_projects / switch_project so the model can choose which workspace to work in. Default False binds neither.

False
project_catalog ProjectCatalog | None

The catalog those two tools read and resolve through, built by the caller (Orchestrator) because assembling it needs the config plus the project/repository stores. None disables the pair regardless of the flag.

None
session_context_reader Callable[[], dict[str, object]] | None

Reads back the session's CURRENT effective context — the most-recent context payload — so a project switch can carry it forward instead of replacing it. Injected because this loop has no transcript access at all: its only context channel is the write-only event_logger, and the truncated recent_events window it does hold would produce a merge that silently drops anything older than the window. None degrades to writing the two switch keys alone.

None
contract DelegationContract | None

The spawner's declared DelegationContract for THIS agent — None/disabled leaves the child unbounded. Checked in the run loop's budget block IN ADDITION TO (never instead of) the shared session budget.

None
verification CommandVerification | None

The spawner's declared ground-truth completion check for THIS agent — None leaves completion ungated. Only RUN when the two-gate self._verification_active holds (master switch on AND a spec AND an execute/all capability mode); otherwise it is carried but inert.

None
verifier_runner VerifierRunner | None

Injected VerifierRunner (defaults to CommandVerifierRunner). A test passes a recording fake so the gate is exercised without a real subprocess.

None
watchdog_sleeper Callable[[float], Awaitable[None]] | None

Injected wait between watchdog sweeps (defaults to asyncio.sleep). A test drives one sweep by passing a fake, so it never has to patch the stdlib module every other coroutine in the process also awaits through.

None
Source code in packages/mewbo_core/src/mewbo_core/loop/tool_use_loop.py
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
def __init__(
    self,
    *,
    agent_context: AgentContext,
    tool_registry: ToolRegistry,
    permission_policy: PermissionPolicy,
    approval_callback: Callable[[ActionStep], bool] | None = None,
    hook_manager: HookManager,
    safety_plane: SafetyPlane | None = None,
    project_instructions: str | None = None,
    user_instructions: str | None = None,
    skill_instructions: str | None = None,
    skill_registry: Any = None,
    agent_registry: Any = None,
    session_tool_registry: SessionToolRegistry | None = None,
    allowed_tools: list[str] | None = None,
    denied_tools: list[str] | None = None,
    strict_tool_scope: bool = False,
    cwd: str | None = None,
    session_id: str | None = None,
    session_capabilities: tuple[str, ...] = (),
    extra_session_tools: list[SessionTool] | None = None,
    enable_skills: bool = True,
    project_autoselect: bool = False,
    project_catalog: ProjectCatalog | None = None,
    session_context_reader: Callable[[], dict[str, object]] | None = None,
    contract: DelegationContract | None = None,
    verification: CommandVerification | None = None,
    verifier_runner: VerifierRunner | None = None,
    watchdog_sleeper: Callable[[float], Awaitable[None]] | None = None,
) -> None:
    """Initialize the tool-use loop.

    Args:
        agent_context: Required — carries model, cancel, logger, registry.
        tool_registry: Registered tools available to this agent.
        permission_policy: Permission rules for tool execution.
        approval_callback: Optional callback for ASK decisions (None for sub-agents).
        hook_manager: Lifecycle hooks.
        safety_plane: Operator-owned tool-call gate and session observer,
            built once by the orchestrator from config + ``.mewbo/`` and
            forwarded by value to every child loop. ``None`` (the default)
            when the deployment-wide switch is off — every safety call
            site is a single ``is not None`` check away from a no-op.
        project_instructions: CLAUDE.md / AGENTS.md content discovered at session start.
        user_instructions: Operator-authored custom instructions, ALREADY RENDERED
            by ``Orchestrator`` from the stored template (``system_instructions/``).
            A plain string — the loop does no Jinja and no store work. Lives on the
            instance, not on a per-call arg, so it survives the in-place system-prompt
            re-render that a model escalation performs.
        skill_instructions: Pre-rendered skill body (from user /skill invocation).
        skill_registry: SkillRegistry for auto-invocation catalog + activate_skill handling.
        agent_registry: AgentRegistry for agent type catalog + spawn_agent type lookup.
        session_tool_registry: Registry of plugin-contributed session-tool
            factories.  Each matching factory (filtered by ``allowed_tools``)
            is instantiated for this agent and added to ``self._session_tools``.
        allowed_tools: The agent's allowlist used to filter which session
            tools the plugin registry should build for this agent.  ``None``
            means "no plugin session tools" (root agents get only the
            built-in ``ExitPlanModeTool``).
        denied_tools: Session-tool ids withheld regardless of which gate in
            ``SessionToolRegistry.build_for`` would otherwise admit them —
            the unconditional auto-surface and the capability auto-surface
            included, and it beats a NAMED ``allowed_tools`` entry too.
            Deliberately NOT three-state like ``allowed_tools``: deny is
            purely subtractive, so ``None`` and ``[]`` are the same "nothing
            denied" set. ``None`` (the default) changes nothing for every
            existing session.
        strict_tool_scope: Whether ``allowed_tools`` is AUTHORITATIVE for this
            agent. ``True`` (spawned leaf sub-agents, wiki-qa/search runs) —
            the allowlist is the whole tool scope, so it also gates
            ``spawn_agent``. ``False`` (the FE default) — ``allowed_tools`` is
            only a PERMISSIVE ceiling over MCP tools (``context.mcp_tools``);
            built-ins and the internal ``spawn_agent`` are NOT scoped by it,
            mirroring the orchestrator's permissive ``filter_specs`` branch.
        cwd: Working directory for this agent (project root).
        session_id: Session identifier — used for plan-mode path scoping.
        session_capabilities: Client-advertised capability tuple from the
            ``X-Mewbo-Capabilities`` header (persisted on the session
            context event). Used to filter capability-gated agents and
            skills out of the system-prompt catalogs and ``activate_skill``
            / ``spawn_agent`` lookups.
        extra_session_tools: Caller-injected ``SessionTool`` instances
            appended to ``self._session_tools`` without a plugin manifest
            (e.g. the structured-response ``emit_result`` tool).
        enable_skills: When ``False``, the auto-invocable ``activate_skill``
            schema is never injected even if the registry holds skills — a
            headless product drive (search/wiki) can opt out so it doesn't
            burn its first step activating a host ``~/.claude`` skill it
            never intended to expose. Default ``True`` (unchanged behavior).
        project_autoselect: When ``True`` (and this is the ROOT agent), bind
            ``list_projects`` / ``switch_project`` so the model can choose
            which workspace to work in. Default ``False`` binds neither.
        project_catalog: The catalog those two tools read and resolve
            through, built by the caller (``Orchestrator``) because
            assembling it needs the config plus the project/repository
            stores. ``None`` disables the pair regardless of the flag.
        session_context_reader: Reads back the session's CURRENT effective
            context — the most-recent ``context`` payload — so a project
            switch can carry it forward instead of replacing it. Injected
            because this loop has no transcript access at all: its only
            context channel is the write-only ``event_logger``, and the
            truncated ``recent_events`` window it does hold would produce a
            merge that silently drops anything older than the window.
            ``None`` degrades to writing the two switch keys alone.
        contract: The spawner's declared ``DelegationContract`` for THIS
            agent — ``None``/disabled leaves the child unbounded.
            Checked in the run loop's budget block
            IN ADDITION TO (never instead of) the shared session budget.
        verification: The spawner's declared ground-truth completion check
            for THIS agent — ``None`` leaves completion ungated. Only
            RUN when the two-gate ``self._verification_active`` holds
            (master switch on AND a spec AND an execute/all capability
            mode); otherwise it is carried but inert.
        verifier_runner: Injected ``VerifierRunner`` (defaults to
            ``CommandVerifierRunner``). A test passes a recording fake so
            the gate is exercised without a real subprocess.
        watchdog_sleeper: Injected wait between watchdog sweeps (defaults
            to ``asyncio.sleep``). A test drives one sweep by passing a
            fake, so it never has to patch the stdlib module every other
            coroutine in the process also awaits through.
    """
    self._ctx = agent_context
    # The run's stop signal, wrapping the SAME predicate the between-turns
    # check has always read. It exists so an in-flight model call or tool
    # execution can be raced against a stop instead of outliving it — see
    # ``cancellation.py`` for why polling alone bounded the stop by a turn.
    self._cancellation = CancellationSignal(agent_context.should_cancel)
    self._contract = contract
    # The model whose per-model prompt overrides + tool variant are ACTIVE.
    # Starts at the configured primary; the fallback ladder promotes it
    # to the escalated model on a sticky switch (see ``_apply_model_escalation``)
    # so the heal becomes behavioural, not just a model swap.
    self._active_model = agent_context.model_name
    self._enable_skills = enable_skills
    self._tool_registry = tool_registry
    self._permission_policy = permission_policy
    self._approval_callback = approval_callback
    self._hook_manager = hook_manager
    # Forwarded BY VALUE from the orchestrator that built the root loop —
    # never re-read from config or disk here. A model switch mid-session
    # re-renders ``messages[0]`` at the safe turn boundary but never touches
    # this reference, and ``_build_child_loop`` (spawn_agent.py) passes the
    # SAME object to every descendant, so an agent editing or deleting the
    # on-disk document mid-session changes nothing about the plane judging
    # its own session.
    self._safety_plane = safety_plane
    self._safety_started_at: float | None = None
    self._safety_tool_calls = 0
    self._safety_write_calls = 0
    self._project_instructions = project_instructions
    self._user_instructions = user_instructions
    self._skill_instructions = skill_instructions
    self._skill_registry = skill_registry
    self._agent_registry = agent_registry
    self._session_tool_registry = session_tool_registry
    self._cwd = cwd
    self._session_id = session_id
    self._session_capabilities = session_capabilities
    # Retained so the tool ceiling can reach the tools this loop injects
    # OUTSIDE ``filter_specs`` — see :meth:`_loop_injected_admitted`.
    self._allowed_tools = allowed_tools
    self._denied_tools = denied_tools
    self._strict_tool_scope = strict_tool_scope
    self._project_catalog = project_catalog
    self._session_context_reader = session_context_reader
    # The catalog key of the project this loop is currently in. ``None``
    # until the first switch — construction supplies a directory, never a
    # key — so it reports "no previous project", never a guessed one.
    self._project_key: str | None = None

    # Deferred-tool partitioning state. ``run()`` stamps all four from the
    # specs it is handed; they are initialized here because
    # ``rebind_workspace`` re-derives the bound set from them and would
    # otherwise depend on the attribute-creation order inside ``run`` — a
    # dependency that fails as an ``AttributeError`` swallowed into a
    # "could not switch" refusal, i.e. a workspace half moved.
    self._tool_specs_full: list[ToolSpec] = []
    self._tool_search_enabled: bool = False
    self._deferred_ids: set[str] = set()
    self._last_active_ids: set[str] = set()
    # A workspace switch performed mid-turn, applied at the NEXT turn
    # boundary (the ``_requested_switch`` pattern): the tool mutates the
    # loop's own state immediately, but the transcript tail stays untouched
    # until the turn ends. ``None`` whenever no switch is pending.
    self._pending_workspace_bind: tuple[list[dict[str, Any]], Any] | None = None

    # Filesystem-containment firebreak. Built ONCE here from
    # three inputs: the enforcement kill-switch (``agent.workspace_
    # enforcement``, staged OFF by default), this agent's narrowed
    # ``workspace_mode``, and the workspace ``cwd``. It is a non-None
    # ``WorkspaceContainment`` ONLY when all three admit containment
    # (enforcement on AND a restrictive tier AND a real cwd) — so with the
    # flag off, a full_access tier, or no cwd, ``self._containment is None``
    # and every path resolves to the plain tenant union.
    # ``self._containment is not None`` is therefore the single "containment
    # active" predicate the loop's root-injection + tool-execution seams read.
    self._containment: WorkspaceContainment | None = self._build_containment(cwd)

    # Verifier-gated completion. Two-gate arming computed ONCE (mirrors the
    # write-progress signal): a ground-truth check only gates an agent that
    # (a) has a spec, (b) runs under the master switch, and (c) could
    # plausibly ACT — capability_mode ∈ {execute, all}. A read-only child,
    # the disabled default, or the staged-off switch leaves the gate inert,
    # so every natural completion is accepted unchecked. The runner
    # is injected (default ``CommandVerifierRunner``) so a test drives the
    # gate with a recording fake and never spawns a real subprocess.
    # The watchdog's poll wait, injected as a collaborator. A caller that
    # must drive a sweep without waiting for one replaces THIS, rather
    # than ``asyncio.sleep`` on the stdlib module — that object is shared
    # by every module and every running loop in the process, so patching
    # it reaches coroutines this loop has nothing to do with.
    self._watchdog_sleep = watchdog_sleeper or asyncio.sleep
    self._verification = verification
    self._verifier_runner: VerifierRunner = verifier_runner or CommandVerifierRunner()
    self._verification_active = verification is not None and CommandVerification.gate_active(
        enabled=bool(get_config_value("agent", "verification_enabled", default=False)),
        capability_mode=agent_context.capability_mode,
    )
    # Latches for the run: whether a check has already passed (never re-run
    # once green) and how many failed re-drives remain.
    self._verify_passed = False
    self._verify_retries_left = int(
        get_config_value("agent", "verification_max_retries", default=2)
    )

    # Dedup cache for read_file: prevents redundant reads when the
    # same file + range hasn't changed on disk (mtime check).
    self._file_read_cache: dict[str, _CachedFileRead] = {}

    # Takes stale images out of history when the list is compacted. A
    # plain field rather than a knob — nothing has asked to tune it.
    self._image_history = ImageHistoryStrip()

    # Plan-mode state (mutable across the loop's lifetime).
    self._current_mode: str = "act"
    # Authoritative token count from the most recent LLM response's
    # usage_metadata.input_tokens. Zero until the first call lands.
    self._last_input_tokens: int = 0

    # In-flight LLM call, for the liveness leg. ``None`` whenever no call is
    # outstanding; a monotonic timestamp while one is. The retry strategy
    # already bounds each attempt with ``asyncio.wait_for``, but that bound
    # can only fire if the awaited coroutine reaches a cancellation point —
    # a provider read wedged below the event loop never does, and one such
    # call sat silent for 24 minutes having emitted ``llm_call_start`` and
    # no end. Nothing noticed it live: the sweepers that would have run once
    # each, at process boot.
    self._llm_call_started_at: float | None = None
    self._llm_call_step: int = 0

    # How many tool schemas the last ``_bind_model`` actually bound, stamped
    # onto ``llm_call_start`` so the bound surface is queryable after the
    # fact. Stored by the binder rather than recomputed at the emit site:
    # a second derivation of "what is bound" is free to drift from the real
    # one, which is the very failure this records.
    self._bound_tool_count: int = 0

    # Error-visibility seam: a bounded, factual note of this run's LLM
    # retry/fallback events, injected as its OWN system-prompt slot so the
    # model can see its own retries. ``_active_resilience_note``
    # tracks the text last baked into ``messages[0]`` so the loop re-renders
    # the prompt only when the note actually changes (a clean run never
    # pays; a healing run re-renders at most once per new event).
    self._resilience_note = ResilienceNote()
    self._active_resilience_note: str = ""

    # Self-steering routing state. The live per-run ``RetryStrategy`` (set in
    # ``run``) is what the ``model_control`` tool reuses for its switch
    # budget + cooldown; ``_requested_switch`` is a model the tool asked to
    # switch to, applied at the NEXT turn boundary (like a sticky fallback,
    # transcript tail untouched); ``_live_messages`` lets the continuity-lock
    # guardrail inspect the in-flight transcript.
    self._retry_strategy: RetryStrategy | None = None
    self._requested_switch: str | None = None
    self._live_messages: list[BaseMessage] | None = None

    # Create SpawnAgentTool when this agent can spawn children — gated on
    # BOTH depth (``can_spawn``) AND tool scope. spawn_agent is injected
    # here rather than through ``filter_specs``, so an explicit ``tools:``
    # allowlist that omits it would otherwise be silently bypassed: a leaf
    # agent scoped to build-and-submit (the st-widget-builder) could still
    # delegate, and misread its own errors as "delegate to a scoped agent",
    # recursing into copies of itself. An ABSENT
    # allowlist (``None``) stays unrestricted (root / ad-hoc spawns); an
    # EMPTY one grants nothing, delegation included. That distinction is
    # load-bearing, not pedantry: a role-bounded viewer's composed allowlist
    # omits the spawn family precisely to disable delegation, and can compose
    # down to empty — reading empty as "unrestricted" would hand delegation
    # back to the principal the ceiling exists to deny.
    #
    # The allowlist gate applies ONLY when the scope is STRICT (a spawned
    # leaf / wiki-qa / search run — where ``allowed_tools`` is the whole
    # authoritative tool scope). Under a PERMISSIVE scope (the FE default),
    # ``allowed_tools`` is ``context.mcp_tools`` — a ceiling over MCP tools
    # only, which never lists the internal spawn_agent; built-ins stay (see
    # the orchestrator's permissive ``filter_specs`` branch), so spawn_agent
    # must stay too. Without this carve-out every console/Aura session that
    # advertised MCP tools had root delegation silently disabled (session
    # 04ea546e…: the st-widget-builder skill mandates spawn_agent, which was
    # unreachable, so the agent violated the skill and built the widget itself).
    spawn_in_scope = (
        not strict_tool_scope
        or allowed_tools is None
        or bool({"spawn_agent", "spawn_agents"} & set(allowed_tools))
    ) and not self._ctx.atomic  # Atomic is a hard firebreak
    self._spawn_agent_tool: Any = None
    if agent_context.can_spawn and spawn_in_scope:
        from mewbo_core.agents.spawn_agent import SpawnAgentTool

        # Sub-agents inherit parent's approval
        # policy so they can execute write/edit/shell tools.
        self._spawn_agent_tool = SpawnAgentTool(
            agent_context=agent_context,
            tool_registry=tool_registry,
            permission_policy=permission_policy,
            approval_callback=approval_callback,
            hook_manager=hook_manager,
            project_instructions=project_instructions,
            user_instructions=user_instructions,
            cwd=cwd,
            agent_registry=agent_registry,
            session_tool_registry=session_tool_registry,
            session_capabilities=session_capabilities,
            enable_skills=enable_skills,
        )

    # Assemble session tools — per-agent stateful handlers that carry
    # their own schema, dispatch, and run-termination flag. The core's
    # built-in ``ExitPlanModeTool`` is always attached to root agents
    # with a session id; plugin-contributed tools are selected by EITHER
    # the agent's ``allowed_tools`` allowlist OR a capability gate (a
    # factory whose ``requires_capabilities`` ⊆ ``session_capabilities``),
    # so a runtime-granted capability surfaces its tools to the
    # root agent without the client listing them explicitly.
    self._session_tools: list[SessionTool] = []
    if agent_context.depth == 0 and session_id is not None:
        # Deliberately NOT ceiling-checked: ``exit_plan_mode`` is the only
        # way out of plan mode, so withholding it from a strict scope that
        # failed to name it would leave the run with no exit at all. It is a
        # structural terminator, never a surface an agent wanders into.
        self._session_tools.append(
            ExitPlanModeTool(
                session_id=session_id,
                event_logger=agent_context.event_logger,
            )
        )
        # Authoritative live todos (act mode): terminal-free, re-emits the
        # FULL statused list as ONE ``todos`` event on each call. Attached
        # inline (not via the plugin factory) so it bypasses the plugin
        # factory's allowlist gate — hence the explicit ceiling check here,
        # or an AgentDef's authoritative ``tools:`` under-states what its
        # agent actually holds.
        if self._loop_injected_admitted("update_todos"):
            self._session_tools.append(
                UpdateTodosTool(
                    session_id=session_id,
                    event_logger=agent_context.event_logger,
                    agent_id=agent_context.agent_id,
                )
            )
        # Project selection (opt-in). Root-only is deliberate and enforced
        # by the enclosing ``depth == 0`` branch: a child is HANDED a
        # workspace by its parent, and letting it re-scope the session would
        # move the ground under every sibling still working in the old
        # directory. Same inline attachment as update_todos, hence the same
        # explicit ceiling check — these bypass the plugin factory's
        # allowlist gate, so an AgentDef's authoritative ``tools:`` would
        # otherwise under-state what its agent holds.
        if project_autoselect and project_catalog is not None:
            from mewbo_core.workspaces.project_switch import (
                LIST_PROJECTS_TOOL_ID,
                SWITCH_PROJECT_TOOL_ID,
                ListProjectsTool,
                SwitchProjectTool,
            )

            if self._loop_injected_admitted(LIST_PROJECTS_TOOL_ID):
                self._session_tools.append(ListProjectsTool(catalog=project_catalog))
            if self._loop_injected_admitted(SWITCH_PROJECT_TOOL_ID):
                self._session_tools.append(
                    SwitchProjectTool(
                        catalog=project_catalog,
                        rebind=self.rebind_workspace,
                    )
                )
    if session_tool_registry is not None and session_id is not None:
        self._session_tools.extend(
            session_tool_registry.build_for(
                allowed_tools,
                session_id=session_id,
                event_logger=agent_context.event_logger,
                session_capabilities=session_capabilities,
                # Deny wins over everything else this call admits —
                # unconditional, the capability auto-surface, even a named
                # allowlist entry. Same list ``ids_for`` (the operator-facing
                # catalog) is given, so the two selections cannot drift.
                denied_tools=denied_tools,
                # df875 law: a PERMISSIVE allowlist (FE mcp_tools) is
                # only an MCP ceiling, so an unconditional tool
                # (schedule_trigger) still surfaces; a STRICT AgentDef scope
                # must name it. Mirrors the spawn_in_scope gate above.
                strict_tool_scope=strict_tool_scope,
                # Delegation privilege ceiling: the loop's effective
                # (already-narrowed) capability_mode gates SESSION tools too,
                # not just registry tools — else a read_only spawn would still
                # receive write-tier session actions (submit/mint/commit/arm).
                # Root is always "all" (no-op); a narrowed sub-agent attenuates.
                capability_mode=agent_context.capability_mode,
            )
        )
    # Caller-injected session tools (e.g. the structured-response emit
    # tool, and every client-declared device tool) — no plugin manifest
    # needed. They terminate / dispatch through the same machinery as
    # plugin tools.
    #
    # They ARE gated on ``capability_mode``, through the same
    # ``capability_mode_admits`` predicate ``build_for`` uses above. This
    # append sits AFTER build_for's gates, so without this it is a hole in
    # the privilege ceiling: a ``read_only`` sub-agent would be handed
    # every client-declared device tool, which now includes a shell at
    # shell UID. A tool that declares no ``capability`` is treated as
    # ``execute`` — session tools are actions, and an undeclared tier must
    # fail closed under a restrictive mode rather than open.
    if extra_session_tools:
        self._session_tools.extend(
            tool
            for tool in extra_session_tools
            if capability_mode_admits(
                agent_context.capability_mode,
                getattr(tool, "capability", None) or "execute",
            )
        )

    # Self-steering model control. Bound whenever the operator opts into
    # self-steering fallback — unlike a task tool it is resilience
    # infrastructure (peer of the automatic fallback ladder), so it is NOT
    # ceiling-checked against ``allowed_tools``: a child pinned by its
    # AgentDef to a model the key rejects (the live failure) must be able to
    # switch even though that AgentDef never named the tool, and is attached
    # at EVERY depth, not just the root. Off by default (self_steering=False)
    # ⇒ zero extra surface, so it never widens a strict scope uninvited.
    if session_id is not None and bool(
        get_config_value("llm", "fallback", "self_steering", default=False)
    ):
        from mewbo_core.tooling.model_control import ModelControlTool

        self._session_tools.append(
            ModelControlTool(
                session_id=session_id,
                agent_id=agent_context.agent_id,
                depth=agent_context.depth,
                get_active_model=lambda: self._active_model,
                get_ladder=lambda: [self._ctx.model_name, *self._ctx.fallback_models],
                get_strategy=lambda: self._retry_strategy,
                has_unanswered_tool_use=self._has_dangling_tool_use,
                apply_switch=self._request_model_switch,
                # Route through ``_emit_event`` (not the raw sink) so a
                # deliberate switch's ``llm_fallback`` is captured by the
                # resilience note too, keeping it the complete switch record.
                event_logger=self._emit_event,
                get_step=lambda: self._llm_call_step,
                max_switches=int(
                    get_config_value("llm", "fallback", "max_switches", default=2)
                ),
                allow_upgrade=bool(
                    get_config_value("llm", "fallback", "allow_upgrade", default=False)
                ),
            )
        )

llm_call_stall_age(now: float) -> float | None

Seconds the in-flight LLM call has been outstanding, else None.

Pure read over (_llm_call_started_at, now) with no clock of its own, so the liveness rule is exercisable at any age without waiting for one — the watchdog supplies time.monotonic(), a test supplies a number.

Source code in packages/mewbo_core/src/mewbo_core/loop/tool_use_loop.py
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
def llm_call_stall_age(self, now: float) -> float | None:
    """Seconds the in-flight LLM call has been outstanding, else ``None``.

    Pure read over ``(_llm_call_started_at, now)`` with no clock of its own,
    so the liveness rule is exercisable at any age without waiting for one
    — the watchdog supplies ``time.monotonic()``, a test supplies a number.
    """
    started = self._llm_call_started_at
    if started is None:
        return None
    return max(0.0, now - started)

rebind_workspace(entry: ProjectEntry) -> dict[str, object]

Re-point this loop at entry's directory. THE one mutation seam.

Everything derived from the working directory moves together here, because a partial switch is worse than none: the agent believes it moved and half the machinery did not. Returns the facts switch_project renders back to the model.

The step that is easiest to miss is the SPAWN seam. :class:~mewbo_core.agents.spawn_agent.SpawnAgentTool holds its OWN copy of the workspace and forwards it verbatim to every child loop it builds, so a switch that updates only self._cwd leaves every sub-agent spawned afterwards working in the OLD project — silently, and with results that look plausible right up until they are applied to the wrong tree.

Two deliberate limits, stated so nobody reads them as bugs:

  • The spec set NARROWS to what this run already held. A switch must never widen the tool ceiling the caller granted at run start, so a new project's own MCP servers are not admitted mid-run; a fresh run against that project resolves them normally, and the context event written here is what makes that fresh run land in the right place.
  • Skills ACCUMULATE rather than being replaced. The registry also holds plugin-contributed and user-level skills that a fresh scan of the new directory alone would drop, and losing those costs more than carrying the previous project's.
Source code in packages/mewbo_core/src/mewbo_core/loop/tool_use_loop.py
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
2028
2029
2030
2031
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
2100
def rebind_workspace(self, entry: ProjectEntry) -> dict[str, object]:
    """Re-point this loop at *entry*'s directory. THE one mutation seam.

    Everything derived from the working directory moves together here,
    because a partial switch is worse than none: the agent believes it moved
    and half the machinery did not. Returns the facts ``switch_project``
    renders back to the model.

    The step that is easiest to miss is the SPAWN seam.
    :class:`~mewbo_core.agents.spawn_agent.SpawnAgentTool` holds its OWN copy of the
    workspace and forwards it verbatim to every child loop it builds, so a
    switch that updates only ``self._cwd`` leaves every sub-agent spawned
    afterwards working in the OLD project — silently, and with results that
    look plausible right up until they are applied to the wrong tree.

    Two deliberate limits, stated so nobody reads them as bugs:

    - The spec set NARROWS to what this run already held. A switch must
      never widen the tool ceiling the caller granted at run start, so a new
      project's own MCP servers are not admitted mid-run; a fresh run
      against that project resolves them normally, and the ``context`` event
      written here is what makes that fresh run land in the right place.
    - Skills ACCUMULATE rather than being replaced. The registry also holds
      plugin-contributed and user-level skills that a fresh scan of the new
      directory alone would drop, and losing those costs more than carrying
      the previous project's.
    """
    new_cwd = entry.path
    if not new_cwd:
        raise ValueError(f"Project '{entry.key}' has no directory to switch into.")

    # Captured before anything moves. ``_project_key`` is ``None`` until the
    # first switch: this loop is handed a DIRECTORY at construction and never
    # a catalog key (the app resolves the key to a path before calling), so
    # there is genuinely no prior key to report on the first move.
    previous_cwd = self._cwd
    previous_project = self._project_key
    self._project_key = entry.key

    self._cwd = new_cwd
    self._containment = self._build_containment(new_cwd)

    # Resolved BEFORE the spawn seam is re-pointed: children inherit the
    # instructions as a plain string captured at build time, so handing the
    # tool the previous project's text would pin every future child to it.
    self._project_instructions = discover_project_instructions(new_cwd)
    if self._spawn_agent_tool is not None:
        self._spawn_agent_tool.rebind_cwd(
            new_cwd, project_instructions=self._project_instructions
        )

    # Cached by cwd, so switching back to a project visited earlier in the
    # run costs a dict lookup rather than a second registry build.
    self._tool_registry = get_or_build_registry(cwd=new_cwd)
    previous_ids = {spec.tool_id for spec in self._tool_specs_full}
    self._tool_specs_full = [
        spec for spec in self._tool_registry.list_specs() if spec.tool_id in previous_ids
    ]

    skills = 0
    if self._enable_skills and self._skill_registry is not None:
        self._skill_registry.load(new_cwd)
        skills = len(self._skill_registry.list_all())

    # Bind now so the tool can report a truthful count, but leave the
    # system-prompt re-render and the swap of the live model to the turn
    # boundary — the same split ``_requested_switch`` makes, which keeps the
    # transcript tail untouched while a tool call is still in flight.
    active_specs = self._select_active_specs(
        self._tool_specs_full,
        discovered=self._discovered_from_messages(self._live_messages or []),
    )
    tool_schemas = self._build_tool_schemas_for_mode(active_specs, self._current_mode)
    self._pending_workspace_bind = (tool_schemas, self._bind_model(tool_schemas))
    self._last_active_ids = {spec.tool_id for spec in active_specs}

    # A plain ``context`` event carrying the two keys the API's session-cwd
    # resolver already walks backwards for. Re-engagement via /query,
    # /message, the diff endpoints and SessionSpec reconstruction therefore
    # all follow the switch with no new resolver and no new event type;
    # inventing one would leave every one of them reading the project the
    # session STARTED in.
    #
    # It carries the PREVIOUS payload forward because a context event is
    # read two incompatible ways in this tree. One camp folds every event
    # (``SessionStoreBase.merge_context_events``); the other takes the most
    # recent payload VERBATIM — the API's ``_load_last_context``, which
    # `/message` re-engage and `/recover` read for the model, the mode, the
    # tool grants and the step budget, and the console's ``getLastContext``,
    # which the model pill and the composer hydrate from. Emitting the two
    # switch keys ALONE would therefore not merely mislabel those readers,
    # it would BLANK them. The same shape has bitten here before: a
    # model-only context event once left every verbatim reader seeing a
    # session with no tool ceiling, no playbook and no cwd.
    #
    # Carrying the LAST payload (not a fold of all of them) is what makes
    # this exactly neutral: before the switch the effective context was that
    # payload, after it is that payload plus the two keys the switch really
    # changed. A fold would instead resurrect a field an earlier turn had
    # deliberately cleared — clients omit a cleared field rather than
    # sending null, which is precisely why the verbatim readers are verbatim.
    payload: dict[str, object] = {"project": entry.key, "cwd": new_cwd}
    if self._session_context_reader is not None:
        try:
            previous = self._session_context_reader()
        except Exception as exc:  # noqa: BLE001 - never fail a switch over carry-forward
            logging.warning("Could not read session context to carry forward: {}", exc)
        else:
            if isinstance(previous, dict):
                payload = {**previous, **payload}
    self._emit_event({"type": "context", "payload": payload})

    return {
        "cwd": new_cwd,
        "previous_cwd": previous_cwd,
        "previous_project": previous_project,
        "project_instructions_found": bool(self._project_instructions),
        "bound_tools": self._bound_tool_count,
        "skills": skills,
    }

run(user_query: str, *, tool_specs: list[ToolSpec], context: ContextSnapshot | None = None, plan: Plan | None = None, mode: str = 'act') -> tuple[TaskQueue, OrchestrationState] async

Run the async tool-use loop and return TaskQueue + OrchestrationState.

Source code in packages/mewbo_core/src/mewbo_core/loop/tool_use_loop.py
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
async def run(
    self,
    user_query: str,
    *,
    tool_specs: list[ToolSpec],
    context: ContextSnapshot | None = None,
    plan: Plan | None = None,
    mode: str = "act",
) -> tuple[TaskQueue, OrchestrationState]:
    """Run the async tool-use loop and return TaskQueue + OrchestrationState."""
    state = OrchestrationState(goal=user_query)
    # Plan-mode is enforced via: (1) filtered tool schema at bind time,
    # (2) path-scoped permission check on edits, (3) the exit_plan_mode
    # approval gate. The loop flips ``_current_mode`` to ``"act"`` after
    # the user approves a plan and re-binds tools.
    self._current_mode = mode if mode in {"plan", "act"} else "act"
    if self._current_mode == "plan" and self._session_id is not None:
        state.plan_path = plan_file_for(self._session_id)
        ensure_plan_dir(self._session_id)
    # Propagate plan context so children inherit session and mode.
    if self._spawn_agent_tool is not None:
        self._spawn_agent_tool.session_id = self._session_id
        self._spawn_agent_tool.parent_mode = self._current_mode
        # The EFFECTIVE set, not the deferral-active subset: deferral strips
        # schemas from the initial bind and re-fetches them through
        # tool_search, so a child seeded from it would lose tools this agent
        # genuinely holds. Children narrow this set; they never widen it.
        self._spawn_agent_tool.parent_tool_specs = list(tool_specs)
    executed_steps: list[ActionStep] = []
    tool_outputs: list[str] = []
    last_error: str | None = None
    final_response: str | None = None

    # Register self in the hypervisor registry.
    # Reuse the handle created by SpawnAgentTool when one already exists for
    # this agent_id — avoids overwriting it and losing the reference held by
    # the lifecycle manager (which later stores AgentResult on the handle).
    existing = await self._ctx.registry.get(self._ctx.agent_id)
    if existing is not None:
        handle = existing
    else:
        handle = AgentHandle(
            agent_id=self._ctx.agent_id,
            parent_id=self._ctx.parent_id,
            depth=self._ctx.depth,
            model_name=self._ctx.model_name,
            task_description=user_query[:200],
            # A spawned child's handle is normally
            # pre-registered (and pre-stamped) by SpawnAgentTool before its
            # loop ever runs, so this branch is the exception (a directly
            # constructed loop, e.g. the root). Stamping ``self._contract``
            # here too keeps the handle self-consistent with whatever this
            # loop was actually built with, instead of silently reading
            # back the disabled default.
            contract=self._contract or DelegationContract(),
        )
        await self._ctx.registry.register(handle)
    handle.status = "running"

    # Background watchdog for stall detection.
    # Code-level reflex — zero token cost. Root only.
    watchdog_task: asyncio.Task[None] | None = None
    if self._ctx.depth == 0:
        watchdog_task = asyncio.create_task(self._watchdog())

    # Langfuse context managers — initialized in try, cleaned in finally.
    agent_span: Any = None
    _agent_span_cm: Any = None
    _propagate_cm: Any = None

    try:
        # Global eye for root agent
        agent_tree = ""
        if self._ctx.depth == 0:
            agent_tree = await self._ctx.registry.render_agent_tree(
                exclude_agent_id=self._ctx.agent_id,
            )
        # Deferred-tool partitioning. When ``agent.tool_search.mode`` is
        # 'on', MCP / metadata.deferred specs are stripped from the
        # initial bind and surfaced by name only via the
        # ``<available-deferred-tools>`` block. The model fetches the
        # schemas it needs through ``tool_search``; the per-turn re-bind
        # hook below grows the bound list as tools are discovered.
        self._tool_search_enabled = self._is_tool_search_enabled(tool_specs)
        self._tool_specs_full = list(tool_specs)
        self._deferred_ids = (
            {s.tool_id for s in tool_specs if is_deferred(s)}
            if self._tool_search_enabled
            else set()
        )
        active_specs = self._select_active_specs(tool_specs, discovered=set())
        messages = self._build_messages(
            user_query,
            context,
            plan,
            agent_tree=agent_tree,
        )
        # Expose the live transcript so the model_control continuity-lock
        # guardrail can inspect it (a mutable reference — it sees appends).
        self._live_messages = messages
        tool_schemas = self._build_tool_schemas_for_mode(
            active_specs,
            self._current_mode,
        )
        model = self._bind_model(tool_schemas)
        self._last_active_ids = {s.tool_id for s in active_specs}

        langfuse_handler = build_langfuse_handler(
            # The loop does NOT author trace identity. The real session and
            # principal arrive from the enclosing session context via
            # ``propagate_attributes`` and override anything the handler
            # carries — so a value invented here is not an override, it is
            # contradictory garbage sitting in every observation's metadata.
            # An empty string attaches no metadata key at all, which is the
            # honest reading of "this seam does not know". The session id is
            # passed only where the loop genuinely holds one.
            user_id="",
            session_id=self._session_id or "",
            trace_name="mewbo-tool-use",
            version=get_version(),
            release=get_config_value("runtime", "envmode", default="Not Specified"),
        )
        invoke_config: dict[str, Any] = {}
        if langfuse_handler is not None:
            invoke_config["callbacks"] = [langfuse_handler]
            metadata = getattr(langfuse_handler, "langfuse_metadata", None)
            if isinstance(metadata, dict) and metadata:
                invoke_config["metadata"] = metadata

        # -- Langfuse: agent-level span + attribute propagation --------
        # Typed ``agent`` so the trace renders as an agent graph, and named
        # from the AgentDef rather than the runtime handle id: a per-run hex
        # id is unbounded cardinality and folds every aggregation into
        # one-bucket-per-run. The handle id keeps its place in metadata.
        _agent_def_name = await self._agent_def_name()
        _agent_span_cm = langfuse_trace_span(
            f"invoke_agent {_agent_def_name}",
            as_type="agent",
            attributes={
                "gen_ai.operation.name": "invoke_agent",
                "gen_ai.agent.name": _agent_def_name,
            },
            metadata={
                "agentid": self._ctx.agent_id[:12],
                "model": self._ctx.model_name,
                "depth": str(self._ctx.depth),
                "mode": self._current_mode,
            },
            input_data={"task": user_query[:200]},
        )
        agent_span = _agent_span_cm.__enter__()
        _propagate_cm = langfuse_propagate(
            tags=[
                "mewbo-tool-use",
                f"model:{self._ctx.model_name}",
                f"depth:{self._ctx.depth}",
            ]
        )
        _propagate_cm.__enter__()

        turns = 0
        self._emit_safety_disclosure()
        # One atomic resilience strategy per run — holds the retry budget,
        # circuit breaker and policy knobs; survives every turn. Retained on
        # the instance so the model_control tool reuses THIS run's budget +
        # circuit breaker for its own switch guardrails.
        retry_strategy = RetryStrategy.from_config()
        self._retry_strategy = retry_strategy
        doom_guard = DoomLoopGuard.from_config(
            extra_poll_rules=self._poll_class_rules(tool_specs)
        )
        # Two-gate arming, computed ONCE for the run: the write-progress
        # signal is only meaningful for an agent that could plausibly
        # WRITE at all — narrowed by its own capability_mode AND actually
        # holding a write-tier tool. Anything else (read_only children, a
        # session with no write tools bound) leaves the signal permanently
        # inert.
        write_capable = self._ctx.capability_mode in {"execute", "all"} and any(
            s.capability_tier() == "write" for s in tool_specs
        )
        write_progress = WriteProgressSignal.from_config(write_capable=write_capable)
        write_progress_stamped = False
        # One-shot latch so the per-agent budget warning
        # (distinct from the session-wide ``loop.budget_warning`` above)
        # fires exactly once per run, not on every turn inside headroom.
        contract_step_warned = False
        # ``(tool_id, code)`` of the most recent blocked-class envelope
        # error that has NOT since been cleared. "Unrecovered" is judged per
        # tool: a later SUCCESS from the same tool means the blocked
        # operation went through after all, and only that clears it —
        # an unrelated tool succeeding says nothing about whether the repo
        # ever became reachable.
        last_blocked: tuple[str, str] | None = None
        # Consecutive promise-as-completion refusals (see the gate below).
        promise_nudges = 0
        # Consecutive required-terminal refusals (see the gate below).
        required_terminal_nudges = 0
        while not state.done:
            # Check cancellation. This is the between-turns read; the same
            # signal also guards the model call and each tool batch below,
            # so a stop never waits out whichever of those is in flight.
            if self._cancellation.requested:
                state.done = True
                state.done_reason = "canceled"
                break

            # Safety-plane observer — the between-turns half of the plane.
            # A scalar check, so its cost does not grow with run length.
            if self._check_safety_turn(turns):
                state.done = True
                state.done_reason = "safety_blocked"
                break

            # Check for interrupt (root agent only).
            if self._ctx.interrupt_step is not None and self._ctx.interrupt_step.is_set():
                self._ctx.interrupt_step.clear()
                messages.append(
                    HumanMessage(
                        content=get_prompt_registry().render("loop.interrupt_marker")
                    )
                )

            # Drain any queued user steering messages (root agent only).
            if self._ctx.message_queue is not None:
                while not self._ctx.message_queue.empty():
                    try:
                        msg = self._ctx.message_queue.get_nowait()
                        messages.append(HumanMessage(content=msg))
                    except _queue_mod.Empty:
                        break

            # Apply a deliberate model switch the model_control tool
            # requested last turn. Delegated to the SAME escalation path a
            # sticky fallback uses (promote active model / re-render prompt /
            # re-derive edit tool / rebind), applied at this turn boundary so
            # the transcript tail is untouched. Pin it on the strategy too, so
            # ``_order_models`` keeps the chosen model at the chain head
            # instead of a prior sticky pin reordering it back out.
            if self._requested_switch is not None:
                _switch_target = self._requested_switch
                self._requested_switch = None
                retry_strategy._pinned_model = _switch_target
                tool_schemas, model = self._apply_model_escalation(
                    _switch_target,
                    messages,
                    context=context,
                    plan=plan,
                    agent_tree=agent_tree,
                    tool_schemas=tool_schemas,
                    model=model,
                )

            # Adopt a workspace switch ``switch_project`` performed mid-turn.
            # ``rebind_workspace`` already moved every piece of loop state
            # and built the new binding; what it deliberately left is the
            # transcript rewrite, taken here at the turn boundary exactly as
            # a requested model switch is. Re-rendering ``messages[0]`` is
            # what carries the new project's instructions, environment block
            # and git context into the model's next generation.
            if self._pending_workspace_bind is not None:
                tool_schemas, model = self._pending_workspace_bind
                self._pending_workspace_bind = None
                messages[0] = SystemMessage(
                    content=self._render_system_prompt(context, plan, agent_tree)
                )
                self._active_resilience_note = self._resilience_note.render()

            # Keep the resilience note current in the system prompt: re-render
            # ``messages[0]`` only when the note text changed since it was last
            # baked in, so a clean run never pays and a healing run re-renders
            # at most once per new retry/fallback event.
            _note = self._resilience_note.render()
            if _note != self._active_resilience_note:
                self._active_resilience_note = _note
                messages[0] = SystemMessage(
                    content=self._render_system_prompt(context, plan, agent_tree)
                )

            # The span stays — it is the natural parent of this turn's
            # generation and tool calls — but its NAME must not carry the
            # turn counter: a per-execution integer in a name is unbounded
            # cardinality, so "how long does a step take" had as many
            # groups as the longest run has turns. The index is a metadata
            # field, which is where a filter can still reach it.
            with langfuse_trace_span(
                "agent_step",
                metadata={
                    "turn": str(turns),
                    "model": self._ctx.model_name,
                },
            ) as span:
                if span is not None:
                    try:
                        span.update_trace(input={"turn": turns, "message_count": len(messages)})
                    except Exception:
                        pass

                # Graduated enforcement: warn as the budget nears, then force
                # ONE wrap-up turn at exhaustion so an unbounded
                # fan-out can't run away — but the agent still gets to
                # answer instead of a bare halt.
                if self._ctx.registry.budget_exhausted():
                    final_response = await self._budget_wrapup_turn(
                        "budget_exhausted",
                        state=state,
                        messages=messages,
                        tool_outputs=tool_outputs,
                        invoke_config=invoke_config,
                    )
                    break
                if self._ctx.registry.budget_warning():
                    messages.append(
                        SystemMessage(
                            content=get_prompt_registry().render("loop.budget_warning")
                        )
                    )

                # DelegationContract. A per-agent bound
                # LAYERED UNDER the session budget just checked above:
                # checked here regardless (a spawner's ceiling applies even
                # when the shared pool has headroom left). Disabled
                # contracts (the default) skip this entirely.
                if self._contract is not None and self._contract.enabled:
                    contract_over = False
                    step_state = await self._ctx.registry.agent_step_state(
                        self._ctx.agent_id
                    )
                    if step_state == "over":
                        contract_over = True
                    elif step_state == "warn" and not contract_step_warned:
                        contract_step_warned = True
                        messages.append(
                            SystemMessage(
                                content=get_prompt_registry().render(
                                    "loop.agent_budget_warning"
                                )
                            )
                        )
                    if not contract_over:
                        token_state = await self._ctx.registry.agent_token_state(
                            self._ctx.agent_id
                        )
                        if token_state == "over":
                            contract_over = True
                    if contract_over:
                        final_response = await self._budget_wrapup_turn(
                            "halted_agent_budget",
                            state=state,
                            messages=messages,
                            tool_outputs=tool_outputs,
                            invoke_config=invoke_config,
                        )
                        break

                # Heartbeat events so clients (console/CLI) can distinguish
                # "waiting on LLM" from a silent hang. ``bound_tools`` is the
                # size of the surface this call carries — the transcript
                # recorded nothing about the bound set anywhere, so a step
                # that silently bound a collapsed one was only diagnosable by
                # reading the model's behaviour back. A COUNT is deliberate:
                # the full name list on every step of every run is payload
                # bloat on a hot path, and a drop shows up in the count.
                self._emit_event(
                    {
                        "type": "llm_call_start",
                        "payload": {
                            "agent_id": self._ctx.agent_id,
                            "depth": self._ctx.depth,
                            "step": turns,
                            "model": self._active_model,
                            "bound_tools": self._bound_tool_count,
                        },
                    }
                )
                # Own clock for the successful ``llm_call_end`` payload's
                # ``duration_ms``, separate from the liveness leg below: that
                # one is re-armed per retry/fallback attempt
                # (``_invoke_with_resilience``), so it cannot bracket the
                # whole logical call the way this single capture does.
                _llm_call_t0 = _time.monotonic()
                # Arm the liveness leg for exactly the window this call is
                # outstanding; the ``finally`` disarms it on every exit so a
                # completed call can never read as a wedged one.
                self._llm_call_started_at = _time.monotonic()
                self._llm_call_step = turns
                try:
                    # Guarded: the resilience ladder can spend minutes across
                    # retries and fallbacks, so an unguarded await would keep
                    # a stop invisible until the ladder settled.
                    response, _final_model = await self._cancellation.guard(
                        self._invoke_with_resilience(
                            primary_model=model,
                            messages=messages,
                            tool_schemas=tool_schemas,
                            turns=turns,
                            invoke_config=invoke_config,
                            strategy=retry_strategy,
                        )
                    )
                except RunCancelled:
                    state.done = True
                    state.done_reason = "canceled"
                    break
                except LlmResilienceExhausted as exhausted:
                    # Clean halt: surface a true failure (never masked as
                    # "completed") so the FE can offer one-click recovery.
                    self._emit_event(
                        {
                            "type": "llm_call_end",
                            "payload": {
                                "agent_id": self._ctx.agent_id,
                                "depth": self._ctx.depth,
                                "step": turns,
                                "success": False,
                                "model": (
                                    exhausted.models_tried[-1]
                                    if exhausted.models_tried
                                    else self._ctx.model_name
                                ),
                                "error_type": exhausted.last_error_type,
                                "reason": exhausted.reason,
                            },
                        }
                    )
                    if span is not None:
                        try:
                            span.update(
                                level="ERROR",
                                # Same substitution as the completion string:
                                # ``str(TimeoutError())`` is empty, and this
                                # span write would otherwise carry a void
                                # status_message for the very failure it marks.
                                status_message=LlmResilienceExhausted.describe_error(
                                    exhausted.last_error
                                ),
                                metadata={
                                    "errortype": exhausted.last_error_type,
                                    "models_tried": ",".join(exhausted.models_tried),
                                    "reason": exhausted.reason,
                                },
                            )
                        except Exception:
                            pass
                    record_span_exception(
                        span,
                        exhausted.last_error,
                        attributes={
                            "errortype": exhausted.last_error_type or "",
                            "reason": exhausted.reason or "",
                        },
                    )
                    raise
                finally:
                    self._llm_call_started_at = None
                # Fallback ladder: if the resilience strategy escalated
                # to (and pinned) a different model, re-render the system
                # prompt + re-derive the edit-tool variant against THAT model
                # so the heal is behavioural, not just a model swap.
                tool_schemas, model = self._apply_model_escalation(
                    _final_model,
                    messages,
                    context=context,
                    plan=plan,
                    agent_tree=agent_tree,
                    tool_schemas=tool_schemas,
                    model=model,
                )
                _step_usage = getattr(response, "usage_metadata", None)
                _h_ref = await self._ctx.registry.get(self._ctx.agent_id)
                # LangChain ``UsageMetadata`` exposes provider cache and
                # reasoning subtotals (Anthropic + OpenAI normalised):
                #   input_token_details.cache_creation — written to cache
                #     this call (Anthropic 5-min: 1.25× input price)
                #   input_token_details.cache_read — served from cache
                #     (Anthropic: 0.1× input; OpenAI: 0.5× input)
                #   output_token_details.reasoning — extended-thinking /
                #     o1 hidden tokens (billed as output)
                # Capturing them per call lets clients show fresh-vs-
                # cached breakdown and an honest billable signal that
                # accounts for cache discounts.
                _in_det = _step_usage.get("input_token_details") or {} if _step_usage else {}
                _out_det = _step_usage.get("output_token_details") or {} if _step_usage else {}
                self._emit_event(
                    {
                        "type": "llm_call_end",
                        "payload": {
                            "agent_id": self._ctx.agent_id,
                            "depth": self._ctx.depth,
                            "step": turns,
                            "success": True,
                            "model": _final_model,
                            "input_tokens": (
                                _step_usage.get("input_tokens", 0) if _step_usage else 0
                            ),
                            "output_tokens": (
                                _step_usage.get("output_tokens", 0) if _step_usage else 0
                            ),
                            "cache_creation_input_tokens": int(
                                _in_det.get("cache_creation", 0) or 0
                            ),
                            "cache_read_input_tokens": int(_in_det.get("cache_read", 0) or 0),
                            "reasoning_output_tokens": int(_out_det.get("reasoning", 0) or 0),
                            "cumulative_input_tokens": (_h_ref.input_tokens if _h_ref else 0),
                            "cumulative_output_tokens": (_h_ref.output_tokens if _h_ref else 0),
                            "duration_ms": int((_time.monotonic() - _llm_call_t0) * 1000),
                        },
                    }
                )
                # Strip thinking blocks from the response before appending
                # to the conversation history.  Anthropic requires a
                # ``signature`` field on thinking blocks when replayed,
                # but proxies (LiteLLM) may not preserve it.
                raw = getattr(response, "content", None)
                if isinstance(raw, list):
                    # A reasoning model's answer arrives as a BARE STRING
                    # element alongside ``{"type": "thinking", ...}``
                    # dicts, not as a proper content part. Left as-is, a
                    # strict OpenAI-shaped backend (a self-hosted Ollama
                    # behind LiteLLM) rejects the replayed history with
                    # 400 "invalid message format" the moment the turn is
                    # replayed on a later request.
                    sanitized: list[str | dict[Any, Any]] = [
                        block if isinstance(block, dict) else {"type": "text", "text": block}
                        for block in raw
                        if not (isinstance(block, dict) and block.get("type") == "thinking")
                    ]
                    # Never leave empty-string assistant content.
                    # Empty assistant turns in
                    # history cause extended-thinking models to
                    # hallucinate framework-style placeholders.
                    if not sanitized:
                        sanitized = [{"type": "text", "text": _NO_CONTENT_PLACEHOLDER}]
                    response = AIMessage(
                        content=sanitized,
                        tool_calls=response.tool_calls,
                        additional_kwargs=response.additional_kwargs,
                        usage_metadata=response.usage_metadata,
                        id=response.id,
                    )
                elif response.tool_calls and (not raw or not str(raw).strip()):
                    # The proxy (LiteLLM) strips thinking blocks itself
                    # and returns ``content=""`` (a STRING, not a list).
                    # Without this branch, the empty string survives
                    # sanitisation and gets replayed in history, causing
                    # the model to hallucinate placeholder meta-text.
                    response = AIMessage(
                        content=_NO_CONTENT_PLACEHOLDER,
                        tool_calls=response.tool_calls,
                        additional_kwargs=response.additional_kwargs,
                        usage_metadata=response.usage_metadata,
                        id=response.id,
                    )
                messages.append(response)

                if not response.tool_calls:
                    # Text response — the model claims completion.
                    content = self._extract_text_content(getattr(response, "content", ""))
                    # Verifier gate: run the ground-truth check BEFORE
                    # accepting the claim (only when armed and not already
                    # green). A failure with a retry left injects the
                    # grounded verifier output and ``continue``s — funnelling
                    # back through the top-of-loop budget checks FIRST, so
                    # retries are bounded by BOTH verification_max_retries AND
                    # the step/wall budget. Exhausted retries accept the text
                    # but flag it honestly (``verification_failed``); a pass
                    # falls through to the normal ``completed`` accept.
                    if self._verification_active and not self._verify_passed:
                        state.verify_attempts += 1
                        outcome = await self._run_verifier(
                            step=turns, attempt=state.verify_attempts
                        )
                        if outcome.passed:
                            self._verify_passed = True
                            state.verified = True
                        elif self._verify_retries_left > 0:
                            self._verify_retries_left -= 1
                            messages.append(
                                SystemMessage(
                                    content=get_prompt_registry().render(
                                        "loop.verification_failed",
                                        output=outcome.feedback,
                                    )
                                )
                            )
                            continue
                        else:
                            final_response = content
                            tool_outputs.append(content)
                            state.done = True
                            state.done_reason = "verification_failed"
                            state.verified = False
                            break
                    # Promise-as-completion gate (root only). A clean
                    # terminal declared while the root's OWN background runs
                    # are still live is a promise about future work — "I'll
                    # check back shortly" with an unfinished probe fleet
                    # behind it — not a completion. Refuse it here, upstream
                    # of the orchestrator's honesty downgrade, and send the
                    # model to await its runs. ``collect_running`` is keyed on
                    # THIS agent so it counts only owned CHILDREN — the seam
                    # deliberately does NOT read the session-wide ownership
                    # index, because the root's own handle is still
                    # ``running`` here (it is marked done only in the loop's
                    # finally), so that index would count the root itself and
                    # refuse every terminal. Bounded, so a model that will not
                    # wait is eventually let through rather than burning the
                    # whole budget spinning.
                    if self._ctx.depth == 0 and promise_nudges < _PROMISE_GATE_MAX_NUDGES:
                        live_owned = await self._ctx.registry.collect_running(
                            self._ctx.agent_id
                        )
                        if live_owned:
                            promise_nudges += 1
                            messages.append(
                                SystemMessage(
                                    content=get_prompt_registry().render(
                                        "loop.agents_still_running",
                                        count=len(live_owned),
                                        ids=", ".join(
                                            h.agent_id[:8] for h in live_owned
                                        ),
                                    )
                                )
                            )
                            continue
                    # Required-terminal gate (root only). A session tool can
                    # declare that ending the run REQUIRES a call to it —
                    # ``wiki_emit_answer`` is the canonical case: a composed
                    # answer delivered as plain text instead of through the
                    # tool is discarded downstream, and the run has no other
                    # channel to tell the user it happened. Refuse the clean
                    # terminal while a declared obligation is unmet and nudge
                    # the model to call it; bounded the same way the promise
                    # gate above is, so a model that genuinely cannot emit is
                    # eventually let through rather than spinning the whole
                    # budget here. UNLIKE the promise gate, exhaustion here
                    # stamps the honest ``unmet_goal`` rather than a claimed
                    # completion the user never saw.
                    if self._ctx.depth == 0:
                        unmet_terminals = self._unmet_required_terminals()
                        if unmet_terminals:
                            if required_terminal_nudges < _REQUIRED_TERMINAL_MAX_NUDGES:
                                required_terminal_nudges += 1
                                messages.append(
                                    SystemMessage(
                                        content=get_prompt_registry().render(
                                            "loop.required_terminal_missing",
                                            tool_ids=", ".join(unmet_terminals),
                                        )
                                    )
                                )
                                continue
                            final_response = content
                            tool_outputs.append(content)
                            state.done = True
                            state.done_reason = "unmet_goal"
                            if last_blocked is not None:
                                state.blocked_code = last_blocked[1]
                            break
                    final_response = content
                    tool_outputs.append(content)
                    state.done = True
                    state.done_reason = "completed"
                    # The accept is NOT unconditional: a run whose clone never
                    # authenticated would otherwise stamp a clean completion
                    # because its LAST turn happened to be text. Consult the
                    # unrecovered blocked-class error instead — it says the
                    # goal was unreachable this run whatever the closing
                    # prose claims. Carried as its OWN field, never as a new
                    # done_reason: that vocabulary is a wire contract shared
                    # with every client, and the status layer owns the
                    # mapping from this code to a user-facing state.
                    if last_blocked is not None:
                        state.blocked_code = last_blocked[1]
                    break

                # The model is repeating the same tool + input with no
                # progress. Halt cleanly and hand back for one-click
                # recovery instead of executing the same call again.
                doom_guard.observe(response.tool_calls)
                if doom_guard.is_stuck():
                    # Keep the transcript valid for resume: the AIMessage
                    # with tool_calls was already appended, so synthesize
                    # interrupted results for its dangling calls.
                    repair_tool_pairing(messages)
                    _repeated = response.tool_calls[0].get("name", "tool")
                    last_error = (
                        f"Halted: model repeated '{_repeated}' "
                        f"{doom_guard.threshold}x with identical input and result "
                        "(no progress)."
                    )
                    halt_payload: RecoveryHaltPayload = {
                        "action": "halt_no_progress",
                        "agent_id": self._ctx.agent_id,
                        "depth": self._ctx.depth,
                        "step": turns,
                        "tool": _repeated,
                    }
                    self._emit_event({"type": "recovery", "payload": halt_payload})
                    state.done = True
                    state.done_reason = "halted_no_progress"
                    break

                # Emit intermediate text as agent_message for trace logs.
                text_content = self._extract_text_content(getattr(response, "content", ""))
                if text_content:
                    self._emit_event(
                        {
                            "type": "agent_message",
                            "payload": {
                                "text": text_content,
                                "agent_id": self._ctx.agent_id,
                                "depth": self._ctx.depth,
                            },
                        }
                    )

                # Execute tool calls with concurrency-aware partitioning.
                # Exclusive tools run alone; concurrent-safe tools are gathered.
                specs_map = {s.tool_id: s for s in tool_specs}
                batches = self._partition_tool_calls(response.tool_calls, specs_map)
                results: list[ToolCallResult] = []
                # Guarded per batch: a step can hold several long tool calls,
                # and an unguarded loop here made the stop wait for ALL of
                # them. Breaking out mid-step leaves the already-executed
                # calls' ``tool_result`` events in the transcript and the rest
                # absent, which is the honest record of what actually ran —
                # nothing downstream re-drives the message list, because the
                # run terminates without another model call.
                cancelled_mid_step = False
                for batch in batches:
                    try:
                        if batch.concurrent:
                            batch_results = await self._cancellation.guard(
                                asyncio.gather(
                                    *[
                                        self._safe_execute(tc, tool_specs)
                                        for tc in batch.calls
                                    ],
                                )
                            )
                            results.extend(batch_results)
                        else:
                            results.append(
                                await self._cancellation.guard(
                                    self._safe_execute(batch.calls[0], tool_specs)
                                )
                            )
                    except RunCancelled:
                        cancelled_mid_step = True
                        break
                if cancelled_mid_step:
                    state.done = True
                    state.done_reason = "canceled"
                    break

                # Feed results back to the doom guard so "no progress" means
                # same input AND same outcome — a call whose result advances
                # (e.g. check_agents as children finish) is healthy progress.
                doom_guard.record_result(results)

                # Track the blocked-class condition across turns so the
                # completion seam can tell "finished" from "gave up".
                for _r in results:
                    if _r.blocked_code:
                        last_blocked = (_r.tool_id, _r.blocked_code)
                    elif (
                        _r.success
                        and last_blocked is not None
                        and _r.tool_id == last_blocked[0]
                    ):
                        last_blocked = None

                # Write-progress signal: distinct from the doom guard —
                # counts consecutive steps for a write-capable agent with
                # no WRITE-tier tool execution (read/execute/search/
                # unknown-tier steps only). Observe-only: crossing the
                # threshold always emits telemetry; the optional reminder
                # never names the signal or its criteria.
                write_progress.observe(results, specs_map)
                if write_progress.threshold_crossed():
                    self._emit_event(
                        {
                            "type": "write_progress_signal",
                            "payload": {
                                "agent_id": self._ctx.agent_id,
                                "depth": self._ctx.depth,
                                "step": turns,
                                "steps_since_write": write_progress.steps_since_write,
                                "threshold": write_progress.threshold,
                            },
                        }
                    )
                    if write_progress.reminder_enabled:
                        messages.append(
                            SystemMessage(
                                content=get_prompt_registry().render(
                                    "loop.task_objective_reminder", goal=state.goal
                                )
                            )
                        )
                elif write_progress.exhausted() and not write_progress_stamped:
                    # Go quiet after the last event — flag it once for the
                    # parent's global eye instead of firing forever.
                    write_progress_stamped = True
                    _h_write_progress = await self._ctx.registry.get(self._ctx.agent_id)
                    if _h_write_progress:
                        _h_write_progress.progress_note = (
                            "No write-tier tool call in the last "
                            f"{write_progress.steps_since_write} steps."
                        )

                for tool_call, result in zip(response.tool_calls, results):
                    messages.append(
                        ToolMessage(
                            # A multimodal result becomes a list of content
                            # parts, which LiteLLM translates into a native
                            # ``tool_result`` carrying the text block then
                            # the image block. Everything else stays the
                            # plain string it always was, so an ordinary
                            # result's cache prefix is byte-identical.
                            content=_tool_message_content(result),
                            tool_call_id=result.tool_call_id,
                        )
                    )
                    tool_outputs.append(f"{result.tool_id}: {result.content}")
                    if not result.success:
                        last_error = result.content
                    # Track as ActionStep for TaskQueue compatibility.
                    action_step = self._tool_call_to_action_step(tool_call)
                    mock = get_mock_speaker()
                    action_step.result = mock(content=result.content)
                    executed_steps.append(action_step)

                # Episodic plan-mode: a session tool (e.g. exit_plan_mode)
                # signals the loop to terminate so the thread exits
                # cleanly. Approval/rejection happens out-of-band via
                # SessionRuntime. Materialise the list so every tool's
                # flag is consumed — ``any`` short-circuits and would
                # leave a second tool's flag set for the next turn.
                # ``terminal_reason()`` lets each tool declare the right
                # done_reason (default: "awaiting_approval"; emit tool: "completed").
                term_tools = [
                    (t, t.should_terminate_run()) for t in self._session_tools
                ]
                terminating = [t for t, flag in term_tools if flag]
                if terminating:
                    state.done = True
                    state.done_reason = terminating[0].terminal_reason()
                    break

                await self._ctx.registry.update_step(
                    self._ctx.agent_id,
                    results[-1].tool_id if results else "",
                )
                turns += 1

                # Auto-update progress
                # for parent monitoring. Zero token cost — direct write.
                if self._ctx.depth > 0 and results:
                    last = results[-1]
                    note = f"turn {turns}: {last.tool_id}"
                    snippet = last.content[:100] if last.content else ""
                    if last.success:
                        note += f" -> {snippet}"
                    else:
                        note += f" -> FAILED: {snippet}"
                    handle = await self._ctx.registry.get(
                        self._ctx.agent_id,
                    )
                    if handle:
                        handle.progress_note = note

                # Inject failure feedback so the model can adapt.
                failures = [r for r in results if not r.success]
                if failures:
                    messages.append(
                        SystemMessage(
                            content=f"{len(failures)}/{len(results)} tool call(s)"
                            " failed this step — adapt your approach."
                        )
                    )

                # Re-bind newly discovered deferred tools. Discovery is
                # derived from the message history each turn, so this is
                # compaction-resilient: whatever survives compaction
                # still drives the bound set on the next iteration.
                if self._tool_search_enabled and self._deferred_ids:
                    discovered = self._discovered_from_messages(messages)
                    active_specs = self._select_active_specs(
                        self._tool_specs_full, discovered=discovered
                    )
                    new_active_ids = {s.tool_id for s in active_specs}
                    if new_active_ids != self._last_active_ids:
                        tool_schemas = self._build_tool_schemas_for_mode(
                            active_specs, self._current_mode
                        )
                        model = self._bind_model(tool_schemas)
                        self._last_active_ids = new_active_ids

                # Proactive mid-loop compaction check.
                if self._should_compact_messages(messages):
                    _compact_info = await self._compact_messages(messages)
                    if _compact_info:
                        await self._ctx.registry.record_compaction(
                            self._ctx.agent_id,
                        )
                        self._emit_event(
                            {
                                "type": "context_compacted",
                                "payload": {
                                    **_compact_info,
                                    "agent_id": self._ctx.agent_id,
                                    "depth": self._ctx.depth,
                                    "mode": "mid_loop",
                                    "turn": turns,
                                },
                            }
                        )

                if span is not None:
                    try:
                        span.update_trace(
                            output={
                                "tool_calls": len(response.tool_calls),
                                "turns": turns,
                            }
                        )
                    except Exception:
                        pass

        # Safety net: currently unreachable (all loop exits set
        # state.done=True), but retained as defensive code for future
        # exit paths that may break without setting state.done.
        if not state.done and final_response is None and messages:
            # Inject child results at synthesis.
            # Root may have non-blocking children still running — give them
            # a brief grace period, then include available results.
            if self._ctx.depth == 0:
                running = await self._ctx.registry.collect_running(
                    self._ctx.agent_id,
                )
                if running:
                    await asyncio.sleep(2.0)
                completed = await self._ctx.registry.collect_completed(
                    self._ctx.agent_id,
                )
                if completed:
                    result_lines = []
                    for h in completed:
                        r = h.result
                        if r:
                            result_lines.append(
                                f"[{h.agent_id[:8]}] {r.status}: {r.summary or r.content[:300]}"
                            )
                    if result_lines:
                        messages.append(
                            SystemMessage(
                                content=get_prompt_registry().render(
                                    "loop.agent_results_header",
                                    joined="\n".join(result_lines),
                                ),
                            )
                        )
                still_running = await self._ctx.registry.collect_running(
                    self._ctx.agent_id,
                )
                if still_running:
                    ids = ", ".join(h.agent_id[:8] for h in still_running)
                    messages.append(
                        SystemMessage(
                            content=get_prompt_registry().render(
                                "loop.agents_still_running",
                                count=len(still_running),
                                ids=ids,
                            ),
                        )
                    )

            final_response = await self._unbound_wrapup_invoke(
                prompt_key="loop.final_answer_synthesis",
                model_name=self._ctx.model_name,
                messages=messages,
                tool_outputs=tool_outputs,
                invoke_config=invoke_config,
            )

        if not state.done:
            state.done = True
            state.done_reason = "completed"

        # Forced closing summary for a CHILD that finished via store
        # side-effects. The heaviest sub-agents did all their real work
        # through tool writes and then stopped on empty text, so the parent
        # — which projects ``task_result`` as the child's summary — received
        # nothing, and a 3.3M-token probe reached its parent blank. One
        # unbound wrap-up turn forces a compressed summary before stop.
        # Root is exempt: its empty terminal is a user-facing turn, not a
        # summary owed upstream. A cancel is exempt too — it is not a
        # completion to summarize. The budget/synthesis paths already filled
        # ``final_response``, so the empty-guard skips them (no double turn).
        if (
            self._ctx.depth > 0
            and state.done_reason != "canceled"
            and not (final_response or "").strip()
        ):
            final_response = await self._unbound_wrapup_invoke(
                prompt_key="loop.final_answer_synthesis",
                model_name=self._active_model,
                messages=messages,
                tool_outputs=tool_outputs,
                invoke_config=invoke_config,
            )

        # Build TaskQueue for compatibility with CLI / API consumers.
        plan_steps = list(plan.steps) if plan and plan.steps else []
        task_queue = TaskQueue(
            plan_steps=plan_steps,
            action_steps=executed_steps,
        )
        # task_result is the LLM's final synthesized text only.
        task_queue.task_result = (final_response or "").strip()
        task_queue.last_error = last_error
        state.tool_results = tool_outputs
        # Every OTHER terminal path (halt, budget wrap-up, verifier
        # exhaustion, plan-mode exit) carries the same fact — a blocked run
        # is blocked regardless of which exit it took.
        if state.blocked_code is None and last_blocked is not None:
            state.blocked_code = last_blocked[1]

    finally:
        # Close Langfuse agent span and propagation context.
        if agent_span is not None:
            try:
                agent_span.update(
                    output={
                        "total_steps": turns,
                        "done_reason": state.done_reason or "unknown",
                    }
                )
            except Exception:
                pass
        if _propagate_cm is not None:
            try:
                _propagate_cm.__exit__(None, None, None)
            except Exception:  # pragma: no cover - defensive
                pass
        if _agent_span_cm is not None:
            try:
                _agent_span_cm.__exit__(None, None, None)
            except Exception:  # pragma: no cover - defensive
                pass

        if watchdog_task is not None and not watchdog_task.done():
            watchdog_task.cancel()

        # Wait for lifecycle managers to complete cleanup.
        if self._spawn_agent_tool is not None:
            await self._spawn_agent_tool.await_lifecycle_managers(timeout=3.0)

        # Cleanup: cancel any child agent that has not reached a terminal.
        # Reads ACTIVE_STATUSES rather than comparing against ``running``
        # alone: a capacity-deferred child sits at ``submitted`` until a slot
        # frees, so a sweep filtering on the one literal walked past exactly
        # the children that had never got to run.
        children = await self._ctx.registry.list_children(self._ctx.agent_id)
        for child in children:
            if child.status in ACTIVE_STATUSES:
                await self._ctx.registry.cancel_agent(child.agent_id)

        # The terminal is PROJECTED, never re-derived here. ``state.done``
        # means the loop stopped, not that the task succeeded — it is True on
        # every exit path, cancellation and doom-halt included — so deriving
        # a status from it directly reported a clean success for a run the
        # user stopped. ``terminal_status()`` is the authority, and this is
        # the only mark a ROOT ever gets; a child was corrected a moment
        # later on the spawn path, which is why only the root kept the wrong
        # value permanently and why the child had a window where a reader
        # disagreed with the final answer. Projecting here settles both.
        await self._ctx.registry.mark_done(
            self._ctx.agent_id,
            state.terminal_status(),
        )

    return task_queue, state

mewbo_core.agents.agent_context

Immutable agent context propagated through the agent hierarchy.

AgentContext is the per-agent state carried by every ToolUseLoop instance. The root agent creates one via AgentContext.root(); child agents receive one via parent_ctx.child().

The hypervisor control plane (AgentHypervisor, AgentHandle) lives in :mod:mewbo_core.agents.hypervisor.

AgentContext dataclass

Immutable context propagated through the agent hierarchy.

Every ToolUseLoop instance requires an AgentContext. The root agent creates one via AgentContext.root(). Child agents receive one via parent_ctx.child().

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_context.py
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
@dataclass(frozen=True, slots=True)
class AgentContext:
    """Immutable context propagated through the agent hierarchy.

    Every ToolUseLoop instance requires an AgentContext. The root agent
    creates one via ``AgentContext.root()``. Child agents receive one
    via ``parent_ctx.child()``.
    """

    agent_id: str
    parent_id: str | None
    depth: int
    max_depth: int
    model_name: str
    should_cancel: Callable[[], bool] | None
    event_logger: Callable[[Event], None] | None
    registry: AgentHypervisor
    fallback_models: tuple[str, ...] = ()
    # Effective delegation privilege ceiling for THIS agent. Propagated
    # like ``fallback_models`` but MONOTONICALLY narrowed
    # at every hop — see :meth:`child`. ``"all"`` (root default) means no
    # capability filtering. A plain ``str`` (not the ``CapabilityMode``
    # Literal): this is hot in-process state that crosses no trust boundary,
    # so the validated Literal lives at the ``SpawnAgentTask`` seam instead.
    capability_mode: str = "all"
    # Effective FILESYSTEM-containment ceiling for THIS agent. The
    # second privilege axis, ORTHOGONAL to ``capability_mode`` and narrowed the
    # same way — monotonically, min-wins, at every ``child()`` hop (see below).
    # ``"full_access"`` (root default) means no path restriction; a narrower tier
    # confines reads (and, above ``read_only``, writes) to the agent's workspace.
    # A plain ``str`` for the same reason as ``capability_mode``: hot in-process
    # state crossing no trust boundary — the validated ``WorkspaceMode`` Literal
    # lives at the ``SpawnAgentTask`` seam. Enforcement is gated on
    # ``agent.workspace_enforcement`` (ON by default); an operator who turns that
    # off leaves this field carried and narrowed but governing nothing.
    workspace_mode: str = "full_access"
    # Delegation firebreak — set by
    # a DelegationContract(autonomy="atomic"). Propagated like
    # ``capability_mode`` but MONOTONICALLY: once set, every descendant stays
    # atomic too (see ``child``), so a grandchild can never re-enable
    # delegation an ancestor gave up.
    atomic: bool = False
    message_queue: queue.Queue[str] | None = None
    interrupt_step: threading.Event | None = None

    # Monotonic privilege ranking for ``capability_mode`` narrowing:
    # lower rank = more restrictive. Mirrors the tiers in
    # ``tool_registry._CAPABILITY_MODE_TIERS`` (kept consistent by a test).
    _CAPABILITY_MODE_RANK: ClassVar[dict[str, int]] = {
        "read_only": 0,
        "execute": 1,
        "all": 2,
    }

    # Monotonic privilege ranking for ``workspace_mode`` narrowing; same
    # min-wins shape, a second axis. Kept in lockstep with
    # ``mewbo_core.workspaces.workspace.WORKSPACE_MODE_RANK`` by a test.
    _WORKSPACE_MODE_RANK: ClassVar[dict[str, int]] = {
        "read_only": 0,
        "workspace_write": 1,
        "full_access": 2,
    }

    @property
    def can_spawn(self) -> bool:
        """True if this agent is allowed to create children."""
        return self.depth < self.max_depth

    @property
    def remaining_depth(self) -> int:
        """Number of spawn levels remaining below this agent."""
        return max(0, self.max_depth - self.depth)

    @staticmethod
    def _narrow(parent: str, requested: str, rank_table: dict[str, int]) -> str:
        """Return the more restrictive of two modes on ONE privilege axis.

        The shared min-wins narrowing behind BOTH ``capability_mode`` and
        ``workspace_mode``: privilege attenuation is monotonic down the
        agent hierarchy — a child can only ever narrow a mode, never widen it,
        so a grandchild under a restrictive ancestor
        can never climb back. An unrecognised mode collapses to the axis's WIDEST
        (no-op) tier — the highest-ranked key in *rank_table* (``all`` for
        capability, ``full_access`` for workspace) — so a stray string neither
        tightens surprisingly nor loosens below the parent; the authoritative
        validation is the ``Literal`` at the ``SpawnAgentTask`` seam upstream.
        """
        widest = max(rank_table, key=lambda k: rank_table[k])
        p = parent if parent in rank_table else widest
        r = requested if requested in rank_table else widest
        return p if rank_table[p] <= rank_table[r] else r

    @staticmethod
    def _narrower_capability_mode(parent: str, requested: str) -> str:
        """Narrow ``capability_mode`` — thin alias over :meth:`_narrow`.

        Retained as the named entry point (referenced by tests + call sites); the
        min-wins / unknown → ``all`` semantics are now the shared ``_narrow``.
        """
        return AgentContext._narrow(
            parent, requested, AgentContext._CAPABILITY_MODE_RANK
        )

    def child(
        self,
        *,
        model_name: str | None = None,
        capability_mode: str = "all",
        workspace_mode: str = "full_access",
        atomic: bool = False,
    ) -> AgentContext:
        """Create a child context with depth+1.

        ``capability_mode`` is the child's REQUESTED delegation privilege
        ceiling; the stored value is the narrower of it and this (parent)
        agent's own ``capability_mode`` — a child can only restrict, so
        an unset request (default ``"all"``) simply inherits the parent's mode.

        ``workspace_mode`` is the SAME story on the filesystem-containment
        axis: the child's requested tier is narrowed min-wins against this
        agent's own, so an unset request (default ``"full_access"``) inherits the
        parent's tier and a grandchild under a ``read_only`` ancestor stays
        ``read_only``.

        ``atomic`` is this child's OWN requested firebreak (from
        its ``DelegationContract``); the stored value is ``self.atomic or
        atomic`` — a simple OR, never narrower — so an atomic ancestor's
        descendants can never climb back to ``open_ended``.

        ``model_name`` falls back to ``self.model_name``, which the inheriting
        child then runs on. That fallback is only correct while this context's
        ``model_name`` is the parent's LIVE model, and a parent that heals down
        its fallback ladder promotes a new one mid-run — on the loop, since this
        context is frozen. So the loop re-seats the spawn seam's context (via
        ``dataclasses.replace``) whenever it escalates: without that, a parent
        that had just escaped a dead model would fan every un-overridden child
        straight back onto it, and the children would die at step 0 while the
        parent ran healthy.

        Raises:
            AgentDepthExceeded: If ``depth + 1 > max_depth``.
        """
        next_depth = self.depth + 1
        if next_depth > self.max_depth:
            raise AgentDepthExceeded(next_depth, self.max_depth)
        return AgentContext(
            agent_id=uuid.uuid4().hex[:12],
            parent_id=self.agent_id,
            depth=next_depth,
            max_depth=self.max_depth,
            model_name=model_name or self.model_name,
            fallback_models=self.fallback_models,
            capability_mode=self._narrow(
                self.capability_mode, capability_mode, self._CAPABILITY_MODE_RANK
            ),
            workspace_mode=self._narrow(
                self.workspace_mode, workspace_mode, self._WORKSPACE_MODE_RANK
            ),
            atomic=self.atomic or atomic,
            should_cancel=self.should_cancel,
            event_logger=self.event_logger,
            registry=self.registry,
            # Bidirectional message passing —
            # each agent gets its own queue for parent→child steering.
            # System→agent NL feedback channel.
            message_queue=queue.Queue(),
            interrupt_step=None,
        )

    @staticmethod
    def root(
        *,
        model_name: str,
        max_depth: int = 5,
        fallback_models: tuple[str, ...] = (),
        should_cancel: Callable[[], bool] | None = None,
        event_logger: Callable[[Event], None] | None = None,
        registry: AgentHypervisor | None = None,
        message_queue: queue.Queue[str] | None = None,
        interrupt_step: threading.Event | None = None,
        workspace_mode: str = "full_access",
        capability_mode: str = "all",
    ) -> AgentContext:
        """Create the root agent context.

        ``workspace_mode`` seeds the root of the filesystem-containment
        axis from ``agent.default_workspace_mode`` (default ``"full_access"``);
        every child narrows from here. Left at the default, the whole tree is
        unrestricted.

        ``capability_mode`` seeds the root of the delegation-privilege axis the
        same way (default ``"all"`` — no filtering). A caller resolving a
        principal's role ceiling passes a narrower tier (``read_only`` for a
        viewer) so the ROOT agent — not only its spawned children — is capped;
        every child then narrows monotonically from here via :meth:`child`. Left
        at ``"all"`` the whole tree is unrestricted.
        """
        reg = registry or AgentHypervisor()
        return AgentContext(
            agent_id=uuid.uuid4().hex[:12],
            parent_id=None,
            depth=0,
            max_depth=max_depth,
            model_name=model_name,
            fallback_models=fallback_models,
            should_cancel=should_cancel,
            event_logger=event_logger,
            registry=reg,
            workspace_mode=workspace_mode,
            capability_mode=capability_mode,
            message_queue=message_queue or queue.Queue(),
            interrupt_step=interrupt_step or threading.Event(),
        )

can_spawn: bool property

True if this agent is allowed to create children.

remaining_depth: int property

Number of spawn levels remaining below this agent.

child(*, model_name: str | None = None, capability_mode: str = 'all', workspace_mode: str = 'full_access', atomic: bool = False) -> AgentContext

Create a child context with depth+1.

capability_mode is the child's REQUESTED delegation privilege ceiling; the stored value is the narrower of it and this (parent) agent's own capability_mode — a child can only restrict, so an unset request (default "all") simply inherits the parent's mode.

workspace_mode is the SAME story on the filesystem-containment axis: the child's requested tier is narrowed min-wins against this agent's own, so an unset request (default "full_access") inherits the parent's tier and a grandchild under a read_only ancestor stays read_only.

atomic is this child's OWN requested firebreak (from its DelegationContract); the stored value is self.atomic or atomic — a simple OR, never narrower — so an atomic ancestor's descendants can never climb back to open_ended.

model_name falls back to self.model_name, which the inheriting child then runs on. That fallback is only correct while this context's model_name is the parent's LIVE model, and a parent that heals down its fallback ladder promotes a new one mid-run — on the loop, since this context is frozen. So the loop re-seats the spawn seam's context (via dataclasses.replace) whenever it escalates: without that, a parent that had just escaped a dead model would fan every un-overridden child straight back onto it, and the children would die at step 0 while the parent ran healthy.

Raises:

Type Description
AgentDepthExceeded

If depth + 1 > max_depth.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_context.py
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
def child(
    self,
    *,
    model_name: str | None = None,
    capability_mode: str = "all",
    workspace_mode: str = "full_access",
    atomic: bool = False,
) -> AgentContext:
    """Create a child context with depth+1.

    ``capability_mode`` is the child's REQUESTED delegation privilege
    ceiling; the stored value is the narrower of it and this (parent)
    agent's own ``capability_mode`` — a child can only restrict, so
    an unset request (default ``"all"``) simply inherits the parent's mode.

    ``workspace_mode`` is the SAME story on the filesystem-containment
    axis: the child's requested tier is narrowed min-wins against this
    agent's own, so an unset request (default ``"full_access"``) inherits the
    parent's tier and a grandchild under a ``read_only`` ancestor stays
    ``read_only``.

    ``atomic`` is this child's OWN requested firebreak (from
    its ``DelegationContract``); the stored value is ``self.atomic or
    atomic`` — a simple OR, never narrower — so an atomic ancestor's
    descendants can never climb back to ``open_ended``.

    ``model_name`` falls back to ``self.model_name``, which the inheriting
    child then runs on. That fallback is only correct while this context's
    ``model_name`` is the parent's LIVE model, and a parent that heals down
    its fallback ladder promotes a new one mid-run — on the loop, since this
    context is frozen. So the loop re-seats the spawn seam's context (via
    ``dataclasses.replace``) whenever it escalates: without that, a parent
    that had just escaped a dead model would fan every un-overridden child
    straight back onto it, and the children would die at step 0 while the
    parent ran healthy.

    Raises:
        AgentDepthExceeded: If ``depth + 1 > max_depth``.
    """
    next_depth = self.depth + 1
    if next_depth > self.max_depth:
        raise AgentDepthExceeded(next_depth, self.max_depth)
    return AgentContext(
        agent_id=uuid.uuid4().hex[:12],
        parent_id=self.agent_id,
        depth=next_depth,
        max_depth=self.max_depth,
        model_name=model_name or self.model_name,
        fallback_models=self.fallback_models,
        capability_mode=self._narrow(
            self.capability_mode, capability_mode, self._CAPABILITY_MODE_RANK
        ),
        workspace_mode=self._narrow(
            self.workspace_mode, workspace_mode, self._WORKSPACE_MODE_RANK
        ),
        atomic=self.atomic or atomic,
        should_cancel=self.should_cancel,
        event_logger=self.event_logger,
        registry=self.registry,
        # Bidirectional message passing —
        # each agent gets its own queue for parent→child steering.
        # System→agent NL feedback channel.
        message_queue=queue.Queue(),
        interrupt_step=None,
    )

root(*, model_name: str, max_depth: int = 5, fallback_models: tuple[str, ...] = (), should_cancel: Callable[[], bool] | None = None, event_logger: Callable[[Event], None] | None = None, registry: AgentHypervisor | None = None, message_queue: queue.Queue[str] | None = None, interrupt_step: threading.Event | None = None, workspace_mode: str = 'full_access', capability_mode: str = 'all') -> AgentContext staticmethod

Create the root agent context.

workspace_mode seeds the root of the filesystem-containment axis from agent.default_workspace_mode (default "full_access"); every child narrows from here. Left at the default, the whole tree is unrestricted.

capability_mode seeds the root of the delegation-privilege axis the same way (default "all" — no filtering). A caller resolving a principal's role ceiling passes a narrower tier (read_only for a viewer) so the ROOT agent — not only its spawned children — is capped; every child then narrows monotonically from here via :meth:child. Left at "all" the whole tree is unrestricted.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_context.py
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
@staticmethod
def root(
    *,
    model_name: str,
    max_depth: int = 5,
    fallback_models: tuple[str, ...] = (),
    should_cancel: Callable[[], bool] | None = None,
    event_logger: Callable[[Event], None] | None = None,
    registry: AgentHypervisor | None = None,
    message_queue: queue.Queue[str] | None = None,
    interrupt_step: threading.Event | None = None,
    workspace_mode: str = "full_access",
    capability_mode: str = "all",
) -> AgentContext:
    """Create the root agent context.

    ``workspace_mode`` seeds the root of the filesystem-containment
    axis from ``agent.default_workspace_mode`` (default ``"full_access"``);
    every child narrows from here. Left at the default, the whole tree is
    unrestricted.

    ``capability_mode`` seeds the root of the delegation-privilege axis the
    same way (default ``"all"`` — no filtering). A caller resolving a
    principal's role ceiling passes a narrower tier (``read_only`` for a
    viewer) so the ROOT agent — not only its spawned children — is capped;
    every child then narrows monotonically from here via :meth:`child`. Left
    at ``"all"`` the whole tree is unrestricted.
    """
    reg = registry or AgentHypervisor()
    return AgentContext(
        agent_id=uuid.uuid4().hex[:12],
        parent_id=None,
        depth=0,
        max_depth=max_depth,
        model_name=model_name,
        fallback_models=fallback_models,
        should_cancel=should_cancel,
        event_logger=event_logger,
        registry=reg,
        workspace_mode=workspace_mode,
        capability_mode=capability_mode,
        message_queue=message_queue or queue.Queue(),
        interrupt_step=interrupt_step or threading.Event(),
    )

AgentDepthExceeded

Bases: Exception

Raised when attempting to spawn beyond max_depth.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_context.py
27
28
29
30
31
32
33
34
class AgentDepthExceeded(Exception):
    """Raised when attempting to spawn beyond max_depth."""

    def __init__(self, attempted: int, maximum: int) -> None:
        """Initialize with the attempted and maximum depth values."""
        super().__init__(f"Agent depth {attempted} exceeds maximum {maximum}")
        self.attempted = attempted
        self.maximum = maximum

__init__(attempted: int, maximum: int) -> None

Initialize with the attempted and maximum depth values.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_context.py
30
31
32
33
34
def __init__(self, attempted: int, maximum: int) -> None:
    """Initialize with the attempted and maximum depth values."""
    super().__init__(f"Agent depth {attempted} exceeds maximum {maximum}")
    self.attempted = attempted
    self.maximum = maximum

AgentHandle dataclass

Mutable runtime state for a single agent — lives in the hypervisor.

Created when an agent registers and updated throughout its lifecycle. Fields are read by the CLI agent tree display and the hypervisor's query/cancellation methods.

Handle starts as submitted, transitions to running when the loop begins, then to a terminal state. last_step_at enables tool-call-granularity stall detection without destroying accumulated context. message_queue enables bidirectional adaptive coordination — the hypervisor injects NL feedback.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
@dataclass(slots=True)
class AgentHandle:
    """Mutable runtime state for a single agent — lives in the hypervisor.

    Created when an agent registers and updated throughout its lifecycle.
    Fields are read by the CLI agent tree display and the hypervisor's
    query/cancellation methods.

    Handle starts as ``submitted``, transitions to
    ``running`` when the loop begins, then to a terminal state.
    ``last_step_at`` enables tool-call-granularity
    stall detection without destroying accumulated context.
    ``message_queue`` enables bidirectional
    adaptive coordination — the hypervisor injects NL feedback.
    """

    agent_id: str
    parent_id: str | None
    depth: int
    model_name: str
    task_description: str
    # The registered AgentDef name this agent was spawned as (e.g.
    # ``scg-path-probe``) — ``None`` for an ad-hoc spawn with no ``agent_type``.
    # Distinct from ``model_name``: the trace projection needs the LANE identity
    # (the def), which a model name can never carry.
    agent_type: str | None = None
    status: AgentStatus = "submitted"  # Start as submitted
    started_at: float = field(default_factory=time.monotonic)
    stopped_at: float | None = None
    steps_completed: int = 0
    last_tool_id: str | None = None
    # Tool-call-granularity timing for stall detection
    last_step_at: float | None = None
    # The tool actually IN FLIGHT right now — stamped at dispatch start and
    # cleared back to ``None`` when the call resolves. Distinct from
    # ``last_tool_id`` (only updated on COMPLETION): during a long-running
    # call, ``last_tool_id`` still names the PREVIOUS finished tool, so the
    # watchdog must read ``active_tool_id`` for stall attribution, never
    # ``last_tool_id``.
    active_tool_id: str | None = None
    error: str | AgentError | None = None
    asyncio_task: asyncio.Task[object] | None = None
    # Bidirectional message passing
    message_queue: queue.Queue[str] | None = None
    # Completed CU stored on handle for async retrieval
    result: AgentResult | None = None
    # Auto-updated progress for monitoring
    progress_note: str | None = None
    # Context compaction tracking — visible in agent tree rendering.
    compaction_count: int = 0
    last_compacted_at: float | None = None
    # Token usage — accumulated from LLM response.usage_metadata per call.
    # Written only by the owning ToolUseLoop coroutine; read by CLI/API.
    input_tokens: int = 0
    output_tokens: int = 0
    # Signaled when agent reaches a terminal state (completed/failed/cancelled).
    done_event: asyncio.Event = field(default_factory=asyncio.Event)
    # Bumped by the spawn bridge's bounded-retry driver each
    # time this same task is re-admitted (1 = first/only attempt). Read by the
    # agent-tree render + check_agents payload so a retried child is visible.
    attempts: int = 1
    # The spawner's own delegation bounds for this child.
    # Additive: the default (disabled) contract is a no-op, so an ad-hoc
    # spawn with"no ``contract`` behaves exactly as before.
    contract: DelegationContract = field(default_factory=DelegationContract)
    # This child's OWN spawn attestation record hash
    # (stamped by ``SpawnAgentTool`` right after registration), threaded back
    # in as the terminal record's ``spawn_hash`` link. "" when no chain is
    # wired for this session.
    attestation_spawn_hash: str = ""
    # Latched when this agent's ONE terminal ``stop`` lifecycle event is written.
    # Several paths can settle the same agent (its own success/cancel/failure
    # handler, and its parent's teardown cascade), and they RACE — the flag is
    # what keeps "exactly one stop per start" true rather than merely likely.
    # Written only through the spawn bridge's terminal emitter.
    terminal_emitted: bool = False

AgentHypervisor

Hypervisor control plane — manages the full agent tree for a session.

Thread-safe via asyncio.Lock. A single instance is shared across all agents in the hierarchy through AgentContext.registry.

Responsibilities
  • Admission control: accept() / accept_batch() / release() gate concurrency through an injected :class:AgentQueue (default 20 running slots) that QUEUES over-cap work instead of dropping it.
  • Registration: register() / unregister() track agent handles keyed by agent_id.
  • Status: update_step() / mark_done() record execution progress and terminal state.
  • Queries: list_children() / list_descendants() / list_all() expose the live tree for display and introspection.
  • Cancellation: cancel_agent() cancels a single agent; cleanup() tears down the entire tree on session exit.
Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
class AgentHypervisor:
    """Hypervisor control plane — manages the full agent tree for a session.

    Thread-safe via ``asyncio.Lock``. A single instance is shared across all
    agents in the hierarchy through ``AgentContext.registry``.

    Responsibilities:
        - **Admission control**: ``accept()`` / ``accept_batch()`` /
          ``release()`` gate concurrency through an injected
          :class:`AgentQueue` (default 20 running slots) that QUEUES over-cap
          work instead of dropping it.
        - **Registration**: ``register()`` / ``unregister()`` track agent
          handles keyed by ``agent_id``.
        - **Status**: ``update_step()`` / ``mark_done()`` record execution
          progress and terminal state.
        - **Queries**: ``list_children()`` / ``list_descendants()`` /
          ``list_all()`` expose the live tree for display and introspection.
        - **Cancellation**: ``cancel_agent()`` cancels a single agent;
          ``cleanup()`` tears down the entire tree on session exit.
    """

    def __init__(
        self,
        *,
        max_concurrent: int = 20,
        session_step_budget: int = 0,
        attestation: AttestationChain | None = None,
        queue: AgentQueue | None = None,
    ) -> None:
        """Initialize hypervisor with concurrency and budget limits.

        Session-wide budget with graduated enforcement.

        Args:
            max_concurrent: Maximum number of concurrently RUNNING agents.
            session_step_budget: Total tool steps allowed across all agents in the
                session. 0 means unlimited.
            attestation: optional provenance hash chain for
                this session's agent tree. The hypervisor never constructs one
                itself (it has no session_id, no store) — the orchestrator
                injects it, seeded from the store's last persisted record, when
                the feature is config-enabled. ``None`` (the default) means no
                attestation is recorded; ``SpawnAgentTool`` reads this
                attribute directly and no-ops when it's absent.
            queue: the admission scheduler. Injected so a deployment (or a
                test) can seat a differently-sized or differently-ordered one
                without the hypervisor learning how scheduling works; the
                default is an :class:`AgentQueue` of ``max_concurrent`` slots.
        """
        self._agents: dict[str, AgentHandle] = {}
        self._lock: asyncio.Lock = asyncio.Lock()
        self._queue: AgentQueue = queue or AgentQueue(capacity=max_concurrent)
        # Set unconditionally so an INJECTED queue is wired too: the queue can
        # detect a failed dispatch but only the hypervisor can settle the
        # handle it stranded.
        self._queue.on_dispatch_failed = self._settle_failed_dispatch
        # Session-wide resource tracking
        self._total_steps: int = 0
        self._session_step_budget: int = session_step_budget
        self.attestation = attestation

    # ------------------------------------------------------------------
    # Admission control
    # ------------------------------------------------------------------

    async def accept(self, unit: SpawnUnit) -> bool:
        """Admit ONE spawn: ``True`` = dispatched now, ``False`` = deferred.

        Never refuses for capacity — that is the whole point of the queue.
        """
        return await self._queue.accept(unit)

    async def accept_batch(self, units: Sequence[SpawnUnit]) -> list[bool]:
        """Admit a whole fan-out at once — see :meth:`AgentQueue.accept_batch`."""
        return await self._queue.accept_batch(units)

    async def release(self) -> None:
        """Settle one running agent's slot, dispatching whatever waits on it.

        Async because it is the DISPATCH PUMP, not a bare counter decrement:
        the freed slot is handed to the next waiting unit, which means starting
        it. This is deliberately the completion path every settling agent
        already calls, and exactly-once is already guaranteed there — a pump
        hung off ``on_agent_stop`` instead would fire while the slot is still
        held and find the queue full every time.
        """
        await self._queue.release()

    async def cancel_pending(self) -> list[str]:
        """Drop every accepted-but-undispatched spawn, settling each as cancelled.

        A waiting agent holds no slot and owns no ``asyncio.Task``, so neither
        ``cancel_agent`` nor the cleanup sweep's task cancellation can reach
        it — it would sit ``submitted`` forever and read as live. Returns the
        ids dropped. Called before teardown so a manager settling on the way
        out cannot pump a brand-new child into a run that is already ending.
        """
        dropped = self._queue.clear()
        if not dropped:
            return []
        now = time.monotonic()
        async with self._lock:
            for spawn in dropped:
                handle = self._agents.get(spawn.agent_id)
                if handle is not None and handle.status in ACTIVE_STATUSES:
                    handle.status = "cancelled"
                    handle.error = "cancelled while waiting for a concurrency slot"
                    handle.stopped_at = now
                    handle.done_event.set()
        return [spawn.agent_id for spawn in dropped]

    async def _settle_failed_dispatch(
        self, spawn: ScheduledSpawn, exc: BaseException
    ) -> None:
        """Settle an agent whose launcher raised — it will never start.

        Its handle was registered at acceptance and its spawn path already
        returned, so nothing downstream owns it: it would sit non-terminal
        forever, keep answering ``collect_running``, and hold the
        promise-as-completion gate open for the rest of the session. The
        message is the exception's TYPE plus a bounded head — a launcher
        failure is infrastructure, and an unbounded provider string has no
        business on a handle read by every client.
        """
        detail = f"{type(exc).__name__}: {exc}"[:200]
        await self.mark_done(
            spawn.agent_id, "failed", error=f"sub-agent never started — {detail}"
        )

    @property
    def free_slots(self) -> int:
        """Slots a spawn could be dispatched into right now."""
        return self._queue.free_slots

    @property
    def pending_dispatch(self) -> int:
        """Accepted spawns not yet started.

        NOT a status: those agents are ``submitted`` like any other, and this
        is only the scheduler's count of how many are still waiting on a slot.
        """
        return self._queue.waiting

    # ------------------------------------------------------------------
    # Registration
    # ------------------------------------------------------------------

    async def register(self, handle: AgentHandle) -> None:
        """Register a new agent in the hypervisor."""
        async with self._lock:
            self._agents[handle.agent_id] = handle

    async def unregister(self, agent_id: str) -> AgentHandle | None:
        """Remove an agent from the hypervisor."""
        async with self._lock:
            handle = self._agents.pop(agent_id, None)
            if handle and handle.status == "running":
                handle.status = "completed"
                handle.stopped_at = time.monotonic()
            return handle

    # ------------------------------------------------------------------
    # Status updates
    # ------------------------------------------------------------------

    async def update_step(self, agent_id: str, tool_id: str) -> None:
        """Record a completed tool execution step.

        Track at tool-call granularity, not agent
        granularity. Updates last_step_at for stall detection and total_steps
        for session budget enforcement.
        """
        async with self._lock:
            handle = self._agents.get(agent_id)
            if handle:
                handle.steps_completed += 1
                handle.last_tool_id = tool_id
                handle.last_step_at = time.monotonic()
                self._total_steps += 1

    async def mark_tool_start(self, agent_id: str, tool_id: str | None) -> None:
        """Stamp (or clear) the tool actually in flight for stall attribution.

        ``update_step`` only stamps ``last_tool_id`` on
        COMPLETION, so a watchdog check firing mid-call would misattribute the
        stall to the previous, already-finished tool. Call this at dispatch
        start with the tool name, and again with ``None`` once the call
        resolves (success, timeout, or exception) so ``active_tool_id`` never
        lingers stale once the agent moves on.
        """
        async with self._lock:
            handle = self._agents.get(agent_id)
            if handle:
                handle.active_tool_id = tool_id

    async def mark_done(
        self,
        agent_id: str,
        status: AgentStatus,
        error: str | AgentError | None = None,
    ) -> None:
        """Mark an agent as done with a terminal status."""
        async with self._lock:
            handle = self._agents.get(agent_id)
            if handle:
                handle.status = status
                handle.error = error
                handle.stopped_at = time.monotonic()
                handle.done_event.set()

    # ------------------------------------------------------------------
    # Queries
    # ------------------------------------------------------------------

    async def get(self, agent_id: str) -> AgentHandle | None:
        """Return a single agent handle, or None."""
        async with self._lock:
            return self._agents.get(agent_id)

    async def list_children(self, parent_id: str) -> list[AgentHandle]:
        """List direct children of a parent. Enforces isolation."""
        async with self._lock:
            return [h for h in self._agents.values() if h.parent_id == parent_id]

    async def list_descendants(self, ancestor_id: str) -> list[AgentHandle]:
        """List all descendants recursively."""
        async with self._lock:
            result: list[AgentHandle] = []
            queue = [ancestor_id]
            while queue:
                pid = queue.pop()
                for h in self._agents.values():
                    if h.parent_id == pid:
                        result.append(h)
                        queue.append(h.agent_id)
            return result

    async def list_all(self) -> list[AgentHandle]:
        """Full tree view — for CLI/API user visibility."""
        async with self._lock:
            return list(self._agents.values())

    async def list_visible(self, exclude_agent_id: str | None = None) -> list[AgentHandle]:
        """Snapshot of agents excluding the caller — used by check_agents."""
        async with self._lock:
            return [h for h in self._agents.values() if h.agent_id != exclude_agent_id]

    # ------------------------------------------------------------------
    # Budget & monitoring
    # ------------------------------------------------------------------

    @property
    def total_steps(self) -> int:
        """Total tool steps executed across all agents in the session."""
        return self._total_steps

    def budget_exhausted(self) -> bool:
        """Check if the session step budget is exhausted (0 = unlimited).

        Graduated enforcement — this is the trigger
        check. The response (NL warning injection) happens in ToolUseLoop.
        """
        return self._session_step_budget > 0 and self._total_steps >= self._session_step_budget

    def budget_remaining(self) -> int:
        """Steps remaining in the session budget (0 = unlimited)."""
        if self._session_step_budget <= 0:
            return -1  # Unlimited
        return max(0, self._session_step_budget - self._total_steps)

    def budget_warning(self, *, headroom: int = 5) -> bool:
        """True when within ``headroom`` steps of the session budget (0 = off)."""
        if self._session_step_budget <= 0:
            return False
        return self.budget_remaining() <= headroom

    async def stalled_agents(self, threshold: float = 120.0) -> list[AgentHandle]:
        """Return agents not making progress within threshold seconds.

        Internal trigger: delegatee
        unresponsive → diagnose → evaluate → intervene.
        """
        now = time.monotonic()
        async with self._lock:
            return [
                h
                for h in self._agents.values()
                if h.status == "running"
                and h.last_step_at is not None
                and (now - h.last_step_at) > threshold
            ]

    async def agent_step_state(self, agent_id: str) -> str:
        """Per-contract step-budget state for one agent.

        ``"ok"`` for an unregistered agent or one whose contract carries no
        step bound — the same inert default as a disabled contract.
        """
        async with self._lock:
            handle = self._agents.get(agent_id)
        if handle is None:
            return "ok"
        return handle.contract.step_state(handle.steps_completed)

    async def agent_token_state(self, agent_id: str) -> str:
        """Per-contract advisory token state for one agent."""
        async with self._lock:
            handle = self._agents.get(agent_id)
        if handle is None:
            return "ok"
        return handle.contract.token_state(handle.input_tokens + handle.output_tokens)

    async def over_wall_deadline_agents(
        self, now: float | None = None
    ) -> list[tuple[AgentHandle, str]]:
        """Running agents whose contract wall-clock deadline is warn/over.

        Mirrors :meth:`stalled_agents` — a per-contract cousin of the stall
        sweep, keyed on ``started_at`` rather than last-progress. ``now``
        arrives as a clock ARG (never read internally), so a caller can drive
        this deterministically without sleeping a real clock.
        """
        clock = now if now is not None else time.monotonic()
        async with self._lock:
            result: list[tuple[AgentHandle, str]] = []
            for h in self._agents.values():
                if h.status != "running" or h.contract.max_wall_s <= 0:
                    continue
                state = h.contract.wall_state(clock - h.started_at)
                if state in ("warn", "over"):
                    result.append((h, state))
            return result

    @staticmethod
    def _contract_marker(handle: AgentHandle) -> str:
        """Compact ``| contract: ...`` suffix for an agent-tree row.

        Empty (no-op) unless the handle's contract is ``enabled``, so a
        contract-less child's row carries no suffix.
        """
        if not handle.contract.enabled:
            return ""
        bits: list[str] = [handle.contract.autonomy]
        if handle.contract.max_steps > 0:
            bits.append(f"{handle.steps_completed}/{handle.contract.max_steps} steps")
        if handle.contract.max_wall_s > 0:
            elapsed = time.monotonic() - handle.started_at
            bits.append(f"{elapsed:.0f}/{handle.contract.max_wall_s:.0f}s")
        return f" | contract: {', '.join(bits)}"

    # ------------------------------------------------------------------
    # Bidirectional messaging
    # ------------------------------------------------------------------

    async def send_message(self, agent_id: str, message: str) -> str | None:
        """Send a steering message to a running agent.

        Returns ``None`` on success, or a diagnostic string on failure.

        Adaptive coordination — the hypervisor
        injects NL feedback (budget warnings, stall nudges) into the agent's
        message queue. The ToolUseLoop drains this queue between steps.
        Bidirectional system↔agent feedback loop.
        """
        async with self._lock:
            handle = self._agents.get(agent_id)
            if not handle:
                return "agent not in registry"
            if handle.status != "running":
                return f"agent status is '{handle.status}'"
            if not handle.message_queue:
                return "no message queue"
            handle.message_queue.put_nowait(message)
            return None

    async def record_compaction(self, agent_id: str) -> None:
        """Record that an agent compacted its context."""
        async with self._lock:
            handle = self._agents.get(agent_id)
            if handle:
                handle.compaction_count += 1
                handle.last_compacted_at = time.monotonic()

    # ------------------------------------------------------------------
    # Global eye
    # ------------------------------------------------------------------

    async def render_agent_tree(
        self,
        *,
        exclude_agent_id: str | None = None,
    ) -> str:
        """Render a concise text summary of the agent tree for the root's system prompt.

        Structural transparency — the root
        agent (the hypervisor's "brain") gets a live view of all agents so
        it can reason about the delegation state and intervene if needed.

        Args:
            exclude_agent_id: If provided, omit this agent from the rendered
                tree.  Used so the calling agent does not see itself listed.
        """
        async with self._lock:
            if not self._agents:
                return ""
            visible = [h for h in self._agents.values() if h.agent_id != exclude_agent_id]
            if not visible:
                return ""
            registry = get_prompt_registry()
            counts: dict[str, int] = {}
            for h in visible:
                counts[h.status] = counts.get(h.status, 0) + 1
            status_parts = [f"{v} {k}" for k, v in sorted(counts.items())]
            budget_str = ""
            if self._session_step_budget > 0:
                budget_str = registry.render(
                    "catalog.agent_tree.budget",
                    total_steps=self._total_steps,
                    session_step_budget=self._session_step_budget,
                )
            header = registry.render(
                "catalog.agent_tree.header",
                status_parts=", ".join(status_parts),
                budget_str=budget_str,
            )
            lines = [header]
            for h in sorted(visible, key=lambda x: (x.depth, x.agent_id)):
                indent = "  " * h.depth
                step_info = f"{h.steps_completed} steps"
                if h.last_tool_id:
                    step_info += registry.render(
                        "catalog.agent_tree.step_info_last", last_tool_id=h.last_tool_id
                    )
                # Surface bounded-retry provenance inline so the
                # root sees a child was re-delegated (omitted when never
                # retried, so an ordinary line carries no attempt count).
                if h.attempts > 1:
                    step_info += f", {h.attempts} attempts"
                status_marker = ""
                if h.status == "completed":
                    status_marker = " -> success"
                elif h.status == "failed":
                    status_marker = " -> FAILED"
                elif h.status == "cancelled":
                    status_marker = " -> cancelled"
                # Progress/result in tree view
                extra = ""
                if h.result and h.result.summary:
                    extra = registry.render(
                        "catalog.agent_tree.result",
                        status=h.result.status,
                        summary=h.result.summary[:120],
                    )
                elif h.progress_note:
                    extra = registry.render(
                        "catalog.agent_tree.progress",
                        progress_note=h.progress_note[:120],
                    )
                compact_marker = (
                    registry.render(
                        "catalog.agent_tree.compact", compaction_count=h.compaction_count
                    )
                    if h.compaction_count
                    else ""
                )
                task_preview = h.task_description[:80]
                line = registry.render(
                    "catalog.agent_tree.line",
                    indent=indent,
                    agent_id_head=h.agent_id[:8],
                    status=h.status,
                    task_preview=task_preview,
                    step_info=step_info,
                    status_marker=status_marker,
                    compact_marker=compact_marker,
                    extra=extra,
                )
                # Appended OUTSIDE the templated line (never
                # a catalog.yaml edit) so a contract-less row's bytes are
                # completely unaffected.
                lines.append(line + self._contract_marker(h))
            return "\n".join(lines)

    # ------------------------------------------------------------------
    # Async delegation queries
    # ------------------------------------------------------------------

    async def collect_completed(self, parent_id: str) -> list[AgentHandle]:
        """Return children that reached a terminal state with a stored result.

        Async CU retrieval — parent reads results
        when ready, not when child finishes.
        """
        terminal = {"completed", "failed", "cancelled"}
        async with self._lock:
            return [
                h
                for h in self._agents.values()
                if h.parent_id == parent_id and h.status in terminal and h.result is not None
            ]

    async def collect_running(self, parent_id: str) -> list[AgentHandle]:
        """Return children of *parent_id* that are still active.

        Reads :data:`ACTIVE_STATUSES` rather than re-listing the states. That
        matters beyond tidiness: this method is the promise-as-completion
        gate's ownership index AND what ``check_agents`` reports, so a
        capacity-deferred child — ``submitted``, registered, not yet started —
        must appear here. A refused unit was registered nowhere, which is
        exactly why a parent that over-subscribed the fleet was told all its
        work was done.
        """
        async with self._lock:
            return [
                h
                for h in self._agents.values()
                if h.parent_id == parent_id and h.status in ACTIVE_STATUSES
            ]

    async def send_to_parent(self, child_agent_id: str, message: str) -> str | None:
        """Route a message from a child agent to its parent's queue.

        Returns ``None`` on success, or a diagnostic string on failure.

        Bidirectional message passing —
        enables lifecycle manager to notify parent on child completion.
        """
        async with self._lock:
            child = self._agents.get(child_agent_id)
            if not child or not child.parent_id:
                return "child not in registry or no parent"
            parent = self._agents.get(child.parent_id)
            if not parent:
                return "parent not in registry"
            if not parent.message_queue:
                return "parent has no message queue"
            parent.message_queue.put_nowait(message)
            return None

    # ------------------------------------------------------------------
    # Cancellation
    # ------------------------------------------------------------------

    async def cancel_agent(self, agent_id: str) -> str | None:
        """Cancel an agent, whether it has started or is still waiting for a slot.

        Returns ``None`` on success, or a diagnostic string on failure.

        **The waiting case is checked FIRST and it is not an optimization.**
        ``asyncio_task`` is populated by the child driver, which runs only
        AFTER dispatch — so a capacity-deferred agent has no task, the
        task-cancellation path reported ``no asyncio task``, and the agent then
        STARTED anyway the moment a sibling freed a slot. Refusing to cancel
        something and then running it is the worst of both answers, and it is
        reachable the instant a fan-out exceeds the concurrency limit.
        Dropping it from the scheduler is what makes the cancel real.
        """
        async with self._lock:
            handle = self._agents.get(agent_id)
            if not handle:
                return "agent not in registry"
            if self._queue.discard(agent_id) is not None:
                handle.status = "cancelled"
                handle.error = "cancelled before it was dispatched"
                handle.stopped_at = time.monotonic()
                handle.done_event.set()
                return None
            if not handle.asyncio_task:
                return "no asyncio task"
            if handle.asyncio_task.done():
                return f"task already done (status: {handle.status})"
            handle.asyncio_task.cancel()
            handle.status = "cancelled"
            handle.stopped_at = time.monotonic()
            handle.done_event.set()
            return None

    # ------------------------------------------------------------------
    # Cleanup
    # ------------------------------------------------------------------

    async def cleanup(self, timeout: float = 5.0) -> None:
        """Cancel all agents with graceful escalation.

        Requests cancellation on all active asyncio tasks, waits up to
        *timeout* seconds for them to finish, then force-marks any still
        pending as cancelled.

        Covers every :data:`ACTIVE_STATUSES` agent, not just ``running``: one
        still at ``submitted`` never had a loop task to cancel — and one still
        waiting for a slot never had a task at all — which is exactly why
        either would otherwise survive shutdown unmarked. Waiting units are
        dropped from the scheduler FIRST, so a manager settling during the
        wait window cannot pump a fresh child into a hypervisor that is being
        torn down.

        Marking here is deliberately IN-MEMORY only — the hypervisor holds no
        event logger, and wiring one in would invert the layering (the session
        owns the transcript, the hypervisor owns the tree). The durable terminal
        is written by whichever peer can still reach the log:

        - While the task is still alive (during the request or the wait):
          cancelling it raises ``CancelledError`` inside the agent's own
          lifecycle handler, which writes the terminal ``stop`` before
          unwinding. That is the normal path.
        - Once the wait times out, at the final force-mark sweep (the task is
          wedged, or the process is being torn down mid-flight): nothing
          in-process can still write, so the durable terminal comes from the
          startup sweep on the NEXT boot, which settles agent spans left open
          under a run that has since ended.
        """
        await self.cancel_pending()

        async with self._lock:
            active = [h for h in self._agents.values() if h.status in ACTIVE_STATUSES]

        if not active:
            async with self._lock:
                self._agents.clear()
            return

        # Request cancellation.
        for handle in active:
            if handle.asyncio_task and not handle.asyncio_task.done():
                handle.asyncio_task.cancel()

        # Wait with timeout.
        tasks = [h.asyncio_task for h in active if h.asyncio_task and not h.asyncio_task.done()]
        if tasks:
            done, pending = await asyncio.wait(
                tasks,
                timeout=timeout,
                return_when=asyncio.ALL_COMPLETED,
            )

            # Force-mark any still-pending as cancelled.
            for task in pending:
                for handle in active:
                    if handle.asyncio_task is task:
                        async with self._lock:
                            if handle.status in ACTIVE_STATUSES:
                                handle.status = "cancelled"
                                handle.error = "Force-cancelled after timeout"
                                handle.stopped_at = time.monotonic()

        # Final sweep: mark any remaining non-terminal agents and clear.
        async with self._lock:
            for h in list(self._agents.values()):
                if h.status in ACTIVE_STATUSES:
                    h.status = "cancelled"
                    h.stopped_at = time.monotonic()
            self._agents.clear()

free_slots: int property

Slots a spawn could be dispatched into right now.

pending_dispatch: int property

Accepted spawns not yet started.

NOT a status: those agents are submitted like any other, and this is only the scheduler's count of how many are still waiting on a slot.

total_steps: int property

Total tool steps executed across all agents in the session.

__init__(*, max_concurrent: int = 20, session_step_budget: int = 0, attestation: AttestationChain | None = None, queue: AgentQueue | None = None) -> None

Initialize hypervisor with concurrency and budget limits.

Session-wide budget with graduated enforcement.

Parameters:

Name Type Description Default
max_concurrent int

Maximum number of concurrently RUNNING agents.

20
session_step_budget int

Total tool steps allowed across all agents in the session. 0 means unlimited.

0
attestation AttestationChain | None

optional provenance hash chain for this session's agent tree. The hypervisor never constructs one itself (it has no session_id, no store) — the orchestrator injects it, seeded from the store's last persisted record, when the feature is config-enabled. None (the default) means no attestation is recorded; SpawnAgentTool reads this attribute directly and no-ops when it's absent.

None
queue AgentQueue | None

the admission scheduler. Injected so a deployment (or a test) can seat a differently-sized or differently-ordered one without the hypervisor learning how scheduling works; the default is an :class:AgentQueue of max_concurrent slots.

None
Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
def __init__(
    self,
    *,
    max_concurrent: int = 20,
    session_step_budget: int = 0,
    attestation: AttestationChain | None = None,
    queue: AgentQueue | None = None,
) -> None:
    """Initialize hypervisor with concurrency and budget limits.

    Session-wide budget with graduated enforcement.

    Args:
        max_concurrent: Maximum number of concurrently RUNNING agents.
        session_step_budget: Total tool steps allowed across all agents in the
            session. 0 means unlimited.
        attestation: optional provenance hash chain for
            this session's agent tree. The hypervisor never constructs one
            itself (it has no session_id, no store) — the orchestrator
            injects it, seeded from the store's last persisted record, when
            the feature is config-enabled. ``None`` (the default) means no
            attestation is recorded; ``SpawnAgentTool`` reads this
            attribute directly and no-ops when it's absent.
        queue: the admission scheduler. Injected so a deployment (or a
            test) can seat a differently-sized or differently-ordered one
            without the hypervisor learning how scheduling works; the
            default is an :class:`AgentQueue` of ``max_concurrent`` slots.
    """
    self._agents: dict[str, AgentHandle] = {}
    self._lock: asyncio.Lock = asyncio.Lock()
    self._queue: AgentQueue = queue or AgentQueue(capacity=max_concurrent)
    # Set unconditionally so an INJECTED queue is wired too: the queue can
    # detect a failed dispatch but only the hypervisor can settle the
    # handle it stranded.
    self._queue.on_dispatch_failed = self._settle_failed_dispatch
    # Session-wide resource tracking
    self._total_steps: int = 0
    self._session_step_budget: int = session_step_budget
    self.attestation = attestation

accept(unit: SpawnUnit) -> bool async

Admit ONE spawn: True = dispatched now, False = deferred.

Never refuses for capacity — that is the whole point of the queue.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
721
722
723
724
725
726
async def accept(self, unit: SpawnUnit) -> bool:
    """Admit ONE spawn: ``True`` = dispatched now, ``False`` = deferred.

    Never refuses for capacity — that is the whole point of the queue.
    """
    return await self._queue.accept(unit)

accept_batch(units: Sequence[SpawnUnit]) -> list[bool] async

Admit a whole fan-out at once — see :meth:AgentQueue.accept_batch.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
728
729
730
async def accept_batch(self, units: Sequence[SpawnUnit]) -> list[bool]:
    """Admit a whole fan-out at once — see :meth:`AgentQueue.accept_batch`."""
    return await self._queue.accept_batch(units)

agent_step_state(agent_id: str) -> str async

Per-contract step-budget state for one agent.

"ok" for an unregistered agent or one whose contract carries no step bound — the same inert default as a disabled contract.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
948
949
950
951
952
953
954
955
956
957
958
async def agent_step_state(self, agent_id: str) -> str:
    """Per-contract step-budget state for one agent.

    ``"ok"`` for an unregistered agent or one whose contract carries no
    step bound — the same inert default as a disabled contract.
    """
    async with self._lock:
        handle = self._agents.get(agent_id)
    if handle is None:
        return "ok"
    return handle.contract.step_state(handle.steps_completed)

agent_token_state(agent_id: str) -> str async

Per-contract advisory token state for one agent.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
960
961
962
963
964
965
966
async def agent_token_state(self, agent_id: str) -> str:
    """Per-contract advisory token state for one agent."""
    async with self._lock:
        handle = self._agents.get(agent_id)
    if handle is None:
        return "ok"
    return handle.contract.token_state(handle.input_tokens + handle.output_tokens)

budget_exhausted() -> bool

Check if the session step budget is exhausted (0 = unlimited).

Graduated enforcement — this is the trigger check. The response (NL warning injection) happens in ToolUseLoop.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
912
913
914
915
916
917
918
def budget_exhausted(self) -> bool:
    """Check if the session step budget is exhausted (0 = unlimited).

    Graduated enforcement — this is the trigger
    check. The response (NL warning injection) happens in ToolUseLoop.
    """
    return self._session_step_budget > 0 and self._total_steps >= self._session_step_budget

budget_remaining() -> int

Steps remaining in the session budget (0 = unlimited).

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
920
921
922
923
924
def budget_remaining(self) -> int:
    """Steps remaining in the session budget (0 = unlimited)."""
    if self._session_step_budget <= 0:
        return -1  # Unlimited
    return max(0, self._session_step_budget - self._total_steps)

budget_warning(*, headroom: int = 5) -> bool

True when within headroom steps of the session budget (0 = off).

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
926
927
928
929
930
def budget_warning(self, *, headroom: int = 5) -> bool:
    """True when within ``headroom`` steps of the session budget (0 = off)."""
    if self._session_step_budget <= 0:
        return False
    return self.budget_remaining() <= headroom

cancel_agent(agent_id: str) -> str | None async

Cancel an agent, whether it has started or is still waiting for a slot.

Returns None on success, or a diagnostic string on failure.

The waiting case is checked FIRST and it is not an optimization. asyncio_task is populated by the child driver, which runs only AFTER dispatch — so a capacity-deferred agent has no task, the task-cancellation path reported no asyncio task, and the agent then STARTED anyway the moment a sibling freed a slot. Refusing to cancel something and then running it is the worst of both answers, and it is reachable the instant a fan-out exceeds the concurrency limit. Dropping it from the scheduler is what makes the cancel real.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
async def cancel_agent(self, agent_id: str) -> str | None:
    """Cancel an agent, whether it has started or is still waiting for a slot.

    Returns ``None`` on success, or a diagnostic string on failure.

    **The waiting case is checked FIRST and it is not an optimization.**
    ``asyncio_task`` is populated by the child driver, which runs only
    AFTER dispatch — so a capacity-deferred agent has no task, the
    task-cancellation path reported ``no asyncio task``, and the agent then
    STARTED anyway the moment a sibling freed a slot. Refusing to cancel
    something and then running it is the worst of both answers, and it is
    reachable the instant a fan-out exceeds the concurrency limit.
    Dropping it from the scheduler is what makes the cancel real.
    """
    async with self._lock:
        handle = self._agents.get(agent_id)
        if not handle:
            return "agent not in registry"
        if self._queue.discard(agent_id) is not None:
            handle.status = "cancelled"
            handle.error = "cancelled before it was dispatched"
            handle.stopped_at = time.monotonic()
            handle.done_event.set()
            return None
        if not handle.asyncio_task:
            return "no asyncio task"
        if handle.asyncio_task.done():
            return f"task already done (status: {handle.status})"
        handle.asyncio_task.cancel()
        handle.status = "cancelled"
        handle.stopped_at = time.monotonic()
        handle.done_event.set()
        return None

cancel_pending() -> list[str] async

Drop every accepted-but-undispatched spawn, settling each as cancelled.

A waiting agent holds no slot and owns no asyncio.Task, so neither cancel_agent nor the cleanup sweep's task cancellation can reach it — it would sit submitted forever and read as live. Returns the ids dropped. Called before teardown so a manager settling on the way out cannot pump a brand-new child into a run that is already ending.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
async def cancel_pending(self) -> list[str]:
    """Drop every accepted-but-undispatched spawn, settling each as cancelled.

    A waiting agent holds no slot and owns no ``asyncio.Task``, so neither
    ``cancel_agent`` nor the cleanup sweep's task cancellation can reach
    it — it would sit ``submitted`` forever and read as live. Returns the
    ids dropped. Called before teardown so a manager settling on the way
    out cannot pump a brand-new child into a run that is already ending.
    """
    dropped = self._queue.clear()
    if not dropped:
        return []
    now = time.monotonic()
    async with self._lock:
        for spawn in dropped:
            handle = self._agents.get(spawn.agent_id)
            if handle is not None and handle.status in ACTIVE_STATUSES:
                handle.status = "cancelled"
                handle.error = "cancelled while waiting for a concurrency slot"
                handle.stopped_at = now
                handle.done_event.set()
    return [spawn.agent_id for spawn in dropped]

cleanup(timeout: float = 5.0) -> None async

Cancel all agents with graceful escalation.

Requests cancellation on all active asyncio tasks, waits up to timeout seconds for them to finish, then force-marks any still pending as cancelled.

Covers every :data:ACTIVE_STATUSES agent, not just running: one still at submitted never had a loop task to cancel — and one still waiting for a slot never had a task at all — which is exactly why either would otherwise survive shutdown unmarked. Waiting units are dropped from the scheduler FIRST, so a manager settling during the wait window cannot pump a fresh child into a hypervisor that is being torn down.

Marking here is deliberately IN-MEMORY only — the hypervisor holds no event logger, and wiring one in would invert the layering (the session owns the transcript, the hypervisor owns the tree). The durable terminal is written by whichever peer can still reach the log:

  • While the task is still alive (during the request or the wait): cancelling it raises CancelledError inside the agent's own lifecycle handler, which writes the terminal stop before unwinding. That is the normal path.
  • Once the wait times out, at the final force-mark sweep (the task is wedged, or the process is being torn down mid-flight): nothing in-process can still write, so the durable terminal comes from the startup sweep on the NEXT boot, which settles agent spans left open under a run that has since ended.
Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
async def cleanup(self, timeout: float = 5.0) -> None:
    """Cancel all agents with graceful escalation.

    Requests cancellation on all active asyncio tasks, waits up to
    *timeout* seconds for them to finish, then force-marks any still
    pending as cancelled.

    Covers every :data:`ACTIVE_STATUSES` agent, not just ``running``: one
    still at ``submitted`` never had a loop task to cancel — and one still
    waiting for a slot never had a task at all — which is exactly why
    either would otherwise survive shutdown unmarked. Waiting units are
    dropped from the scheduler FIRST, so a manager settling during the
    wait window cannot pump a fresh child into a hypervisor that is being
    torn down.

    Marking here is deliberately IN-MEMORY only — the hypervisor holds no
    event logger, and wiring one in would invert the layering (the session
    owns the transcript, the hypervisor owns the tree). The durable terminal
    is written by whichever peer can still reach the log:

    - While the task is still alive (during the request or the wait):
      cancelling it raises ``CancelledError`` inside the agent's own
      lifecycle handler, which writes the terminal ``stop`` before
      unwinding. That is the normal path.
    - Once the wait times out, at the final force-mark sweep (the task is
      wedged, or the process is being torn down mid-flight): nothing
      in-process can still write, so the durable terminal comes from the
      startup sweep on the NEXT boot, which settles agent spans left open
      under a run that has since ended.
    """
    await self.cancel_pending()

    async with self._lock:
        active = [h for h in self._agents.values() if h.status in ACTIVE_STATUSES]

    if not active:
        async with self._lock:
            self._agents.clear()
        return

    # Request cancellation.
    for handle in active:
        if handle.asyncio_task and not handle.asyncio_task.done():
            handle.asyncio_task.cancel()

    # Wait with timeout.
    tasks = [h.asyncio_task for h in active if h.asyncio_task and not h.asyncio_task.done()]
    if tasks:
        done, pending = await asyncio.wait(
            tasks,
            timeout=timeout,
            return_when=asyncio.ALL_COMPLETED,
        )

        # Force-mark any still-pending as cancelled.
        for task in pending:
            for handle in active:
                if handle.asyncio_task is task:
                    async with self._lock:
                        if handle.status in ACTIVE_STATUSES:
                            handle.status = "cancelled"
                            handle.error = "Force-cancelled after timeout"
                            handle.stopped_at = time.monotonic()

    # Final sweep: mark any remaining non-terminal agents and clear.
    async with self._lock:
        for h in list(self._agents.values()):
            if h.status in ACTIVE_STATUSES:
                h.status = "cancelled"
                h.stopped_at = time.monotonic()
        self._agents.clear()

collect_completed(parent_id: str) -> list[AgentHandle] async

Return children that reached a terminal state with a stored result.

Async CU retrieval — parent reads results when ready, not when child finishes.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
async def collect_completed(self, parent_id: str) -> list[AgentHandle]:
    """Return children that reached a terminal state with a stored result.

    Async CU retrieval — parent reads results
    when ready, not when child finishes.
    """
    terminal = {"completed", "failed", "cancelled"}
    async with self._lock:
        return [
            h
            for h in self._agents.values()
            if h.parent_id == parent_id and h.status in terminal and h.result is not None
        ]

collect_running(parent_id: str) -> list[AgentHandle] async

Return children of parent_id that are still active.

Reads :data:ACTIVE_STATUSES rather than re-listing the states. That matters beyond tidiness: this method is the promise-as-completion gate's ownership index AND what check_agents reports, so a capacity-deferred child — submitted, registered, not yet started — must appear here. A refused unit was registered nowhere, which is exactly why a parent that over-subscribed the fleet was told all its work was done.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
async def collect_running(self, parent_id: str) -> list[AgentHandle]:
    """Return children of *parent_id* that are still active.

    Reads :data:`ACTIVE_STATUSES` rather than re-listing the states. That
    matters beyond tidiness: this method is the promise-as-completion
    gate's ownership index AND what ``check_agents`` reports, so a
    capacity-deferred child — ``submitted``, registered, not yet started —
    must appear here. A refused unit was registered nowhere, which is
    exactly why a parent that over-subscribed the fleet was told all its
    work was done.
    """
    async with self._lock:
        return [
            h
            for h in self._agents.values()
            if h.parent_id == parent_id and h.status in ACTIVE_STATUSES
        ]

get(agent_id: str) -> AgentHandle | None async

Return a single agent handle, or None.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
870
871
872
873
async def get(self, agent_id: str) -> AgentHandle | None:
    """Return a single agent handle, or None."""
    async with self._lock:
        return self._agents.get(agent_id)

list_all() -> list[AgentHandle] async

Full tree view — for CLI/API user visibility.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
893
894
895
896
async def list_all(self) -> list[AgentHandle]:
    """Full tree view — for CLI/API user visibility."""
    async with self._lock:
        return list(self._agents.values())

list_children(parent_id: str) -> list[AgentHandle] async

List direct children of a parent. Enforces isolation.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
875
876
877
878
async def list_children(self, parent_id: str) -> list[AgentHandle]:
    """List direct children of a parent. Enforces isolation."""
    async with self._lock:
        return [h for h in self._agents.values() if h.parent_id == parent_id]

list_descendants(ancestor_id: str) -> list[AgentHandle] async

List all descendants recursively.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
880
881
882
883
884
885
886
887
888
889
890
891
async def list_descendants(self, ancestor_id: str) -> list[AgentHandle]:
    """List all descendants recursively."""
    async with self._lock:
        result: list[AgentHandle] = []
        queue = [ancestor_id]
        while queue:
            pid = queue.pop()
            for h in self._agents.values():
                if h.parent_id == pid:
                    result.append(h)
                    queue.append(h.agent_id)
        return result

list_visible(exclude_agent_id: str | None = None) -> list[AgentHandle] async

Snapshot of agents excluding the caller — used by check_agents.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
898
899
900
901
async def list_visible(self, exclude_agent_id: str | None = None) -> list[AgentHandle]:
    """Snapshot of agents excluding the caller — used by check_agents."""
    async with self._lock:
        return [h for h in self._agents.values() if h.agent_id != exclude_agent_id]

mark_done(agent_id: str, status: AgentStatus, error: str | AgentError | None = None) -> None async

Mark an agent as done with a terminal status.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
851
852
853
854
855
856
857
858
859
860
861
862
863
864
async def mark_done(
    self,
    agent_id: str,
    status: AgentStatus,
    error: str | AgentError | None = None,
) -> None:
    """Mark an agent as done with a terminal status."""
    async with self._lock:
        handle = self._agents.get(agent_id)
        if handle:
            handle.status = status
            handle.error = error
            handle.stopped_at = time.monotonic()
            handle.done_event.set()

mark_tool_start(agent_id: str, tool_id: str | None) -> None async

Stamp (or clear) the tool actually in flight for stall attribution.

update_step only stamps last_tool_id on COMPLETION, so a watchdog check firing mid-call would misattribute the stall to the previous, already-finished tool. Call this at dispatch start with the tool name, and again with None once the call resolves (success, timeout, or exception) so active_tool_id never lingers stale once the agent moves on.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
836
837
838
839
840
841
842
843
844
845
846
847
848
849
async def mark_tool_start(self, agent_id: str, tool_id: str | None) -> None:
    """Stamp (or clear) the tool actually in flight for stall attribution.

    ``update_step`` only stamps ``last_tool_id`` on
    COMPLETION, so a watchdog check firing mid-call would misattribute the
    stall to the previous, already-finished tool. Call this at dispatch
    start with the tool name, and again with ``None`` once the call
    resolves (success, timeout, or exception) so ``active_tool_id`` never
    lingers stale once the agent moves on.
    """
    async with self._lock:
        handle = self._agents.get(agent_id)
        if handle:
            handle.active_tool_id = tool_id

over_wall_deadline_agents(now: float | None = None) -> list[tuple[AgentHandle, str]] async

Running agents whose contract wall-clock deadline is warn/over.

Mirrors :meth:stalled_agents — a per-contract cousin of the stall sweep, keyed on started_at rather than last-progress. now arrives as a clock ARG (never read internally), so a caller can drive this deterministically without sleeping a real clock.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
async def over_wall_deadline_agents(
    self, now: float | None = None
) -> list[tuple[AgentHandle, str]]:
    """Running agents whose contract wall-clock deadline is warn/over.

    Mirrors :meth:`stalled_agents` — a per-contract cousin of the stall
    sweep, keyed on ``started_at`` rather than last-progress. ``now``
    arrives as a clock ARG (never read internally), so a caller can drive
    this deterministically without sleeping a real clock.
    """
    clock = now if now is not None else time.monotonic()
    async with self._lock:
        result: list[tuple[AgentHandle, str]] = []
        for h in self._agents.values():
            if h.status != "running" or h.contract.max_wall_s <= 0:
                continue
            state = h.contract.wall_state(clock - h.started_at)
            if state in ("warn", "over"):
                result.append((h, state))
        return result

record_compaction(agent_id: str) -> None async

Record that an agent compacted its context.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1031
1032
1033
1034
1035
1036
1037
async def record_compaction(self, agent_id: str) -> None:
    """Record that an agent compacted its context."""
    async with self._lock:
        handle = self._agents.get(agent_id)
        if handle:
            handle.compaction_count += 1
            handle.last_compacted_at = time.monotonic()

register(handle: AgentHandle) -> None async

Register a new agent in the hypervisor.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
803
804
805
806
async def register(self, handle: AgentHandle) -> None:
    """Register a new agent in the hypervisor."""
    async with self._lock:
        self._agents[handle.agent_id] = handle

release() -> None async

Settle one running agent's slot, dispatching whatever waits on it.

Async because it is the DISPATCH PUMP, not a bare counter decrement: the freed slot is handed to the next waiting unit, which means starting it. This is deliberately the completion path every settling agent already calls, and exactly-once is already guaranteed there — a pump hung off on_agent_stop instead would fire while the slot is still held and find the queue full every time.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
732
733
734
735
736
737
738
739
740
741
742
async def release(self) -> None:
    """Settle one running agent's slot, dispatching whatever waits on it.

    Async because it is the DISPATCH PUMP, not a bare counter decrement:
    the freed slot is handed to the next waiting unit, which means starting
    it. This is deliberately the completion path every settling agent
    already calls, and exactly-once is already guaranteed there — a pump
    hung off ``on_agent_stop`` instead would fire while the slot is still
    held and find the queue full every time.
    """
    await self._queue.release()

render_agent_tree(*, exclude_agent_id: str | None = None) -> str async

Render a concise text summary of the agent tree for the root's system prompt.

Structural transparency — the root agent (the hypervisor's "brain") gets a live view of all agents so it can reason about the delegation state and intervene if needed.

Parameters:

Name Type Description Default
exclude_agent_id str | None

If provided, omit this agent from the rendered tree. Used so the calling agent does not see itself listed.

None
Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
async def render_agent_tree(
    self,
    *,
    exclude_agent_id: str | None = None,
) -> str:
    """Render a concise text summary of the agent tree for the root's system prompt.

    Structural transparency — the root
    agent (the hypervisor's "brain") gets a live view of all agents so
    it can reason about the delegation state and intervene if needed.

    Args:
        exclude_agent_id: If provided, omit this agent from the rendered
            tree.  Used so the calling agent does not see itself listed.
    """
    async with self._lock:
        if not self._agents:
            return ""
        visible = [h for h in self._agents.values() if h.agent_id != exclude_agent_id]
        if not visible:
            return ""
        registry = get_prompt_registry()
        counts: dict[str, int] = {}
        for h in visible:
            counts[h.status] = counts.get(h.status, 0) + 1
        status_parts = [f"{v} {k}" for k, v in sorted(counts.items())]
        budget_str = ""
        if self._session_step_budget > 0:
            budget_str = registry.render(
                "catalog.agent_tree.budget",
                total_steps=self._total_steps,
                session_step_budget=self._session_step_budget,
            )
        header = registry.render(
            "catalog.agent_tree.header",
            status_parts=", ".join(status_parts),
            budget_str=budget_str,
        )
        lines = [header]
        for h in sorted(visible, key=lambda x: (x.depth, x.agent_id)):
            indent = "  " * h.depth
            step_info = f"{h.steps_completed} steps"
            if h.last_tool_id:
                step_info += registry.render(
                    "catalog.agent_tree.step_info_last", last_tool_id=h.last_tool_id
                )
            # Surface bounded-retry provenance inline so the
            # root sees a child was re-delegated (omitted when never
            # retried, so an ordinary line carries no attempt count).
            if h.attempts > 1:
                step_info += f", {h.attempts} attempts"
            status_marker = ""
            if h.status == "completed":
                status_marker = " -> success"
            elif h.status == "failed":
                status_marker = " -> FAILED"
            elif h.status == "cancelled":
                status_marker = " -> cancelled"
            # Progress/result in tree view
            extra = ""
            if h.result and h.result.summary:
                extra = registry.render(
                    "catalog.agent_tree.result",
                    status=h.result.status,
                    summary=h.result.summary[:120],
                )
            elif h.progress_note:
                extra = registry.render(
                    "catalog.agent_tree.progress",
                    progress_note=h.progress_note[:120],
                )
            compact_marker = (
                registry.render(
                    "catalog.agent_tree.compact", compaction_count=h.compaction_count
                )
                if h.compaction_count
                else ""
            )
            task_preview = h.task_description[:80]
            line = registry.render(
                "catalog.agent_tree.line",
                indent=indent,
                agent_id_head=h.agent_id[:8],
                status=h.status,
                task_preview=task_preview,
                step_info=step_info,
                status_marker=status_marker,
                compact_marker=compact_marker,
                extra=extra,
            )
            # Appended OUTSIDE the templated line (never
            # a catalog.yaml edit) so a contract-less row's bytes are
            # completely unaffected.
            lines.append(line + self._contract_marker(h))
        return "\n".join(lines)

send_message(agent_id: str, message: str) -> str | None async

Send a steering message to a running agent.

Returns None on success, or a diagnostic string on failure.

Adaptive coordination — the hypervisor injects NL feedback (budget warnings, stall nudges) into the agent's message queue. The ToolUseLoop drains this queue between steps. Bidirectional system↔agent feedback loop.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
async def send_message(self, agent_id: str, message: str) -> str | None:
    """Send a steering message to a running agent.

    Returns ``None`` on success, or a diagnostic string on failure.

    Adaptive coordination — the hypervisor
    injects NL feedback (budget warnings, stall nudges) into the agent's
    message queue. The ToolUseLoop drains this queue between steps.
    Bidirectional system↔agent feedback loop.
    """
    async with self._lock:
        handle = self._agents.get(agent_id)
        if not handle:
            return "agent not in registry"
        if handle.status != "running":
            return f"agent status is '{handle.status}'"
        if not handle.message_queue:
            return "no message queue"
        handle.message_queue.put_nowait(message)
        return None

send_to_parent(child_agent_id: str, message: str) -> str | None async

Route a message from a child agent to its parent's queue.

Returns None on success, or a diagnostic string on failure.

Bidirectional message passing — enables lifecycle manager to notify parent on child completion.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
async def send_to_parent(self, child_agent_id: str, message: str) -> str | None:
    """Route a message from a child agent to its parent's queue.

    Returns ``None`` on success, or a diagnostic string on failure.

    Bidirectional message passing —
    enables lifecycle manager to notify parent on child completion.
    """
    async with self._lock:
        child = self._agents.get(child_agent_id)
        if not child or not child.parent_id:
            return "child not in registry or no parent"
        parent = self._agents.get(child.parent_id)
        if not parent:
            return "parent not in registry"
        if not parent.message_queue:
            return "parent has no message queue"
        parent.message_queue.put_nowait(message)
        return None

stalled_agents(threshold: float = 120.0) -> list[AgentHandle] async

Return agents not making progress within threshold seconds.

Internal trigger: delegatee unresponsive → diagnose → evaluate → intervene.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
async def stalled_agents(self, threshold: float = 120.0) -> list[AgentHandle]:
    """Return agents not making progress within threshold seconds.

    Internal trigger: delegatee
    unresponsive → diagnose → evaluate → intervene.
    """
    now = time.monotonic()
    async with self._lock:
        return [
            h
            for h in self._agents.values()
            if h.status == "running"
            and h.last_step_at is not None
            and (now - h.last_step_at) > threshold
        ]

unregister(agent_id: str) -> AgentHandle | None async

Remove an agent from the hypervisor.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
808
809
810
811
812
813
814
815
async def unregister(self, agent_id: str) -> AgentHandle | None:
    """Remove an agent from the hypervisor."""
    async with self._lock:
        handle = self._agents.pop(agent_id, None)
        if handle and handle.status == "running":
            handle.status = "completed"
            handle.stopped_at = time.monotonic()
        return handle

update_step(agent_id: str, tool_id: str) -> None async

Record a completed tool execution step.

Track at tool-call granularity, not agent granularity. Updates last_step_at for stall detection and total_steps for session budget enforcement.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
821
822
823
824
825
826
827
828
829
830
831
832
833
834
async def update_step(self, agent_id: str, tool_id: str) -> None:
    """Record a completed tool execution step.

    Track at tool-call granularity, not agent
    granularity. Updates last_step_at for stall detection and total_steps
    for session budget enforcement.
    """
    async with self._lock:
        handle = self._agents.get(agent_id)
        if handle:
            handle.steps_completed += 1
            handle.last_tool_id = tool_id
            handle.last_step_at = time.monotonic()
            self._total_steps += 1

mewbo_core.agents.hypervisor

Agent hypervisor — active governor for multi-agent delegation.

The hypervisor is the centralized control plane that manages the full agent tree for a session. It acts as an active governor (not just a registry), mediating every delegation decision and result handoff.

Scientific grounding: - Adaptive coordination cycle — the hypervisor monitors agents and intervenes via NL feedback when triggers fire. - Structural transparency — configurable monitoring with lifecycle events at each phase transition. - Graduated enforcement — warn, throttle, feedback (never kill first; killing destroys 31-48% of accumulated context). - Task state machine — agents progress through submitted → running → completed/failed/cancelled/rejected. - Communication Units — structured AgentResult with compressed summary field enables inter-agent context passing.

Responsibilities: - Admission control — bounding concurrent agents via semaphore. - Lifecycle tracking — AgentHandle with delegation phase awareness. - Active monitoring — budget tracking, stall detection, NL interventions. - Global eye — render_agent_tree() gives the root agent a live view. - Bidirectional messaging — send_message() enables parent→child steering. - Structured results — AgentResult carries Communication Units between agents. - Graceful shutdown — 3-phase escalation: cancel → wait → force-mark.

AgentHandle dataclass

Mutable runtime state for a single agent — lives in the hypervisor.

Created when an agent registers and updated throughout its lifecycle. Fields are read by the CLI agent tree display and the hypervisor's query/cancellation methods.

Handle starts as submitted, transitions to running when the loop begins, then to a terminal state. last_step_at enables tool-call-granularity stall detection without destroying accumulated context. message_queue enables bidirectional adaptive coordination — the hypervisor injects NL feedback.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
@dataclass(slots=True)
class AgentHandle:
    """Mutable runtime state for a single agent — lives in the hypervisor.

    Created when an agent registers and updated throughout its lifecycle.
    Fields are read by the CLI agent tree display and the hypervisor's
    query/cancellation methods.

    Handle starts as ``submitted``, transitions to
    ``running`` when the loop begins, then to a terminal state.
    ``last_step_at`` enables tool-call-granularity
    stall detection without destroying accumulated context.
    ``message_queue`` enables bidirectional
    adaptive coordination — the hypervisor injects NL feedback.
    """

    agent_id: str
    parent_id: str | None
    depth: int
    model_name: str
    task_description: str
    # The registered AgentDef name this agent was spawned as (e.g.
    # ``scg-path-probe``) — ``None`` for an ad-hoc spawn with no ``agent_type``.
    # Distinct from ``model_name``: the trace projection needs the LANE identity
    # (the def), which a model name can never carry.
    agent_type: str | None = None
    status: AgentStatus = "submitted"  # Start as submitted
    started_at: float = field(default_factory=time.monotonic)
    stopped_at: float | None = None
    steps_completed: int = 0
    last_tool_id: str | None = None
    # Tool-call-granularity timing for stall detection
    last_step_at: float | None = None
    # The tool actually IN FLIGHT right now — stamped at dispatch start and
    # cleared back to ``None`` when the call resolves. Distinct from
    # ``last_tool_id`` (only updated on COMPLETION): during a long-running
    # call, ``last_tool_id`` still names the PREVIOUS finished tool, so the
    # watchdog must read ``active_tool_id`` for stall attribution, never
    # ``last_tool_id``.
    active_tool_id: str | None = None
    error: str | AgentError | None = None
    asyncio_task: asyncio.Task[object] | None = None
    # Bidirectional message passing
    message_queue: queue.Queue[str] | None = None
    # Completed CU stored on handle for async retrieval
    result: AgentResult | None = None
    # Auto-updated progress for monitoring
    progress_note: str | None = None
    # Context compaction tracking — visible in agent tree rendering.
    compaction_count: int = 0
    last_compacted_at: float | None = None
    # Token usage — accumulated from LLM response.usage_metadata per call.
    # Written only by the owning ToolUseLoop coroutine; read by CLI/API.
    input_tokens: int = 0
    output_tokens: int = 0
    # Signaled when agent reaches a terminal state (completed/failed/cancelled).
    done_event: asyncio.Event = field(default_factory=asyncio.Event)
    # Bumped by the spawn bridge's bounded-retry driver each
    # time this same task is re-admitted (1 = first/only attempt). Read by the
    # agent-tree render + check_agents payload so a retried child is visible.
    attempts: int = 1
    # The spawner's own delegation bounds for this child.
    # Additive: the default (disabled) contract is a no-op, so an ad-hoc
    # spawn with"no ``contract`` behaves exactly as before.
    contract: DelegationContract = field(default_factory=DelegationContract)
    # This child's OWN spawn attestation record hash
    # (stamped by ``SpawnAgentTool`` right after registration), threaded back
    # in as the terminal record's ``spawn_hash`` link. "" when no chain is
    # wired for this session.
    attestation_spawn_hash: str = ""
    # Latched when this agent's ONE terminal ``stop`` lifecycle event is written.
    # Several paths can settle the same agent (its own success/cancel/failure
    # handler, and its parent's teardown cascade), and they RACE — the flag is
    # what keeps "exactly one stop per start" true rather than merely likely.
    # Written only through the spawn bridge's terminal emitter.
    terminal_emitted: bool = False

AgentHypervisor

Hypervisor control plane — manages the full agent tree for a session.

Thread-safe via asyncio.Lock. A single instance is shared across all agents in the hierarchy through AgentContext.registry.

Responsibilities
  • Admission control: accept() / accept_batch() / release() gate concurrency through an injected :class:AgentQueue (default 20 running slots) that QUEUES over-cap work instead of dropping it.
  • Registration: register() / unregister() track agent handles keyed by agent_id.
  • Status: update_step() / mark_done() record execution progress and terminal state.
  • Queries: list_children() / list_descendants() / list_all() expose the live tree for display and introspection.
  • Cancellation: cancel_agent() cancels a single agent; cleanup() tears down the entire tree on session exit.
Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
class AgentHypervisor:
    """Hypervisor control plane — manages the full agent tree for a session.

    Thread-safe via ``asyncio.Lock``. A single instance is shared across all
    agents in the hierarchy through ``AgentContext.registry``.

    Responsibilities:
        - **Admission control**: ``accept()`` / ``accept_batch()`` /
          ``release()`` gate concurrency through an injected
          :class:`AgentQueue` (default 20 running slots) that QUEUES over-cap
          work instead of dropping it.
        - **Registration**: ``register()`` / ``unregister()`` track agent
          handles keyed by ``agent_id``.
        - **Status**: ``update_step()`` / ``mark_done()`` record execution
          progress and terminal state.
        - **Queries**: ``list_children()`` / ``list_descendants()`` /
          ``list_all()`` expose the live tree for display and introspection.
        - **Cancellation**: ``cancel_agent()`` cancels a single agent;
          ``cleanup()`` tears down the entire tree on session exit.
    """

    def __init__(
        self,
        *,
        max_concurrent: int = 20,
        session_step_budget: int = 0,
        attestation: AttestationChain | None = None,
        queue: AgentQueue | None = None,
    ) -> None:
        """Initialize hypervisor with concurrency and budget limits.

        Session-wide budget with graduated enforcement.

        Args:
            max_concurrent: Maximum number of concurrently RUNNING agents.
            session_step_budget: Total tool steps allowed across all agents in the
                session. 0 means unlimited.
            attestation: optional provenance hash chain for
                this session's agent tree. The hypervisor never constructs one
                itself (it has no session_id, no store) — the orchestrator
                injects it, seeded from the store's last persisted record, when
                the feature is config-enabled. ``None`` (the default) means no
                attestation is recorded; ``SpawnAgentTool`` reads this
                attribute directly and no-ops when it's absent.
            queue: the admission scheduler. Injected so a deployment (or a
                test) can seat a differently-sized or differently-ordered one
                without the hypervisor learning how scheduling works; the
                default is an :class:`AgentQueue` of ``max_concurrent`` slots.
        """
        self._agents: dict[str, AgentHandle] = {}
        self._lock: asyncio.Lock = asyncio.Lock()
        self._queue: AgentQueue = queue or AgentQueue(capacity=max_concurrent)
        # Set unconditionally so an INJECTED queue is wired too: the queue can
        # detect a failed dispatch but only the hypervisor can settle the
        # handle it stranded.
        self._queue.on_dispatch_failed = self._settle_failed_dispatch
        # Session-wide resource tracking
        self._total_steps: int = 0
        self._session_step_budget: int = session_step_budget
        self.attestation = attestation

    # ------------------------------------------------------------------
    # Admission control
    # ------------------------------------------------------------------

    async def accept(self, unit: SpawnUnit) -> bool:
        """Admit ONE spawn: ``True`` = dispatched now, ``False`` = deferred.

        Never refuses for capacity — that is the whole point of the queue.
        """
        return await self._queue.accept(unit)

    async def accept_batch(self, units: Sequence[SpawnUnit]) -> list[bool]:
        """Admit a whole fan-out at once — see :meth:`AgentQueue.accept_batch`."""
        return await self._queue.accept_batch(units)

    async def release(self) -> None:
        """Settle one running agent's slot, dispatching whatever waits on it.

        Async because it is the DISPATCH PUMP, not a bare counter decrement:
        the freed slot is handed to the next waiting unit, which means starting
        it. This is deliberately the completion path every settling agent
        already calls, and exactly-once is already guaranteed there — a pump
        hung off ``on_agent_stop`` instead would fire while the slot is still
        held and find the queue full every time.
        """
        await self._queue.release()

    async def cancel_pending(self) -> list[str]:
        """Drop every accepted-but-undispatched spawn, settling each as cancelled.

        A waiting agent holds no slot and owns no ``asyncio.Task``, so neither
        ``cancel_agent`` nor the cleanup sweep's task cancellation can reach
        it — it would sit ``submitted`` forever and read as live. Returns the
        ids dropped. Called before teardown so a manager settling on the way
        out cannot pump a brand-new child into a run that is already ending.
        """
        dropped = self._queue.clear()
        if not dropped:
            return []
        now = time.monotonic()
        async with self._lock:
            for spawn in dropped:
                handle = self._agents.get(spawn.agent_id)
                if handle is not None and handle.status in ACTIVE_STATUSES:
                    handle.status = "cancelled"
                    handle.error = "cancelled while waiting for a concurrency slot"
                    handle.stopped_at = now
                    handle.done_event.set()
        return [spawn.agent_id for spawn in dropped]

    async def _settle_failed_dispatch(
        self, spawn: ScheduledSpawn, exc: BaseException
    ) -> None:
        """Settle an agent whose launcher raised — it will never start.

        Its handle was registered at acceptance and its spawn path already
        returned, so nothing downstream owns it: it would sit non-terminal
        forever, keep answering ``collect_running``, and hold the
        promise-as-completion gate open for the rest of the session. The
        message is the exception's TYPE plus a bounded head — a launcher
        failure is infrastructure, and an unbounded provider string has no
        business on a handle read by every client.
        """
        detail = f"{type(exc).__name__}: {exc}"[:200]
        await self.mark_done(
            spawn.agent_id, "failed", error=f"sub-agent never started — {detail}"
        )

    @property
    def free_slots(self) -> int:
        """Slots a spawn could be dispatched into right now."""
        return self._queue.free_slots

    @property
    def pending_dispatch(self) -> int:
        """Accepted spawns not yet started.

        NOT a status: those agents are ``submitted`` like any other, and this
        is only the scheduler's count of how many are still waiting on a slot.
        """
        return self._queue.waiting

    # ------------------------------------------------------------------
    # Registration
    # ------------------------------------------------------------------

    async def register(self, handle: AgentHandle) -> None:
        """Register a new agent in the hypervisor."""
        async with self._lock:
            self._agents[handle.agent_id] = handle

    async def unregister(self, agent_id: str) -> AgentHandle | None:
        """Remove an agent from the hypervisor."""
        async with self._lock:
            handle = self._agents.pop(agent_id, None)
            if handle and handle.status == "running":
                handle.status = "completed"
                handle.stopped_at = time.monotonic()
            return handle

    # ------------------------------------------------------------------
    # Status updates
    # ------------------------------------------------------------------

    async def update_step(self, agent_id: str, tool_id: str) -> None:
        """Record a completed tool execution step.

        Track at tool-call granularity, not agent
        granularity. Updates last_step_at for stall detection and total_steps
        for session budget enforcement.
        """
        async with self._lock:
            handle = self._agents.get(agent_id)
            if handle:
                handle.steps_completed += 1
                handle.last_tool_id = tool_id
                handle.last_step_at = time.monotonic()
                self._total_steps += 1

    async def mark_tool_start(self, agent_id: str, tool_id: str | None) -> None:
        """Stamp (or clear) the tool actually in flight for stall attribution.

        ``update_step`` only stamps ``last_tool_id`` on
        COMPLETION, so a watchdog check firing mid-call would misattribute the
        stall to the previous, already-finished tool. Call this at dispatch
        start with the tool name, and again with ``None`` once the call
        resolves (success, timeout, or exception) so ``active_tool_id`` never
        lingers stale once the agent moves on.
        """
        async with self._lock:
            handle = self._agents.get(agent_id)
            if handle:
                handle.active_tool_id = tool_id

    async def mark_done(
        self,
        agent_id: str,
        status: AgentStatus,
        error: str | AgentError | None = None,
    ) -> None:
        """Mark an agent as done with a terminal status."""
        async with self._lock:
            handle = self._agents.get(agent_id)
            if handle:
                handle.status = status
                handle.error = error
                handle.stopped_at = time.monotonic()
                handle.done_event.set()

    # ------------------------------------------------------------------
    # Queries
    # ------------------------------------------------------------------

    async def get(self, agent_id: str) -> AgentHandle | None:
        """Return a single agent handle, or None."""
        async with self._lock:
            return self._agents.get(agent_id)

    async def list_children(self, parent_id: str) -> list[AgentHandle]:
        """List direct children of a parent. Enforces isolation."""
        async with self._lock:
            return [h for h in self._agents.values() if h.parent_id == parent_id]

    async def list_descendants(self, ancestor_id: str) -> list[AgentHandle]:
        """List all descendants recursively."""
        async with self._lock:
            result: list[AgentHandle] = []
            queue = [ancestor_id]
            while queue:
                pid = queue.pop()
                for h in self._agents.values():
                    if h.parent_id == pid:
                        result.append(h)
                        queue.append(h.agent_id)
            return result

    async def list_all(self) -> list[AgentHandle]:
        """Full tree view — for CLI/API user visibility."""
        async with self._lock:
            return list(self._agents.values())

    async def list_visible(self, exclude_agent_id: str | None = None) -> list[AgentHandle]:
        """Snapshot of agents excluding the caller — used by check_agents."""
        async with self._lock:
            return [h for h in self._agents.values() if h.agent_id != exclude_agent_id]

    # ------------------------------------------------------------------
    # Budget & monitoring
    # ------------------------------------------------------------------

    @property
    def total_steps(self) -> int:
        """Total tool steps executed across all agents in the session."""
        return self._total_steps

    def budget_exhausted(self) -> bool:
        """Check if the session step budget is exhausted (0 = unlimited).

        Graduated enforcement — this is the trigger
        check. The response (NL warning injection) happens in ToolUseLoop.
        """
        return self._session_step_budget > 0 and self._total_steps >= self._session_step_budget

    def budget_remaining(self) -> int:
        """Steps remaining in the session budget (0 = unlimited)."""
        if self._session_step_budget <= 0:
            return -1  # Unlimited
        return max(0, self._session_step_budget - self._total_steps)

    def budget_warning(self, *, headroom: int = 5) -> bool:
        """True when within ``headroom`` steps of the session budget (0 = off)."""
        if self._session_step_budget <= 0:
            return False
        return self.budget_remaining() <= headroom

    async def stalled_agents(self, threshold: float = 120.0) -> list[AgentHandle]:
        """Return agents not making progress within threshold seconds.

        Internal trigger: delegatee
        unresponsive → diagnose → evaluate → intervene.
        """
        now = time.monotonic()
        async with self._lock:
            return [
                h
                for h in self._agents.values()
                if h.status == "running"
                and h.last_step_at is not None
                and (now - h.last_step_at) > threshold
            ]

    async def agent_step_state(self, agent_id: str) -> str:
        """Per-contract step-budget state for one agent.

        ``"ok"`` for an unregistered agent or one whose contract carries no
        step bound — the same inert default as a disabled contract.
        """
        async with self._lock:
            handle = self._agents.get(agent_id)
        if handle is None:
            return "ok"
        return handle.contract.step_state(handle.steps_completed)

    async def agent_token_state(self, agent_id: str) -> str:
        """Per-contract advisory token state for one agent."""
        async with self._lock:
            handle = self._agents.get(agent_id)
        if handle is None:
            return "ok"
        return handle.contract.token_state(handle.input_tokens + handle.output_tokens)

    async def over_wall_deadline_agents(
        self, now: float | None = None
    ) -> list[tuple[AgentHandle, str]]:
        """Running agents whose contract wall-clock deadline is warn/over.

        Mirrors :meth:`stalled_agents` — a per-contract cousin of the stall
        sweep, keyed on ``started_at`` rather than last-progress. ``now``
        arrives as a clock ARG (never read internally), so a caller can drive
        this deterministically without sleeping a real clock.
        """
        clock = now if now is not None else time.monotonic()
        async with self._lock:
            result: list[tuple[AgentHandle, str]] = []
            for h in self._agents.values():
                if h.status != "running" or h.contract.max_wall_s <= 0:
                    continue
                state = h.contract.wall_state(clock - h.started_at)
                if state in ("warn", "over"):
                    result.append((h, state))
            return result

    @staticmethod
    def _contract_marker(handle: AgentHandle) -> str:
        """Compact ``| contract: ...`` suffix for an agent-tree row.

        Empty (no-op) unless the handle's contract is ``enabled``, so a
        contract-less child's row carries no suffix.
        """
        if not handle.contract.enabled:
            return ""
        bits: list[str] = [handle.contract.autonomy]
        if handle.contract.max_steps > 0:
            bits.append(f"{handle.steps_completed}/{handle.contract.max_steps} steps")
        if handle.contract.max_wall_s > 0:
            elapsed = time.monotonic() - handle.started_at
            bits.append(f"{elapsed:.0f}/{handle.contract.max_wall_s:.0f}s")
        return f" | contract: {', '.join(bits)}"

    # ------------------------------------------------------------------
    # Bidirectional messaging
    # ------------------------------------------------------------------

    async def send_message(self, agent_id: str, message: str) -> str | None:
        """Send a steering message to a running agent.

        Returns ``None`` on success, or a diagnostic string on failure.

        Adaptive coordination — the hypervisor
        injects NL feedback (budget warnings, stall nudges) into the agent's
        message queue. The ToolUseLoop drains this queue between steps.
        Bidirectional system↔agent feedback loop.
        """
        async with self._lock:
            handle = self._agents.get(agent_id)
            if not handle:
                return "agent not in registry"
            if handle.status != "running":
                return f"agent status is '{handle.status}'"
            if not handle.message_queue:
                return "no message queue"
            handle.message_queue.put_nowait(message)
            return None

    async def record_compaction(self, agent_id: str) -> None:
        """Record that an agent compacted its context."""
        async with self._lock:
            handle = self._agents.get(agent_id)
            if handle:
                handle.compaction_count += 1
                handle.last_compacted_at = time.monotonic()

    # ------------------------------------------------------------------
    # Global eye
    # ------------------------------------------------------------------

    async def render_agent_tree(
        self,
        *,
        exclude_agent_id: str | None = None,
    ) -> str:
        """Render a concise text summary of the agent tree for the root's system prompt.

        Structural transparency — the root
        agent (the hypervisor's "brain") gets a live view of all agents so
        it can reason about the delegation state and intervene if needed.

        Args:
            exclude_agent_id: If provided, omit this agent from the rendered
                tree.  Used so the calling agent does not see itself listed.
        """
        async with self._lock:
            if not self._agents:
                return ""
            visible = [h for h in self._agents.values() if h.agent_id != exclude_agent_id]
            if not visible:
                return ""
            registry = get_prompt_registry()
            counts: dict[str, int] = {}
            for h in visible:
                counts[h.status] = counts.get(h.status, 0) + 1
            status_parts = [f"{v} {k}" for k, v in sorted(counts.items())]
            budget_str = ""
            if self._session_step_budget > 0:
                budget_str = registry.render(
                    "catalog.agent_tree.budget",
                    total_steps=self._total_steps,
                    session_step_budget=self._session_step_budget,
                )
            header = registry.render(
                "catalog.agent_tree.header",
                status_parts=", ".join(status_parts),
                budget_str=budget_str,
            )
            lines = [header]
            for h in sorted(visible, key=lambda x: (x.depth, x.agent_id)):
                indent = "  " * h.depth
                step_info = f"{h.steps_completed} steps"
                if h.last_tool_id:
                    step_info += registry.render(
                        "catalog.agent_tree.step_info_last", last_tool_id=h.last_tool_id
                    )
                # Surface bounded-retry provenance inline so the
                # root sees a child was re-delegated (omitted when never
                # retried, so an ordinary line carries no attempt count).
                if h.attempts > 1:
                    step_info += f", {h.attempts} attempts"
                status_marker = ""
                if h.status == "completed":
                    status_marker = " -> success"
                elif h.status == "failed":
                    status_marker = " -> FAILED"
                elif h.status == "cancelled":
                    status_marker = " -> cancelled"
                # Progress/result in tree view
                extra = ""
                if h.result and h.result.summary:
                    extra = registry.render(
                        "catalog.agent_tree.result",
                        status=h.result.status,
                        summary=h.result.summary[:120],
                    )
                elif h.progress_note:
                    extra = registry.render(
                        "catalog.agent_tree.progress",
                        progress_note=h.progress_note[:120],
                    )
                compact_marker = (
                    registry.render(
                        "catalog.agent_tree.compact", compaction_count=h.compaction_count
                    )
                    if h.compaction_count
                    else ""
                )
                task_preview = h.task_description[:80]
                line = registry.render(
                    "catalog.agent_tree.line",
                    indent=indent,
                    agent_id_head=h.agent_id[:8],
                    status=h.status,
                    task_preview=task_preview,
                    step_info=step_info,
                    status_marker=status_marker,
                    compact_marker=compact_marker,
                    extra=extra,
                )
                # Appended OUTSIDE the templated line (never
                # a catalog.yaml edit) so a contract-less row's bytes are
                # completely unaffected.
                lines.append(line + self._contract_marker(h))
            return "\n".join(lines)

    # ------------------------------------------------------------------
    # Async delegation queries
    # ------------------------------------------------------------------

    async def collect_completed(self, parent_id: str) -> list[AgentHandle]:
        """Return children that reached a terminal state with a stored result.

        Async CU retrieval — parent reads results
        when ready, not when child finishes.
        """
        terminal = {"completed", "failed", "cancelled"}
        async with self._lock:
            return [
                h
                for h in self._agents.values()
                if h.parent_id == parent_id and h.status in terminal and h.result is not None
            ]

    async def collect_running(self, parent_id: str) -> list[AgentHandle]:
        """Return children of *parent_id* that are still active.

        Reads :data:`ACTIVE_STATUSES` rather than re-listing the states. That
        matters beyond tidiness: this method is the promise-as-completion
        gate's ownership index AND what ``check_agents`` reports, so a
        capacity-deferred child — ``submitted``, registered, not yet started —
        must appear here. A refused unit was registered nowhere, which is
        exactly why a parent that over-subscribed the fleet was told all its
        work was done.
        """
        async with self._lock:
            return [
                h
                for h in self._agents.values()
                if h.parent_id == parent_id and h.status in ACTIVE_STATUSES
            ]

    async def send_to_parent(self, child_agent_id: str, message: str) -> str | None:
        """Route a message from a child agent to its parent's queue.

        Returns ``None`` on success, or a diagnostic string on failure.

        Bidirectional message passing —
        enables lifecycle manager to notify parent on child completion.
        """
        async with self._lock:
            child = self._agents.get(child_agent_id)
            if not child or not child.parent_id:
                return "child not in registry or no parent"
            parent = self._agents.get(child.parent_id)
            if not parent:
                return "parent not in registry"
            if not parent.message_queue:
                return "parent has no message queue"
            parent.message_queue.put_nowait(message)
            return None

    # ------------------------------------------------------------------
    # Cancellation
    # ------------------------------------------------------------------

    async def cancel_agent(self, agent_id: str) -> str | None:
        """Cancel an agent, whether it has started or is still waiting for a slot.

        Returns ``None`` on success, or a diagnostic string on failure.

        **The waiting case is checked FIRST and it is not an optimization.**
        ``asyncio_task`` is populated by the child driver, which runs only
        AFTER dispatch — so a capacity-deferred agent has no task, the
        task-cancellation path reported ``no asyncio task``, and the agent then
        STARTED anyway the moment a sibling freed a slot. Refusing to cancel
        something and then running it is the worst of both answers, and it is
        reachable the instant a fan-out exceeds the concurrency limit.
        Dropping it from the scheduler is what makes the cancel real.
        """
        async with self._lock:
            handle = self._agents.get(agent_id)
            if not handle:
                return "agent not in registry"
            if self._queue.discard(agent_id) is not None:
                handle.status = "cancelled"
                handle.error = "cancelled before it was dispatched"
                handle.stopped_at = time.monotonic()
                handle.done_event.set()
                return None
            if not handle.asyncio_task:
                return "no asyncio task"
            if handle.asyncio_task.done():
                return f"task already done (status: {handle.status})"
            handle.asyncio_task.cancel()
            handle.status = "cancelled"
            handle.stopped_at = time.monotonic()
            handle.done_event.set()
            return None

    # ------------------------------------------------------------------
    # Cleanup
    # ------------------------------------------------------------------

    async def cleanup(self, timeout: float = 5.0) -> None:
        """Cancel all agents with graceful escalation.

        Requests cancellation on all active asyncio tasks, waits up to
        *timeout* seconds for them to finish, then force-marks any still
        pending as cancelled.

        Covers every :data:`ACTIVE_STATUSES` agent, not just ``running``: one
        still at ``submitted`` never had a loop task to cancel — and one still
        waiting for a slot never had a task at all — which is exactly why
        either would otherwise survive shutdown unmarked. Waiting units are
        dropped from the scheduler FIRST, so a manager settling during the
        wait window cannot pump a fresh child into a hypervisor that is being
        torn down.

        Marking here is deliberately IN-MEMORY only — the hypervisor holds no
        event logger, and wiring one in would invert the layering (the session
        owns the transcript, the hypervisor owns the tree). The durable terminal
        is written by whichever peer can still reach the log:

        - While the task is still alive (during the request or the wait):
          cancelling it raises ``CancelledError`` inside the agent's own
          lifecycle handler, which writes the terminal ``stop`` before
          unwinding. That is the normal path.
        - Once the wait times out, at the final force-mark sweep (the task is
          wedged, or the process is being torn down mid-flight): nothing
          in-process can still write, so the durable terminal comes from the
          startup sweep on the NEXT boot, which settles agent spans left open
          under a run that has since ended.
        """
        await self.cancel_pending()

        async with self._lock:
            active = [h for h in self._agents.values() if h.status in ACTIVE_STATUSES]

        if not active:
            async with self._lock:
                self._agents.clear()
            return

        # Request cancellation.
        for handle in active:
            if handle.asyncio_task and not handle.asyncio_task.done():
                handle.asyncio_task.cancel()

        # Wait with timeout.
        tasks = [h.asyncio_task for h in active if h.asyncio_task and not h.asyncio_task.done()]
        if tasks:
            done, pending = await asyncio.wait(
                tasks,
                timeout=timeout,
                return_when=asyncio.ALL_COMPLETED,
            )

            # Force-mark any still-pending as cancelled.
            for task in pending:
                for handle in active:
                    if handle.asyncio_task is task:
                        async with self._lock:
                            if handle.status in ACTIVE_STATUSES:
                                handle.status = "cancelled"
                                handle.error = "Force-cancelled after timeout"
                                handle.stopped_at = time.monotonic()

        # Final sweep: mark any remaining non-terminal agents and clear.
        async with self._lock:
            for h in list(self._agents.values()):
                if h.status in ACTIVE_STATUSES:
                    h.status = "cancelled"
                    h.stopped_at = time.monotonic()
            self._agents.clear()

free_slots: int property

Slots a spawn could be dispatched into right now.

pending_dispatch: int property

Accepted spawns not yet started.

NOT a status: those agents are submitted like any other, and this is only the scheduler's count of how many are still waiting on a slot.

total_steps: int property

Total tool steps executed across all agents in the session.

__init__(*, max_concurrent: int = 20, session_step_budget: int = 0, attestation: AttestationChain | None = None, queue: AgentQueue | None = None) -> None

Initialize hypervisor with concurrency and budget limits.

Session-wide budget with graduated enforcement.

Parameters:

Name Type Description Default
max_concurrent int

Maximum number of concurrently RUNNING agents.

20
session_step_budget int

Total tool steps allowed across all agents in the session. 0 means unlimited.

0
attestation AttestationChain | None

optional provenance hash chain for this session's agent tree. The hypervisor never constructs one itself (it has no session_id, no store) — the orchestrator injects it, seeded from the store's last persisted record, when the feature is config-enabled. None (the default) means no attestation is recorded; SpawnAgentTool reads this attribute directly and no-ops when it's absent.

None
queue AgentQueue | None

the admission scheduler. Injected so a deployment (or a test) can seat a differently-sized or differently-ordered one without the hypervisor learning how scheduling works; the default is an :class:AgentQueue of max_concurrent slots.

None
Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
def __init__(
    self,
    *,
    max_concurrent: int = 20,
    session_step_budget: int = 0,
    attestation: AttestationChain | None = None,
    queue: AgentQueue | None = None,
) -> None:
    """Initialize hypervisor with concurrency and budget limits.

    Session-wide budget with graduated enforcement.

    Args:
        max_concurrent: Maximum number of concurrently RUNNING agents.
        session_step_budget: Total tool steps allowed across all agents in the
            session. 0 means unlimited.
        attestation: optional provenance hash chain for
            this session's agent tree. The hypervisor never constructs one
            itself (it has no session_id, no store) — the orchestrator
            injects it, seeded from the store's last persisted record, when
            the feature is config-enabled. ``None`` (the default) means no
            attestation is recorded; ``SpawnAgentTool`` reads this
            attribute directly and no-ops when it's absent.
        queue: the admission scheduler. Injected so a deployment (or a
            test) can seat a differently-sized or differently-ordered one
            without the hypervisor learning how scheduling works; the
            default is an :class:`AgentQueue` of ``max_concurrent`` slots.
    """
    self._agents: dict[str, AgentHandle] = {}
    self._lock: asyncio.Lock = asyncio.Lock()
    self._queue: AgentQueue = queue or AgentQueue(capacity=max_concurrent)
    # Set unconditionally so an INJECTED queue is wired too: the queue can
    # detect a failed dispatch but only the hypervisor can settle the
    # handle it stranded.
    self._queue.on_dispatch_failed = self._settle_failed_dispatch
    # Session-wide resource tracking
    self._total_steps: int = 0
    self._session_step_budget: int = session_step_budget
    self.attestation = attestation

accept(unit: SpawnUnit) -> bool async

Admit ONE spawn: True = dispatched now, False = deferred.

Never refuses for capacity — that is the whole point of the queue.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
721
722
723
724
725
726
async def accept(self, unit: SpawnUnit) -> bool:
    """Admit ONE spawn: ``True`` = dispatched now, ``False`` = deferred.

    Never refuses for capacity — that is the whole point of the queue.
    """
    return await self._queue.accept(unit)

accept_batch(units: Sequence[SpawnUnit]) -> list[bool] async

Admit a whole fan-out at once — see :meth:AgentQueue.accept_batch.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
728
729
730
async def accept_batch(self, units: Sequence[SpawnUnit]) -> list[bool]:
    """Admit a whole fan-out at once — see :meth:`AgentQueue.accept_batch`."""
    return await self._queue.accept_batch(units)

agent_step_state(agent_id: str) -> str async

Per-contract step-budget state for one agent.

"ok" for an unregistered agent or one whose contract carries no step bound — the same inert default as a disabled contract.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
948
949
950
951
952
953
954
955
956
957
958
async def agent_step_state(self, agent_id: str) -> str:
    """Per-contract step-budget state for one agent.

    ``"ok"`` for an unregistered agent or one whose contract carries no
    step bound — the same inert default as a disabled contract.
    """
    async with self._lock:
        handle = self._agents.get(agent_id)
    if handle is None:
        return "ok"
    return handle.contract.step_state(handle.steps_completed)

agent_token_state(agent_id: str) -> str async

Per-contract advisory token state for one agent.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
960
961
962
963
964
965
966
async def agent_token_state(self, agent_id: str) -> str:
    """Per-contract advisory token state for one agent."""
    async with self._lock:
        handle = self._agents.get(agent_id)
    if handle is None:
        return "ok"
    return handle.contract.token_state(handle.input_tokens + handle.output_tokens)

budget_exhausted() -> bool

Check if the session step budget is exhausted (0 = unlimited).

Graduated enforcement — this is the trigger check. The response (NL warning injection) happens in ToolUseLoop.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
912
913
914
915
916
917
918
def budget_exhausted(self) -> bool:
    """Check if the session step budget is exhausted (0 = unlimited).

    Graduated enforcement — this is the trigger
    check. The response (NL warning injection) happens in ToolUseLoop.
    """
    return self._session_step_budget > 0 and self._total_steps >= self._session_step_budget

budget_remaining() -> int

Steps remaining in the session budget (0 = unlimited).

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
920
921
922
923
924
def budget_remaining(self) -> int:
    """Steps remaining in the session budget (0 = unlimited)."""
    if self._session_step_budget <= 0:
        return -1  # Unlimited
    return max(0, self._session_step_budget - self._total_steps)

budget_warning(*, headroom: int = 5) -> bool

True when within headroom steps of the session budget (0 = off).

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
926
927
928
929
930
def budget_warning(self, *, headroom: int = 5) -> bool:
    """True when within ``headroom`` steps of the session budget (0 = off)."""
    if self._session_step_budget <= 0:
        return False
    return self.budget_remaining() <= headroom

cancel_agent(agent_id: str) -> str | None async

Cancel an agent, whether it has started or is still waiting for a slot.

Returns None on success, or a diagnostic string on failure.

The waiting case is checked FIRST and it is not an optimization. asyncio_task is populated by the child driver, which runs only AFTER dispatch — so a capacity-deferred agent has no task, the task-cancellation path reported no asyncio task, and the agent then STARTED anyway the moment a sibling freed a slot. Refusing to cancel something and then running it is the worst of both answers, and it is reachable the instant a fan-out exceeds the concurrency limit. Dropping it from the scheduler is what makes the cancel real.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
async def cancel_agent(self, agent_id: str) -> str | None:
    """Cancel an agent, whether it has started or is still waiting for a slot.

    Returns ``None`` on success, or a diagnostic string on failure.

    **The waiting case is checked FIRST and it is not an optimization.**
    ``asyncio_task`` is populated by the child driver, which runs only
    AFTER dispatch — so a capacity-deferred agent has no task, the
    task-cancellation path reported ``no asyncio task``, and the agent then
    STARTED anyway the moment a sibling freed a slot. Refusing to cancel
    something and then running it is the worst of both answers, and it is
    reachable the instant a fan-out exceeds the concurrency limit.
    Dropping it from the scheduler is what makes the cancel real.
    """
    async with self._lock:
        handle = self._agents.get(agent_id)
        if not handle:
            return "agent not in registry"
        if self._queue.discard(agent_id) is not None:
            handle.status = "cancelled"
            handle.error = "cancelled before it was dispatched"
            handle.stopped_at = time.monotonic()
            handle.done_event.set()
            return None
        if not handle.asyncio_task:
            return "no asyncio task"
        if handle.asyncio_task.done():
            return f"task already done (status: {handle.status})"
        handle.asyncio_task.cancel()
        handle.status = "cancelled"
        handle.stopped_at = time.monotonic()
        handle.done_event.set()
        return None

cancel_pending() -> list[str] async

Drop every accepted-but-undispatched spawn, settling each as cancelled.

A waiting agent holds no slot and owns no asyncio.Task, so neither cancel_agent nor the cleanup sweep's task cancellation can reach it — it would sit submitted forever and read as live. Returns the ids dropped. Called before teardown so a manager settling on the way out cannot pump a brand-new child into a run that is already ending.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
async def cancel_pending(self) -> list[str]:
    """Drop every accepted-but-undispatched spawn, settling each as cancelled.

    A waiting agent holds no slot and owns no ``asyncio.Task``, so neither
    ``cancel_agent`` nor the cleanup sweep's task cancellation can reach
    it — it would sit ``submitted`` forever and read as live. Returns the
    ids dropped. Called before teardown so a manager settling on the way
    out cannot pump a brand-new child into a run that is already ending.
    """
    dropped = self._queue.clear()
    if not dropped:
        return []
    now = time.monotonic()
    async with self._lock:
        for spawn in dropped:
            handle = self._agents.get(spawn.agent_id)
            if handle is not None and handle.status in ACTIVE_STATUSES:
                handle.status = "cancelled"
                handle.error = "cancelled while waiting for a concurrency slot"
                handle.stopped_at = now
                handle.done_event.set()
    return [spawn.agent_id for spawn in dropped]

cleanup(timeout: float = 5.0) -> None async

Cancel all agents with graceful escalation.

Requests cancellation on all active asyncio tasks, waits up to timeout seconds for them to finish, then force-marks any still pending as cancelled.

Covers every :data:ACTIVE_STATUSES agent, not just running: one still at submitted never had a loop task to cancel — and one still waiting for a slot never had a task at all — which is exactly why either would otherwise survive shutdown unmarked. Waiting units are dropped from the scheduler FIRST, so a manager settling during the wait window cannot pump a fresh child into a hypervisor that is being torn down.

Marking here is deliberately IN-MEMORY only — the hypervisor holds no event logger, and wiring one in would invert the layering (the session owns the transcript, the hypervisor owns the tree). The durable terminal is written by whichever peer can still reach the log:

  • While the task is still alive (during the request or the wait): cancelling it raises CancelledError inside the agent's own lifecycle handler, which writes the terminal stop before unwinding. That is the normal path.
  • Once the wait times out, at the final force-mark sweep (the task is wedged, or the process is being torn down mid-flight): nothing in-process can still write, so the durable terminal comes from the startup sweep on the NEXT boot, which settles agent spans left open under a run that has since ended.
Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
async def cleanup(self, timeout: float = 5.0) -> None:
    """Cancel all agents with graceful escalation.

    Requests cancellation on all active asyncio tasks, waits up to
    *timeout* seconds for them to finish, then force-marks any still
    pending as cancelled.

    Covers every :data:`ACTIVE_STATUSES` agent, not just ``running``: one
    still at ``submitted`` never had a loop task to cancel — and one still
    waiting for a slot never had a task at all — which is exactly why
    either would otherwise survive shutdown unmarked. Waiting units are
    dropped from the scheduler FIRST, so a manager settling during the
    wait window cannot pump a fresh child into a hypervisor that is being
    torn down.

    Marking here is deliberately IN-MEMORY only — the hypervisor holds no
    event logger, and wiring one in would invert the layering (the session
    owns the transcript, the hypervisor owns the tree). The durable terminal
    is written by whichever peer can still reach the log:

    - While the task is still alive (during the request or the wait):
      cancelling it raises ``CancelledError`` inside the agent's own
      lifecycle handler, which writes the terminal ``stop`` before
      unwinding. That is the normal path.
    - Once the wait times out, at the final force-mark sweep (the task is
      wedged, or the process is being torn down mid-flight): nothing
      in-process can still write, so the durable terminal comes from the
      startup sweep on the NEXT boot, which settles agent spans left open
      under a run that has since ended.
    """
    await self.cancel_pending()

    async with self._lock:
        active = [h for h in self._agents.values() if h.status in ACTIVE_STATUSES]

    if not active:
        async with self._lock:
            self._agents.clear()
        return

    # Request cancellation.
    for handle in active:
        if handle.asyncio_task and not handle.asyncio_task.done():
            handle.asyncio_task.cancel()

    # Wait with timeout.
    tasks = [h.asyncio_task for h in active if h.asyncio_task and not h.asyncio_task.done()]
    if tasks:
        done, pending = await asyncio.wait(
            tasks,
            timeout=timeout,
            return_when=asyncio.ALL_COMPLETED,
        )

        # Force-mark any still-pending as cancelled.
        for task in pending:
            for handle in active:
                if handle.asyncio_task is task:
                    async with self._lock:
                        if handle.status in ACTIVE_STATUSES:
                            handle.status = "cancelled"
                            handle.error = "Force-cancelled after timeout"
                            handle.stopped_at = time.monotonic()

    # Final sweep: mark any remaining non-terminal agents and clear.
    async with self._lock:
        for h in list(self._agents.values()):
            if h.status in ACTIVE_STATUSES:
                h.status = "cancelled"
                h.stopped_at = time.monotonic()
        self._agents.clear()

collect_completed(parent_id: str) -> list[AgentHandle] async

Return children that reached a terminal state with a stored result.

Async CU retrieval — parent reads results when ready, not when child finishes.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
async def collect_completed(self, parent_id: str) -> list[AgentHandle]:
    """Return children that reached a terminal state with a stored result.

    Async CU retrieval — parent reads results
    when ready, not when child finishes.
    """
    terminal = {"completed", "failed", "cancelled"}
    async with self._lock:
        return [
            h
            for h in self._agents.values()
            if h.parent_id == parent_id and h.status in terminal and h.result is not None
        ]

collect_running(parent_id: str) -> list[AgentHandle] async

Return children of parent_id that are still active.

Reads :data:ACTIVE_STATUSES rather than re-listing the states. That matters beyond tidiness: this method is the promise-as-completion gate's ownership index AND what check_agents reports, so a capacity-deferred child — submitted, registered, not yet started — must appear here. A refused unit was registered nowhere, which is exactly why a parent that over-subscribed the fleet was told all its work was done.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
async def collect_running(self, parent_id: str) -> list[AgentHandle]:
    """Return children of *parent_id* that are still active.

    Reads :data:`ACTIVE_STATUSES` rather than re-listing the states. That
    matters beyond tidiness: this method is the promise-as-completion
    gate's ownership index AND what ``check_agents`` reports, so a
    capacity-deferred child — ``submitted``, registered, not yet started —
    must appear here. A refused unit was registered nowhere, which is
    exactly why a parent that over-subscribed the fleet was told all its
    work was done.
    """
    async with self._lock:
        return [
            h
            for h in self._agents.values()
            if h.parent_id == parent_id and h.status in ACTIVE_STATUSES
        ]

get(agent_id: str) -> AgentHandle | None async

Return a single agent handle, or None.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
870
871
872
873
async def get(self, agent_id: str) -> AgentHandle | None:
    """Return a single agent handle, or None."""
    async with self._lock:
        return self._agents.get(agent_id)

list_all() -> list[AgentHandle] async

Full tree view — for CLI/API user visibility.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
893
894
895
896
async def list_all(self) -> list[AgentHandle]:
    """Full tree view — for CLI/API user visibility."""
    async with self._lock:
        return list(self._agents.values())

list_children(parent_id: str) -> list[AgentHandle] async

List direct children of a parent. Enforces isolation.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
875
876
877
878
async def list_children(self, parent_id: str) -> list[AgentHandle]:
    """List direct children of a parent. Enforces isolation."""
    async with self._lock:
        return [h for h in self._agents.values() if h.parent_id == parent_id]

list_descendants(ancestor_id: str) -> list[AgentHandle] async

List all descendants recursively.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
880
881
882
883
884
885
886
887
888
889
890
891
async def list_descendants(self, ancestor_id: str) -> list[AgentHandle]:
    """List all descendants recursively."""
    async with self._lock:
        result: list[AgentHandle] = []
        queue = [ancestor_id]
        while queue:
            pid = queue.pop()
            for h in self._agents.values():
                if h.parent_id == pid:
                    result.append(h)
                    queue.append(h.agent_id)
        return result

list_visible(exclude_agent_id: str | None = None) -> list[AgentHandle] async

Snapshot of agents excluding the caller — used by check_agents.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
898
899
900
901
async def list_visible(self, exclude_agent_id: str | None = None) -> list[AgentHandle]:
    """Snapshot of agents excluding the caller — used by check_agents."""
    async with self._lock:
        return [h for h in self._agents.values() if h.agent_id != exclude_agent_id]

mark_done(agent_id: str, status: AgentStatus, error: str | AgentError | None = None) -> None async

Mark an agent as done with a terminal status.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
851
852
853
854
855
856
857
858
859
860
861
862
863
864
async def mark_done(
    self,
    agent_id: str,
    status: AgentStatus,
    error: str | AgentError | None = None,
) -> None:
    """Mark an agent as done with a terminal status."""
    async with self._lock:
        handle = self._agents.get(agent_id)
        if handle:
            handle.status = status
            handle.error = error
            handle.stopped_at = time.monotonic()
            handle.done_event.set()

mark_tool_start(agent_id: str, tool_id: str | None) -> None async

Stamp (or clear) the tool actually in flight for stall attribution.

update_step only stamps last_tool_id on COMPLETION, so a watchdog check firing mid-call would misattribute the stall to the previous, already-finished tool. Call this at dispatch start with the tool name, and again with None once the call resolves (success, timeout, or exception) so active_tool_id never lingers stale once the agent moves on.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
836
837
838
839
840
841
842
843
844
845
846
847
848
849
async def mark_tool_start(self, agent_id: str, tool_id: str | None) -> None:
    """Stamp (or clear) the tool actually in flight for stall attribution.

    ``update_step`` only stamps ``last_tool_id`` on
    COMPLETION, so a watchdog check firing mid-call would misattribute the
    stall to the previous, already-finished tool. Call this at dispatch
    start with the tool name, and again with ``None`` once the call
    resolves (success, timeout, or exception) so ``active_tool_id`` never
    lingers stale once the agent moves on.
    """
    async with self._lock:
        handle = self._agents.get(agent_id)
        if handle:
            handle.active_tool_id = tool_id

over_wall_deadline_agents(now: float | None = None) -> list[tuple[AgentHandle, str]] async

Running agents whose contract wall-clock deadline is warn/over.

Mirrors :meth:stalled_agents — a per-contract cousin of the stall sweep, keyed on started_at rather than last-progress. now arrives as a clock ARG (never read internally), so a caller can drive this deterministically without sleeping a real clock.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
async def over_wall_deadline_agents(
    self, now: float | None = None
) -> list[tuple[AgentHandle, str]]:
    """Running agents whose contract wall-clock deadline is warn/over.

    Mirrors :meth:`stalled_agents` — a per-contract cousin of the stall
    sweep, keyed on ``started_at`` rather than last-progress. ``now``
    arrives as a clock ARG (never read internally), so a caller can drive
    this deterministically without sleeping a real clock.
    """
    clock = now if now is not None else time.monotonic()
    async with self._lock:
        result: list[tuple[AgentHandle, str]] = []
        for h in self._agents.values():
            if h.status != "running" or h.contract.max_wall_s <= 0:
                continue
            state = h.contract.wall_state(clock - h.started_at)
            if state in ("warn", "over"):
                result.append((h, state))
        return result

record_compaction(agent_id: str) -> None async

Record that an agent compacted its context.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1031
1032
1033
1034
1035
1036
1037
async def record_compaction(self, agent_id: str) -> None:
    """Record that an agent compacted its context."""
    async with self._lock:
        handle = self._agents.get(agent_id)
        if handle:
            handle.compaction_count += 1
            handle.last_compacted_at = time.monotonic()

register(handle: AgentHandle) -> None async

Register a new agent in the hypervisor.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
803
804
805
806
async def register(self, handle: AgentHandle) -> None:
    """Register a new agent in the hypervisor."""
    async with self._lock:
        self._agents[handle.agent_id] = handle

release() -> None async

Settle one running agent's slot, dispatching whatever waits on it.

Async because it is the DISPATCH PUMP, not a bare counter decrement: the freed slot is handed to the next waiting unit, which means starting it. This is deliberately the completion path every settling agent already calls, and exactly-once is already guaranteed there — a pump hung off on_agent_stop instead would fire while the slot is still held and find the queue full every time.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
732
733
734
735
736
737
738
739
740
741
742
async def release(self) -> None:
    """Settle one running agent's slot, dispatching whatever waits on it.

    Async because it is the DISPATCH PUMP, not a bare counter decrement:
    the freed slot is handed to the next waiting unit, which means starting
    it. This is deliberately the completion path every settling agent
    already calls, and exactly-once is already guaranteed there — a pump
    hung off ``on_agent_stop`` instead would fire while the slot is still
    held and find the queue full every time.
    """
    await self._queue.release()

render_agent_tree(*, exclude_agent_id: str | None = None) -> str async

Render a concise text summary of the agent tree for the root's system prompt.

Structural transparency — the root agent (the hypervisor's "brain") gets a live view of all agents so it can reason about the delegation state and intervene if needed.

Parameters:

Name Type Description Default
exclude_agent_id str | None

If provided, omit this agent from the rendered tree. Used so the calling agent does not see itself listed.

None
Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
async def render_agent_tree(
    self,
    *,
    exclude_agent_id: str | None = None,
) -> str:
    """Render a concise text summary of the agent tree for the root's system prompt.

    Structural transparency — the root
    agent (the hypervisor's "brain") gets a live view of all agents so
    it can reason about the delegation state and intervene if needed.

    Args:
        exclude_agent_id: If provided, omit this agent from the rendered
            tree.  Used so the calling agent does not see itself listed.
    """
    async with self._lock:
        if not self._agents:
            return ""
        visible = [h for h in self._agents.values() if h.agent_id != exclude_agent_id]
        if not visible:
            return ""
        registry = get_prompt_registry()
        counts: dict[str, int] = {}
        for h in visible:
            counts[h.status] = counts.get(h.status, 0) + 1
        status_parts = [f"{v} {k}" for k, v in sorted(counts.items())]
        budget_str = ""
        if self._session_step_budget > 0:
            budget_str = registry.render(
                "catalog.agent_tree.budget",
                total_steps=self._total_steps,
                session_step_budget=self._session_step_budget,
            )
        header = registry.render(
            "catalog.agent_tree.header",
            status_parts=", ".join(status_parts),
            budget_str=budget_str,
        )
        lines = [header]
        for h in sorted(visible, key=lambda x: (x.depth, x.agent_id)):
            indent = "  " * h.depth
            step_info = f"{h.steps_completed} steps"
            if h.last_tool_id:
                step_info += registry.render(
                    "catalog.agent_tree.step_info_last", last_tool_id=h.last_tool_id
                )
            # Surface bounded-retry provenance inline so the
            # root sees a child was re-delegated (omitted when never
            # retried, so an ordinary line carries no attempt count).
            if h.attempts > 1:
                step_info += f", {h.attempts} attempts"
            status_marker = ""
            if h.status == "completed":
                status_marker = " -> success"
            elif h.status == "failed":
                status_marker = " -> FAILED"
            elif h.status == "cancelled":
                status_marker = " -> cancelled"
            # Progress/result in tree view
            extra = ""
            if h.result and h.result.summary:
                extra = registry.render(
                    "catalog.agent_tree.result",
                    status=h.result.status,
                    summary=h.result.summary[:120],
                )
            elif h.progress_note:
                extra = registry.render(
                    "catalog.agent_tree.progress",
                    progress_note=h.progress_note[:120],
                )
            compact_marker = (
                registry.render(
                    "catalog.agent_tree.compact", compaction_count=h.compaction_count
                )
                if h.compaction_count
                else ""
            )
            task_preview = h.task_description[:80]
            line = registry.render(
                "catalog.agent_tree.line",
                indent=indent,
                agent_id_head=h.agent_id[:8],
                status=h.status,
                task_preview=task_preview,
                step_info=step_info,
                status_marker=status_marker,
                compact_marker=compact_marker,
                extra=extra,
            )
            # Appended OUTSIDE the templated line (never
            # a catalog.yaml edit) so a contract-less row's bytes are
            # completely unaffected.
            lines.append(line + self._contract_marker(h))
        return "\n".join(lines)

send_message(agent_id: str, message: str) -> str | None async

Send a steering message to a running agent.

Returns None on success, or a diagnostic string on failure.

Adaptive coordination — the hypervisor injects NL feedback (budget warnings, stall nudges) into the agent's message queue. The ToolUseLoop drains this queue between steps. Bidirectional system↔agent feedback loop.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
async def send_message(self, agent_id: str, message: str) -> str | None:
    """Send a steering message to a running agent.

    Returns ``None`` on success, or a diagnostic string on failure.

    Adaptive coordination — the hypervisor
    injects NL feedback (budget warnings, stall nudges) into the agent's
    message queue. The ToolUseLoop drains this queue between steps.
    Bidirectional system↔agent feedback loop.
    """
    async with self._lock:
        handle = self._agents.get(agent_id)
        if not handle:
            return "agent not in registry"
        if handle.status != "running":
            return f"agent status is '{handle.status}'"
        if not handle.message_queue:
            return "no message queue"
        handle.message_queue.put_nowait(message)
        return None

send_to_parent(child_agent_id: str, message: str) -> str | None async

Route a message from a child agent to its parent's queue.

Returns None on success, or a diagnostic string on failure.

Bidirectional message passing — enables lifecycle manager to notify parent on child completion.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
async def send_to_parent(self, child_agent_id: str, message: str) -> str | None:
    """Route a message from a child agent to its parent's queue.

    Returns ``None`` on success, or a diagnostic string on failure.

    Bidirectional message passing —
    enables lifecycle manager to notify parent on child completion.
    """
    async with self._lock:
        child = self._agents.get(child_agent_id)
        if not child or not child.parent_id:
            return "child not in registry or no parent"
        parent = self._agents.get(child.parent_id)
        if not parent:
            return "parent not in registry"
        if not parent.message_queue:
            return "parent has no message queue"
        parent.message_queue.put_nowait(message)
        return None

stalled_agents(threshold: float = 120.0) -> list[AgentHandle] async

Return agents not making progress within threshold seconds.

Internal trigger: delegatee unresponsive → diagnose → evaluate → intervene.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
async def stalled_agents(self, threshold: float = 120.0) -> list[AgentHandle]:
    """Return agents not making progress within threshold seconds.

    Internal trigger: delegatee
    unresponsive → diagnose → evaluate → intervene.
    """
    now = time.monotonic()
    async with self._lock:
        return [
            h
            for h in self._agents.values()
            if h.status == "running"
            and h.last_step_at is not None
            and (now - h.last_step_at) > threshold
        ]

unregister(agent_id: str) -> AgentHandle | None async

Remove an agent from the hypervisor.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
808
809
810
811
812
813
814
815
async def unregister(self, agent_id: str) -> AgentHandle | None:
    """Remove an agent from the hypervisor."""
    async with self._lock:
        handle = self._agents.pop(agent_id, None)
        if handle and handle.status == "running":
            handle.status = "completed"
            handle.stopped_at = time.monotonic()
        return handle

update_step(agent_id: str, tool_id: str) -> None async

Record a completed tool execution step.

Track at tool-call granularity, not agent granularity. Updates last_step_at for stall detection and total_steps for session budget enforcement.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
821
822
823
824
825
826
827
828
829
830
831
832
833
834
async def update_step(self, agent_id: str, tool_id: str) -> None:
    """Record a completed tool execution step.

    Track at tool-call granularity, not agent
    granularity. Updates last_step_at for stall detection and total_steps
    for session budget enforcement.
    """
    async with self._lock:
        handle = self._agents.get(agent_id)
        if handle:
            handle.steps_completed += 1
            handle.last_tool_id = tool_id
            handle.last_step_at = time.monotonic()
            self._total_steps += 1

AgentQueue

Admission scheduler — bounds concurrency WITHOUT dropping work.

Admission is not a two-outcome question. Take-a-slot-or-be-refused makes "the fleet is busy" indistinguishable from "this task was impossible", and silently loses a wide fan-out's surplus. There is a third answer, and it is the one a scheduler owes its caller: accept now, dispatch later.

THE LAW: waiting is free; only RUNNING consumes a slot. A unit sitting in :attr:waiting holds nothing, so a parent blocked on a child can never be part of a capacity cycle — which is what makes deferred admission safe here where it would otherwise deadlock a tree of nested delegations.

Note the vocabulary: units here are waiting or dispatched, never "queued". The agent-facing lifecycle has no such state, and it must not grow one — a waiting unit is submitted like any other accepted agent.

No lock, and none is needed: every capacity decision below is a synchronous read-modify-write on the event-loop thread, with awaits confined to the launch calls that follow them. Bounded concurrency is enforcement, and enforcement that discards work is not graduated, it is a kill.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
class AgentQueue:
    """Admission scheduler — bounds concurrency WITHOUT dropping work.

    Admission is not a two-outcome question. Take-a-slot-or-be-refused makes
    "the fleet is busy" indistinguishable from "this task was impossible", and
    silently loses a wide fan-out's surplus. There is a third answer, and it is
    the one a scheduler owes its caller: **accept now, dispatch later**.

    THE LAW: **waiting is free; only RUNNING consumes a slot.** A unit sitting
    in :attr:`waiting` holds nothing, so a parent blocked on a child can never
    be part of a capacity cycle — which is what makes deferred admission safe
    here where it would otherwise deadlock a tree of nested delegations.

    Note the vocabulary: units here are *waiting* or *dispatched*, never
    "queued". The agent-facing lifecycle has no such state, and it must not
    grow one — a waiting unit is ``submitted`` like any other accepted agent.

    No lock, and none is needed: every capacity decision below is a synchronous
    read-modify-write on the event-loop thread, with awaits confined to the
    launch calls that follow them. Bounded concurrency is enforcement, and
    enforcement that discards work is not graduated, it is a kill.
    """

    def __init__(self, *, capacity: int = 20) -> None:
        """Initialize with the number of agents that may RUN at once."""
        self._capacity = max(1, int(capacity))
        self._running = 0
        self._pending: list[SpawnUnit] = []
        # Injected by the owning hypervisor. A launcher that raises leaves a
        # REGISTERED agent nobody will ever start; scheduling is this class's
        # business and handles are not, so the failure is handed back rather
        # than swallowed. ``None`` degrades to log-and-continue.
        self.on_dispatch_failed: (
            Callable[[ScheduledSpawn, BaseException], Awaitable[None]] | None
        ) = None

    @property
    def capacity(self) -> int:
        """Maximum number of simultaneously RUNNING agents."""
        return self._capacity

    @property
    def running(self) -> int:
        """Units dispatched and not yet settled."""
        return self._running

    @property
    def waiting(self) -> int:
        """Units accepted and waiting for a slot."""
        return len(self._pending)

    @property
    def free_slots(self) -> int:
        """Slots a unit could be dispatched into right now."""
        return max(0, self._capacity - self._running)

    async def accept(self, unit: SpawnUnit) -> bool:
        """Accept ONE unit. ``True`` = dispatched now, ``False`` = deferred.

        Never refuses — see :meth:`accept_batch`, whose single-entry case this
        is (one implementation, so a single spawn and a batch entry can never
        be admitted under different rules).
        """
        return (await self.accept_batch((unit,)))[0]

    async def accept_batch(self, units: Sequence[SpawnUnit]) -> list[bool]:
        """Accept EVERY unit; dispatch what fits, queue the rest.

        Returns one flag per unit, positionally aligned: ``True`` dispatched,
        ``False`` deferred. **Atomic in acceptance, staggered in dispatch** — the
        capacity decisions run as one synchronous pass with no await between
        them, so a settle landing mid-batch can only add dispatches and can
        never split the batch or refuse part of it. Every unit is accepted
        either way; the return value says only which ones started immediately.
        """
        launching: list[SpawnUnit] = []
        for unit in units:
            if self._running >= self._capacity:
                self._pending.append(unit)
                continue
            self._running += 1
            launching.append(unit)

        # Built from what ACTUALLY started, never from what was planned. A
        # launcher that raises must not be reported as running — the caller
        # turns these flags into "N agents started now", and one of them not
        # existing is the same class of lie as the dropped surplus this
        # scheduler replaced. The compensating pump can also promote a unit
        # deferred moments ago, so that unit's flag has to be corrected UP.
        dispatched: set[str] = set()
        for unit in launching:
            if await self._dispatch(unit):
                dispatched.add(unit.spawn.agent_id)
            else:
                promoted = await self._pump()
                if promoted is not None:
                    dispatched.add(promoted)
        return [unit.spawn.agent_id in dispatched for unit in units]

    async def release(self) -> None:
        """Settle one dispatched unit — and pump the queue. THE DISPATCH PUMP."""
        await self._pump()

    async def _pump(self) -> str | None:
        """Hand the slot this queue holds to the next waiting unit.

        Returns the id of whatever started, or ``None`` when the slot was
        simply given back because nothing was waiting.

        The slot is handed DIRECTLY to the next waiting unit rather than
        released and re-acquired: a release/reacquire pair has an await point
        in the middle through which a newly-arriving spawn can overtake a unit
        that has been waiting since before it existed.

        A launcher that raises is logged, its agent settled through
        :attr:`on_dispatch_failed`, and the slot moves on to the next waiting
        unit instead of being stranded — losing it would shrink fleet capacity
        by one, permanently, for the rest of the session.
        """
        while True:
            unit = self._take_next()
            if unit is None:
                self._running = max(0, self._running - 1)
                return None
            if await self._dispatch(unit):
                return unit.spawn.agent_id

    def discard(self, agent_id: str) -> ScheduledSpawn | None:
        """Remove ONE waiting unit, returning its record — ``None`` if not waiting.

        The cancellation seam for a unit that has not been dispatched. Such a
        unit owns no ``asyncio.Task``, so task cancellation cannot touch it and
        it would otherwise start minutes later on a slot a sibling freed —
        after its canceller was told the cancel had failed.
        """
        for index, unit in enumerate(self._pending):
            if unit.spawn.agent_id == agent_id:
                return self._pending.pop(index).spawn
        return None

    def clear(self) -> list[ScheduledSpawn]:
        """Drop every waiting unit, returning the records so a caller settles them.

        A waiting unit holds no slot and owns no task, so nothing else can end
        it: teardown has to reach in here or those agents stay ``submitted``
        forever and read as live to every liveness consumer.
        """
        dropped = [unit.spawn for unit in self._pending]
        self._pending.clear()
        return dropped

    def _take_next(self) -> SpawnUnit | None:
        """Pop the unit with precedence — the order is the RECORD's to define."""
        if not self._pending:
            return None
        self._pending.sort(key=lambda unit: unit.spawn.ordering_key())
        return self._pending.pop(0)

    async def _dispatch(self, unit: SpawnUnit) -> bool:
        """Start one unit on a slot this queue already holds. ``False`` = it did not.

        A raising launcher must not strand the slot AND must not strand the
        AGENT: its handle is already registered, so with nothing else settling
        it the run's own completion gate would wait on it forever. The slot is
        recycled here; the handle is settled by :attr:`on_dispatch_failed`,
        which the hypervisor owns.
        """
        try:
            await unit.launch()
        except Exception as exc:  # noqa: BLE001 — a launcher must never strand a slot
            logging.error("Deferred sub-agent {} failed to start: {}", unit.spawn.agent_id, exc)
            if self.on_dispatch_failed is not None:
                await self.on_dispatch_failed(unit.spawn, exc)
            return False
        logging.debug(
            "Dispatched sub-agent {} after {:.1f}s waiting",
            unit.spawn.agent_id[:8],
            unit.spawn.waited_for(time.monotonic()),
        )
        return True

capacity: int property

Maximum number of simultaneously RUNNING agents.

free_slots: int property

Slots a unit could be dispatched into right now.

running: int property

Units dispatched and not yet settled.

waiting: int property

Units accepted and waiting for a slot.

__init__(*, capacity: int = 20) -> None

Initialize with the number of agents that may RUN at once.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
497
498
499
500
501
502
503
504
505
506
507
508
def __init__(self, *, capacity: int = 20) -> None:
    """Initialize with the number of agents that may RUN at once."""
    self._capacity = max(1, int(capacity))
    self._running = 0
    self._pending: list[SpawnUnit] = []
    # Injected by the owning hypervisor. A launcher that raises leaves a
    # REGISTERED agent nobody will ever start; scheduling is this class's
    # business and handles are not, so the failure is handed back rather
    # than swallowed. ``None`` degrades to log-and-continue.
    self.on_dispatch_failed: (
        Callable[[ScheduledSpawn, BaseException], Awaitable[None]] | None
    ) = None

accept(unit: SpawnUnit) -> bool async

Accept ONE unit. True = dispatched now, False = deferred.

Never refuses — see :meth:accept_batch, whose single-entry case this is (one implementation, so a single spawn and a batch entry can never be admitted under different rules).

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
530
531
532
533
534
535
536
537
async def accept(self, unit: SpawnUnit) -> bool:
    """Accept ONE unit. ``True`` = dispatched now, ``False`` = deferred.

    Never refuses — see :meth:`accept_batch`, whose single-entry case this
    is (one implementation, so a single spawn and a batch entry can never
    be admitted under different rules).
    """
    return (await self.accept_batch((unit,)))[0]

accept_batch(units: Sequence[SpawnUnit]) -> list[bool] async

Accept EVERY unit; dispatch what fits, queue the rest.

Returns one flag per unit, positionally aligned: True dispatched, False deferred. Atomic in acceptance, staggered in dispatch — the capacity decisions run as one synchronous pass with no await between them, so a settle landing mid-batch can only add dispatches and can never split the batch or refuse part of it. Every unit is accepted either way; the return value says only which ones started immediately.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
async def accept_batch(self, units: Sequence[SpawnUnit]) -> list[bool]:
    """Accept EVERY unit; dispatch what fits, queue the rest.

    Returns one flag per unit, positionally aligned: ``True`` dispatched,
    ``False`` deferred. **Atomic in acceptance, staggered in dispatch** — the
    capacity decisions run as one synchronous pass with no await between
    them, so a settle landing mid-batch can only add dispatches and can
    never split the batch or refuse part of it. Every unit is accepted
    either way; the return value says only which ones started immediately.
    """
    launching: list[SpawnUnit] = []
    for unit in units:
        if self._running >= self._capacity:
            self._pending.append(unit)
            continue
        self._running += 1
        launching.append(unit)

    # Built from what ACTUALLY started, never from what was planned. A
    # launcher that raises must not be reported as running — the caller
    # turns these flags into "N agents started now", and one of them not
    # existing is the same class of lie as the dropped surplus this
    # scheduler replaced. The compensating pump can also promote a unit
    # deferred moments ago, so that unit's flag has to be corrected UP.
    dispatched: set[str] = set()
    for unit in launching:
        if await self._dispatch(unit):
            dispatched.add(unit.spawn.agent_id)
        else:
            promoted = await self._pump()
            if promoted is not None:
                dispatched.add(promoted)
    return [unit.spawn.agent_id in dispatched for unit in units]

clear() -> list[ScheduledSpawn]

Drop every waiting unit, returning the records so a caller settles them.

A waiting unit holds no slot and owns no task, so nothing else can end it: teardown has to reach in here or those agents stay submitted forever and read as live to every liveness consumer.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
614
615
616
617
618
619
620
621
622
623
def clear(self) -> list[ScheduledSpawn]:
    """Drop every waiting unit, returning the records so a caller settles them.

    A waiting unit holds no slot and owns no task, so nothing else can end
    it: teardown has to reach in here or those agents stay ``submitted``
    forever and read as live to every liveness consumer.
    """
    dropped = [unit.spawn for unit in self._pending]
    self._pending.clear()
    return dropped

discard(agent_id: str) -> ScheduledSpawn | None

Remove ONE waiting unit, returning its record — None if not waiting.

The cancellation seam for a unit that has not been dispatched. Such a unit owns no asyncio.Task, so task cancellation cannot touch it and it would otherwise start minutes later on a slot a sibling freed — after its canceller was told the cancel had failed.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
601
602
603
604
605
606
607
608
609
610
611
612
def discard(self, agent_id: str) -> ScheduledSpawn | None:
    """Remove ONE waiting unit, returning its record — ``None`` if not waiting.

    The cancellation seam for a unit that has not been dispatched. Such a
    unit owns no ``asyncio.Task``, so task cancellation cannot touch it and
    it would otherwise start minutes later on a slot a sibling freed —
    after its canceller was told the cancel had failed.
    """
    for index, unit in enumerate(self._pending):
        if unit.spawn.agent_id == agent_id:
            return self._pending.pop(index).spawn
    return None

release() -> None async

Settle one dispatched unit — and pump the queue. THE DISPATCH PUMP.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
573
574
575
async def release(self) -> None:
    """Settle one dispatched unit — and pump the queue. THE DISPATCH PUMP."""
    await self._pump()

AgentResult dataclass

Structured result from a sub-agent — the Communication Unit.

Each agent produces a CU that grows with relevant info and drops irrelevant content, preventing context explosion in chains. cannot_solve status enables explicit failure admission as a first-class outcome, saving downstream waste. summary serves as a checkpoint snapshot — even on failure, partial work survives for retry.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
@dataclass
class AgentResult:
    """Structured result from a sub-agent — the Communication Unit.

    Each agent produces a CU that grows with relevant info
    and drops irrelevant content, preventing context explosion in chains.
    ``cannot_solve`` status enables explicit failure
    admission as a first-class outcome, saving downstream waste.
    ``summary`` serves as a checkpoint
    snapshot — even on failure, partial work survives for retry.
    """

    content: str  # Primary output text
    status: AgentResultStatus
    steps_used: int  # Tool steps consumed
    artifacts: list[str] = field(default_factory=list)  # Files touched
    warnings: list[str] = field(default_factory=list)  # Non-fatal issues
    summary: str = ""  # Compressed CU for parent context
    # Honest retry provenance. Total times this task was
    # admitted+run, incl. the first attempt (1 = never retried). Surfaced so the
    # parent sees a workstream was transparently recovered rather than silently
    # dropped. Additive (default 1): a no-retry spawn reports one attempt.
    attempts: int = 1
    # Task-typed CU shape, stamped from
    # the spawn caller's optional `summary_kind`. Additive (default
    # "generic"): a spawn that never declares a kind gets the untyped shape.
    summary_kind: SummaryKind = "generic"
    # The terminal ``AttestationChain`` record's hash for
    # this result, or "" when no chain is wired (the common case today) or
    # the record failed to persist. Additive: a spawn under a session with no
    # attestation configured simply carries an empty hash.
    attestation_hash: str = ""
    # Verifier-gated completion provenance, projected from the child's
    # ``OrchestrationState``. ``verified is None`` (default) = the gate never
    # ran for this child (no spec, master switch off, or a below-execute
    # capability_mode), so nothing was checked;
    # ``True``/``False`` = a ground-truth check passed / exhausted its retries.
    # ``verify_attempts`` is how many verifier runs the child drove. Surfaced so
    # a spawner never reads a claimed completion as an invisible null.
    verified: bool | None = None
    verify_attempts: int = 0

DelegationContract dataclass

Per-spawn delegation bounds a caller can put on ONE child.

Every field's zero/off value leaves the child unbounded, so a spawn that never declares a contract is constrained by nothing here — this is an opt-in ceiling, never a default constraint. Distinct from the hypervisor's SESSION-wide session_step_budget: that is a shared pool across the whole agent tree; this is one caller's bound on one child, checked in addition to (never instead of) the session budget.

Privilege attenuation — autonomy is the hard delegation firebreak (an atomic child can never itself spawn, and the bit only ever narrows down the tree, see AgentContext.child). Graduated enforcement — step_state/ wall_state/token_state each expose a warn tier before the hard stop, mirroring the session budget's own warn-then-halt shape.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
@dataclass(frozen=True)
class DelegationContract:
    """Per-spawn delegation bounds a caller can put on ONE child.

    Every field's zero/off value leaves the child unbounded, so a spawn that
    never declares a ``contract`` is constrained by nothing here — this is an
    opt-in ceiling, never a default constraint. Distinct from the
    hypervisor's SESSION-wide ``session_step_budget``: that is a shared pool
    across the whole agent tree; this is one caller's bound on one child,
    checked in addition to (never instead of) the session budget.

    Privilege attenuation — ``autonomy`` is
    the hard delegation firebreak (an atomic child can never itself spawn,
    and the bit only ever narrows down the tree, see ``AgentContext.child``).
    Graduated enforcement — ``step_state``/
    ``wall_state``/``token_state`` each expose a warn tier before the hard
    stop, mirroring the session budget's own warn-then-halt shape.
    """

    max_steps: int = 0  # 0 = unlimited, layered UNDER the session budget
    max_wall_s: float = 0.0  # 0 = unlimited
    max_tokens: int = 0  # 0 = no signal — ADVISORY, best-effort only
    autonomy: AutonomyTier = "open_ended"
    model_tier: ModelTier | None = None
    step_warn_headroom: int = 3

    @property
    def atomic(self) -> bool:
        """True when this contract strips the child's own delegation rights."""
        return self.autonomy == "atomic"

    @property
    def enabled(self) -> bool:
        """True when the contract carries at least one real constraint."""
        return bool(
            self.max_steps > 0
            or self.max_wall_s > 0
            or self.max_tokens > 0
            or self.atomic
            or self.model_tier is not None
        )

    @classmethod
    def from_value(cls, value: object) -> DelegationContract:
        """Parse + validate the schema ``contract`` object. Unset/invalid → OFF.

        Validation is total (never raises), mirroring ``RetryPolicy.from_value``:
        a malformed field degrades to the safe default rather than failing a
        spawn, since ``contract`` is an optional caller-declared ceiling, not a
        correctness contract. Unknown keys are dropped after ONE warning per
        parse (never per-key) so a typo'd field doesn't spam the log.
        """
        if not isinstance(value, Mapping):
            return cls()

        known = {
            "max_steps",
            "max_wall_s",
            "max_tokens",
            "autonomy",
            "model_tier",
            "step_warn_headroom",
        }
        unknown = sorted(set(value.keys()) - known)
        if unknown:
            logging.warning("DelegationContract.from_value: dropping unknown keys {}", unknown)

        try:
            max_steps = max(0, int(value.get("max_steps", 0)))
        except (TypeError, ValueError):
            max_steps = 0
        try:
            max_wall_s = max(0.0, float(value.get("max_wall_s", 0.0)))
        except (TypeError, ValueError):
            max_wall_s = 0.0
        try:
            max_tokens = max(0, int(value.get("max_tokens", 0)))
        except (TypeError, ValueError):
            max_tokens = 0
        raw_autonomy = value.get("autonomy", "open_ended")
        autonomy = raw_autonomy if raw_autonomy in _AUTONOMY_TIERS else "open_ended"
        raw_model_tier = value.get("model_tier")
        model_tier = raw_model_tier if raw_model_tier in _MODEL_TIERS else None
        try:
            step_warn_headroom = max(0, int(value.get("step_warn_headroom", 3)))
        except (TypeError, ValueError):
            step_warn_headroom = 3

        return cls(
            max_steps=max_steps,
            max_wall_s=max_wall_s,
            max_tokens=max_tokens,
            autonomy=cast("AutonomyTier", autonomy),
            model_tier=cast("ModelTier | None", model_tier),
            step_warn_headroom=step_warn_headroom,
        )

    def step_state(self, steps_completed: int) -> Literal["ok", "warn", "over"]:
        """Graduated step-budget state — ``ok`` when unset (0 = unlimited)."""
        if self.max_steps <= 0:
            return "ok"
        if steps_completed >= self.max_steps:
            return "over"
        if steps_completed >= self.max_steps - self.step_warn_headroom:
            return "warn"
        return "ok"

    def wall_state(self, elapsed_s: float) -> Literal["ok", "warn", "over"]:
        """Graduated wall-clock state — warn at 80%, over at 100% of the bound."""
        if self.max_wall_s <= 0:
            return "ok"
        if elapsed_s >= self.max_wall_s:
            return "over"
        if elapsed_s >= self.max_wall_s * 0.8:
            return "warn"
        return "ok"

    def token_state(self, total_tokens: int) -> Literal["ok", "warn", "over"]:
        """Feature-detected advisory token state — best-effort, never a hard promise.

        ``total_tokens <= 0`` means the caller has no usage signal at all (a
        model/proxy that never surfaced ``usage_metadata``), so absence of
        data reads as ``ok``, never ``over`` — the same feature-detection
        posture as ``_UsageNormalizingLiteLLM``. Likewise a contract with no
        ``max_tokens`` declared is always ``ok``: this axis is advisory only,
        cost accounting is out of scope.
        """
        if total_tokens <= 0 or self.max_tokens <= 0:
            return "ok"
        return "over" if total_tokens >= self.max_tokens else "ok"

    def resolve_model_override(
        self,
        explicit: str | None,
        tier_map: Mapping[str, str],
        allowed: Any = None,
    ) -> str | None:
        """Resolve ``model_tier`` against a deployment's tier→model map.

        Returns ``None`` (fall through to the caller's existing resolution)
        when: an ``explicit`` model was already given (it always wins — this
        is the LOWEST-priority model source); no ``model_tier`` is declared;
        the tier has no map entry; or the mapped model is excluded by
        ``allowed`` (when non-empty). Otherwise returns the mapped model id.
        """
        if explicit:
            return None
        if self.model_tier is None:
            return None
        mapped = tier_map.get(self.model_tier)
        if not mapped:
            return None
        if allowed and mapped not in allowed:
            return None
        return mapped

    def snapshot(self) -> dict[str, Any]:
        """Bounded-scalar serialization.

        The ONE shared shape read by ``check_agents`` and by attestation
        provenance.
        """
        return asdict(self)

atomic: bool property

True when this contract strips the child's own delegation rights.

enabled: bool property

True when the contract carries at least one real constraint.

from_value(value: object) -> DelegationContract classmethod

Parse + validate the schema contract object. Unset/invalid → OFF.

Validation is total (never raises), mirroring RetryPolicy.from_value: a malformed field degrades to the safe default rather than failing a spawn, since contract is an optional caller-declared ceiling, not a correctness contract. Unknown keys are dropped after ONE warning per parse (never per-key) so a typo'd field doesn't spam the log.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
@classmethod
def from_value(cls, value: object) -> DelegationContract:
    """Parse + validate the schema ``contract`` object. Unset/invalid → OFF.

    Validation is total (never raises), mirroring ``RetryPolicy.from_value``:
    a malformed field degrades to the safe default rather than failing a
    spawn, since ``contract`` is an optional caller-declared ceiling, not a
    correctness contract. Unknown keys are dropped after ONE warning per
    parse (never per-key) so a typo'd field doesn't spam the log.
    """
    if not isinstance(value, Mapping):
        return cls()

    known = {
        "max_steps",
        "max_wall_s",
        "max_tokens",
        "autonomy",
        "model_tier",
        "step_warn_headroom",
    }
    unknown = sorted(set(value.keys()) - known)
    if unknown:
        logging.warning("DelegationContract.from_value: dropping unknown keys {}", unknown)

    try:
        max_steps = max(0, int(value.get("max_steps", 0)))
    except (TypeError, ValueError):
        max_steps = 0
    try:
        max_wall_s = max(0.0, float(value.get("max_wall_s", 0.0)))
    except (TypeError, ValueError):
        max_wall_s = 0.0
    try:
        max_tokens = max(0, int(value.get("max_tokens", 0)))
    except (TypeError, ValueError):
        max_tokens = 0
    raw_autonomy = value.get("autonomy", "open_ended")
    autonomy = raw_autonomy if raw_autonomy in _AUTONOMY_TIERS else "open_ended"
    raw_model_tier = value.get("model_tier")
    model_tier = raw_model_tier if raw_model_tier in _MODEL_TIERS else None
    try:
        step_warn_headroom = max(0, int(value.get("step_warn_headroom", 3)))
    except (TypeError, ValueError):
        step_warn_headroom = 3

    return cls(
        max_steps=max_steps,
        max_wall_s=max_wall_s,
        max_tokens=max_tokens,
        autonomy=cast("AutonomyTier", autonomy),
        model_tier=cast("ModelTier | None", model_tier),
        step_warn_headroom=step_warn_headroom,
    )

resolve_model_override(explicit: str | None, tier_map: Mapping[str, str], allowed: Any = None) -> str | None

Resolve model_tier against a deployment's tier→model map.

Returns None (fall through to the caller's existing resolution) when: an explicit model was already given (it always wins — this is the LOWEST-priority model source); no model_tier is declared; the tier has no map entry; or the mapped model is excluded by allowed (when non-empty). Otherwise returns the mapped model id.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
def resolve_model_override(
    self,
    explicit: str | None,
    tier_map: Mapping[str, str],
    allowed: Any = None,
) -> str | None:
    """Resolve ``model_tier`` against a deployment's tier→model map.

    Returns ``None`` (fall through to the caller's existing resolution)
    when: an ``explicit`` model was already given (it always wins — this
    is the LOWEST-priority model source); no ``model_tier`` is declared;
    the tier has no map entry; or the mapped model is excluded by
    ``allowed`` (when non-empty). Otherwise returns the mapped model id.
    """
    if explicit:
        return None
    if self.model_tier is None:
        return None
    mapped = tier_map.get(self.model_tier)
    if not mapped:
        return None
    if allowed and mapped not in allowed:
        return None
    return mapped

snapshot() -> dict[str, Any]

Bounded-scalar serialization.

The ONE shared shape read by check_agents and by attestation provenance.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
333
334
335
336
337
338
339
def snapshot(self) -> dict[str, Any]:
    """Bounded-scalar serialization.

    The ONE shared shape read by ``check_agents`` and by attestation
    provenance.
    """
    return asdict(self)

step_state(steps_completed: int) -> Literal['ok', 'warn', 'over']

Graduated step-budget state — ok when unset (0 = unlimited).

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
274
275
276
277
278
279
280
281
282
def step_state(self, steps_completed: int) -> Literal["ok", "warn", "over"]:
    """Graduated step-budget state — ``ok`` when unset (0 = unlimited)."""
    if self.max_steps <= 0:
        return "ok"
    if steps_completed >= self.max_steps:
        return "over"
    if steps_completed >= self.max_steps - self.step_warn_headroom:
        return "warn"
    return "ok"

token_state(total_tokens: int) -> Literal['ok', 'warn', 'over']

Feature-detected advisory token state — best-effort, never a hard promise.

total_tokens <= 0 means the caller has no usage signal at all (a model/proxy that never surfaced usage_metadata), so absence of data reads as ok, never over — the same feature-detection posture as _UsageNormalizingLiteLLM. Likewise a contract with no max_tokens declared is always ok: this axis is advisory only, cost accounting is out of scope.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
294
295
296
297
298
299
300
301
302
303
304
305
306
def token_state(self, total_tokens: int) -> Literal["ok", "warn", "over"]:
    """Feature-detected advisory token state — best-effort, never a hard promise.

    ``total_tokens <= 0`` means the caller has no usage signal at all (a
    model/proxy that never surfaced ``usage_metadata``), so absence of
    data reads as ``ok``, never ``over`` — the same feature-detection
    posture as ``_UsageNormalizingLiteLLM``. Likewise a contract with no
    ``max_tokens`` declared is always ``ok``: this axis is advisory only,
    cost accounting is out of scope.
    """
    if total_tokens <= 0 or self.max_tokens <= 0:
        return "ok"
    return "over" if total_tokens >= self.max_tokens else "ok"

wall_state(elapsed_s: float) -> Literal['ok', 'warn', 'over']

Graduated wall-clock state — warn at 80%, over at 100% of the bound.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
284
285
286
287
288
289
290
291
292
def wall_state(self, elapsed_s: float) -> Literal["ok", "warn", "over"]:
    """Graduated wall-clock state — warn at 80%, over at 100% of the bound."""
    if self.max_wall_s <= 0:
        return "ok"
    if elapsed_s >= self.max_wall_s:
        return "over"
    if elapsed_s >= self.max_wall_s * 0.8:
        return "warn"
    return "ok"

ScheduledSpawn

Bases: BaseModel

One accepted sub-agent, as the admission scheduler sees it.

The agent_id is minted at ACCEPTANCE, not at dispatch: a unit waiting for a slot is already a real agent, registered and submitted, which is what lets collect_running and check_agents see it and what makes "accepted" a promise the scheduler must keep rather than a hope. A refused unit registered nowhere at all would leave a parent asking check_agents(wait=true) told everything was done — that blindness, not the refusal itself, is what makes a loss silent.

Ordering lives ON the record (:meth:ordering_key) rather than in a comparator beside the queue — a scheduling policy kept apart from the data it orders drifts from it the moment a field is added, the same reason TriggerSpec owns its own due-ness. Clocks arrive as VALUES too: enqueued_at is stamped by the caller that owns one and :meth:waited_for takes now as an argument, so nothing here reads a clock and a test needs no patching to drive it.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
class ScheduledSpawn(BaseModel):
    """One accepted sub-agent, as the admission scheduler sees it.

    The ``agent_id`` is minted at ACCEPTANCE, not at dispatch: a unit waiting
    for a slot is already a real agent, registered and ``submitted``, which is
    what lets ``collect_running`` and ``check_agents`` see it and what makes
    "accepted" a promise the scheduler must keep rather than a hope. A refused
    unit registered nowhere at all would leave a parent asking
    ``check_agents(wait=true)`` told everything was done — that blindness, not
    the refusal itself, is what makes a loss silent.

    Ordering lives ON the record (:meth:`ordering_key`) rather than in a
    comparator beside the queue — a scheduling policy kept apart from the data
    it orders drifts from it the moment a field is added, the same reason
    ``TriggerSpec`` owns its own due-ness. Clocks arrive as VALUES too:
    ``enqueued_at`` is stamped by the caller that owns one and
    :meth:`waited_for` takes ``now`` as an argument, so nothing here reads a
    clock and a test needs no patching to drive it.
    """

    model_config = ConfigDict(extra="forbid", frozen=True)

    agent_id: str = Field(min_length=1)
    # Position in the ``spawn_agents`` array this entry came from (0 for a
    # single spawn) — the tie-break that preserves a fan-out's declared order
    # when several entries are enqueued against the same clock reading.
    batch_index: int = Field(default=0, ge=0)
    priority: SpawnPriority = "normal"
    # Monotonic reading taken by the caller at acceptance.
    enqueued_at: float = Field(ge=0.0)

    def ordering_key(self) -> tuple[int, float, int]:
        """Sort key: priority first, then arrival, then declared batch order."""
        rank = 0 if self.priority == "critical" else 1
        return (rank, self.enqueued_at, self.batch_index)

    def waited_for(self, now: float) -> float:
        """Seconds spent waiting, measured against a caller-supplied clock."""
        return max(0.0, now - self.enqueued_at)

ordering_key() -> tuple[int, float, int]

Sort key: priority first, then arrival, then declared batch order.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
451
452
453
454
def ordering_key(self) -> tuple[int, float, int]:
    """Sort key: priority first, then arrival, then declared batch order."""
    rank = 0 if self.priority == "critical" else 1
    return (rank, self.enqueued_at, self.batch_index)

waited_for(now: float) -> float

Seconds spent waiting, measured against a caller-supplied clock.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
456
457
458
def waited_for(self, now: float) -> float:
    """Seconds spent waiting, measured against a caller-supplied clock."""
    return max(0.0, now - self.enqueued_at)

SpawnUnit dataclass

A scheduled spawn paired with the callable that starts it.

The record is a validated contract; the launcher is a live closure over in-process spawn state that crosses no trust boundary — the same split the package already draws between a persisted spec and a hot RunHandle.

Source code in packages/mewbo_core/src/mewbo_core/agents/hypervisor.py
461
462
463
464
465
466
467
468
469
470
471
@dataclass(frozen=True)
class SpawnUnit:
    """A scheduled spawn paired with the callable that starts it.

    The record is a validated contract; the launcher is a live closure over
    in-process spawn state that crosses no trust boundary — the same split the
    package already draws between a persisted spec and a hot ``RunHandle``.
    """

    spawn: ScheduledSpawn
    launch: Callable[[], Awaitable[None]]

mewbo_core.agents.spawn_agent

Sub-agent spawning tool for the agent hypervisor.

SpawnAgentTool creates a child ToolUseLoop instance, registers it in the AgentHypervisor, runs it to completion, and returns the result. Tool scoping follows the "filter before binding" pattern: denied tools are removed from the child's bind_tools() list so the child LLM never sees them.

AgentError dataclass

Structured error context from a failed sub-agent.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
@dataclass
class AgentError:
    """Structured error context from a failed sub-agent."""

    agent_id: str
    depth: int
    task: str  # First 200 chars of task description
    error: str  # Exception message
    last_tool: str | None = None
    steps_completed: int = 0

    def __str__(self) -> str:  # noqa: D105
        parts = [f"Agent {self.agent_id} (depth={self.depth})"]
        parts.append(f"failed after {self.steps_completed} steps")
        if self.last_tool:
            parts.append(f"at tool '{self.last_tool}'")
        parts.append(f": {self.error}")
        return " ".join(parts)

ChildWorkspace dataclass

The directory ONE child runs in, paired with the rules that govern it.

The two are ONE fact, so they travel as one value: a child pointed at another project must read THAT project's instruction files, and a pair free to drift is how a fleet ends up working in one repository under another's rules — a miss that produces plausible work rather than an error.

Resolved ONCE per spawn, at admission, rather than read off the tool at each attempt: :meth:SpawnAgentTool.rebind_cwd may move the workspace between a child's retry attempts, and a child must finish where it started.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
@dataclass(frozen=True)
class ChildWorkspace:
    """The directory ONE child runs in, paired with the rules that govern it.

    The two are ONE fact, so they travel as one value: a child pointed at
    another project must read THAT project's instruction files, and a pair free
    to drift is how a fleet ends up working in one repository under another's
    rules — a miss that produces plausible work rather than an error.

    Resolved ONCE per spawn, at admission, rather than read off the tool at each
    attempt: :meth:`SpawnAgentTool.rebind_cwd` may move the workspace between a
    child's retry attempts, and a child must finish where it started.
    """

    cwd: str | None
    project_instructions: str | None

RetryPolicy dataclass

Bounded auto-retry / re-delegation policy for a spawned sub-agent.

Opt-in via the retry spawn-schema field; DEFAULT OFF (max == 0) so an unset/absent retry runs the child exactly once. On a retryable terminal failure the spawn bridge re-delegates the same task on the same handle (retaining the one already held semaphore slot for the whole sequence) up to max extra attempts, sleeping :meth:backoff_for with exponential growth between them.

Model-level causes are deliberately NOT re-escalated here: every child ToolUseLoop run already drives the fallback ladder internally, so a fresh attempt gets a fresh ladder — this layer only re-runs a whole child whose loop died. rejected (declined at admission, before the loop) and cancelled (parent-cancelled → CancelledError, re-raised, never retried) are structurally unreachable by the retry loop.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
@dataclass(frozen=True)
class RetryPolicy:
    """Bounded auto-retry / re-delegation policy for a spawned sub-agent.

    Opt-in via the ``retry`` spawn-schema field; **DEFAULT OFF** (``max == 0``)
    so an unset/absent ``retry`` runs the child exactly once. On a *retryable*
    terminal failure the spawn bridge
    re-delegates the **same task on the same handle** (retaining the one already
    held semaphore slot for the whole sequence) up to ``max`` extra attempts,
    sleeping :meth:`backoff_for` with exponential growth between them.

    Model-level causes are deliberately NOT re-escalated here: every child
    ``ToolUseLoop`` run already drives the fallback ladder internally, so a
    fresh attempt gets a fresh ladder — this layer only re-runs a whole child
    whose loop died. ``rejected`` (declined at admission, before the loop) and
    ``cancelled`` (parent-cancelled → ``CancelledError``, re-raised, never
    retried) are structurally unreachable by the retry loop.
    """

    max: int = 0
    on: tuple[str, ...] = ("timeout", "failed")
    backoff: float = 1.0

    # The coarse retry-cause vocabulary. ``failed`` is the catch-all transient
    # terminal failure; ``timeout`` is the timeout-flavoured subset (mapped via
    # the shared classifier so this layer never re-derives provider semantics).
    _CAUSES: frozenset[str] = frozenset({"timeout", "failed"})

    @classmethod
    def from_value(cls, value: object) -> RetryPolicy:
        """Parse + validate the schema ``retry`` object. Unset/invalid → OFF.

        Validation is total (never raises): a malformed field degrades to the
        safe default rather than failing a spawn, since ``retry`` is an optional
        resilience hint, not a correctness contract.
        """
        if not isinstance(value, Mapping):
            return cls()
        raw_max: Any = value.get("max", 0)
        try:
            max_retries = max(0, int(raw_max))
        except (TypeError, ValueError):
            max_retries = 0
        on_val = value.get("on")
        if isinstance(on_val, (list, tuple)):
            on = tuple(str(x) for x in on_val if str(x) in cls._CAUSES)
        else:
            on = ("timeout", "failed")
        if not on:  # an explicit-but-empty/invalid list falls back to both
            on = ("timeout", "failed")
        raw_backoff: Any = value.get("backoff", 1.0)
        try:
            backoff = max(0.0, float(raw_backoff))
        except (TypeError, ValueError):
            backoff = 1.0
        return cls(max=max_retries, on=on, backoff=backoff)

    @property
    def enabled(self) -> bool:
        """True when at least one retry is permitted."""
        return self.max > 0

    def should_retry(self, cause: str, attempt: int) -> bool:
        """True when another attempt is allowed for this failure ``cause``.

        ``attempt`` is the number of attempts made so far (the one that just
        failed). Total attempts are bounded at ``max + 1``.
        """
        return attempt <= self.max and cause in self.on

    def backoff_for(self, attempt: int) -> float:
        """Exponential backoff (seconds) before the next attempt.

        ``attempt`` is the failed attempt's index (1-based), so the first retry
        waits ``backoff``, the second ``2 * backoff``, etc.
        """
        return self.backoff * (2 ** (max(1, attempt) - 1))

    @staticmethod
    def classify_cause(exc: BaseException) -> str:
        """Map a child-loop exception to a coarse retry cause.

        Reuses the ``RetryStrategy`` classifier's reason taxonomy (DRY — the
        delegation layer never re-derives provider/timeout semantics): a
        timeout/deadline-flavoured failure is ``"timeout"``; everything else
        (transient or otherwise) collapses to the generic ``"failed"``.
        """
        from mewbo_core.llm.llm_resilience import LlmResilienceExhausted, RetryStrategy

        reason = ""
        inner: BaseException = exc
        if isinstance(exc, LlmResilienceExhausted):
            reason = exc.reason or ""
            inner = exc.last_error or exc
        if reason not in ("timeout", "deadline"):
            reason = RetryStrategy.classify(inner).reason
        return "timeout" if reason in ("timeout", "deadline") else "failed"

enabled: bool property

True when at least one retry is permitted.

backoff_for(attempt: int) -> float

Exponential backoff (seconds) before the next attempt.

attempt is the failed attempt's index (1-based), so the first retry waits backoff, the second 2 * backoff, etc.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
393
394
395
396
397
398
399
def backoff_for(self, attempt: int) -> float:
    """Exponential backoff (seconds) before the next attempt.

    ``attempt`` is the failed attempt's index (1-based), so the first retry
    waits ``backoff``, the second ``2 * backoff``, etc.
    """
    return self.backoff * (2 ** (max(1, attempt) - 1))

classify_cause(exc: BaseException) -> str staticmethod

Map a child-loop exception to a coarse retry cause.

Reuses the RetryStrategy classifier's reason taxonomy (DRY — the delegation layer never re-derives provider/timeout semantics): a timeout/deadline-flavoured failure is "timeout"; everything else (transient or otherwise) collapses to the generic "failed".

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
@staticmethod
def classify_cause(exc: BaseException) -> str:
    """Map a child-loop exception to a coarse retry cause.

    Reuses the ``RetryStrategy`` classifier's reason taxonomy (DRY — the
    delegation layer never re-derives provider/timeout semantics): a
    timeout/deadline-flavoured failure is ``"timeout"``; everything else
    (transient or otherwise) collapses to the generic ``"failed"``.
    """
    from mewbo_core.llm.llm_resilience import LlmResilienceExhausted, RetryStrategy

    reason = ""
    inner: BaseException = exc
    if isinstance(exc, LlmResilienceExhausted):
        reason = exc.reason or ""
        inner = exc.last_error or exc
    if reason not in ("timeout", "deadline"):
        reason = RetryStrategy.classify(inner).reason
    return "timeout" if reason in ("timeout", "deadline") else "failed"

from_value(value: object) -> RetryPolicy classmethod

Parse + validate the schema retry object. Unset/invalid → OFF.

Validation is total (never raises): a malformed field degrades to the safe default rather than failing a spawn, since retry is an optional resilience hint, not a correctness contract.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
@classmethod
def from_value(cls, value: object) -> RetryPolicy:
    """Parse + validate the schema ``retry`` object. Unset/invalid → OFF.

    Validation is total (never raises): a malformed field degrades to the
    safe default rather than failing a spawn, since ``retry`` is an optional
    resilience hint, not a correctness contract.
    """
    if not isinstance(value, Mapping):
        return cls()
    raw_max: Any = value.get("max", 0)
    try:
        max_retries = max(0, int(raw_max))
    except (TypeError, ValueError):
        max_retries = 0
    on_val = value.get("on")
    if isinstance(on_val, (list, tuple)):
        on = tuple(str(x) for x in on_val if str(x) in cls._CAUSES)
    else:
        on = ("timeout", "failed")
    if not on:  # an explicit-but-empty/invalid list falls back to both
        on = ("timeout", "failed")
    raw_backoff: Any = value.get("backoff", 1.0)
    try:
        backoff = max(0.0, float(raw_backoff))
    except (TypeError, ValueError):
        backoff = 1.0
    return cls(max=max_retries, on=on, backoff=backoff)

should_retry(cause: str, attempt: int) -> bool

True when another attempt is allowed for this failure cause.

attempt is the number of attempts made so far (the one that just failed). Total attempts are bounded at max + 1.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
385
386
387
388
389
390
391
def should_retry(self, cause: str, attempt: int) -> bool:
    """True when another attempt is allowed for this failure ``cause``.

    ``attempt`` is the number of attempts made so far (the one that just
    failed). Total attempts are bounded at ``max + 1``.
    """
    return attempt <= self.max and cause in self.on

SpawnAgentTask

Bases: BaseModel

One entry in a spawn_agents batch.

Carries the SAME per-task fields as the single spawn_agent schema, but validated at definition: extra="forbid" rejects stray keys so a malformed fan-out fails fast instead of silently dropping a field, and a blank task is refused (an empty delegation is never intentional).

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
class SpawnAgentTask(BaseModel):
    """One entry in a ``spawn_agents`` batch.

    Carries the SAME per-task fields as the single ``spawn_agent`` schema, but
    validated at definition: ``extra="forbid"`` rejects stray keys so a
    malformed fan-out fails fast instead of silently dropping a field, and a
    blank ``task`` is refused (an empty delegation is never intentional).
    """

    model_config = ConfigDict(extra="forbid")

    task: str = Field(min_length=1)
    model: str | None = None
    allowed_tools: list[str] | None = None
    denied_tools: list[str] | None = None
    # Deprecated — retained for schema/prompt compatibility, never enforced.
    max_steps: int | None = None
    acceptance_criteria: str | None = None
    agent_type: str | None = None
    # The workspace THIS child runs in, as a project key the catalog lists.
    # Absent (the default) inherits the parent's directory. Because the batch
    # and the single
    # spawn share one core, a single ``spawn_agents`` call can point each entry
    # at a different project — one fleet per repository under one hypervisor.
    # A key that resolves to nothing REFUSES the spawn (see
    # ``SpawnAgentTool._child_workspace``); it never degrades to the parent's
    # directory.
    project: str | None = Field(default=None, min_length=1)
    # Opt-in bounded auto-retry, parsed downstream by ``RetryPolicy``.
    # A batch entry can carry it just like a single spawn — one transient
    # failure in a wide fan-out then re-delegates instead of dropping a lane.
    retry: dict[str, Any] | None = None
    # Opt-in per-agent delegation bounds, parsed downstream by
    # ``DelegationContract``. A batch entry gets the same LAYERED-UNDER-the-
    # session-budget ceiling as a single spawn.
    contract: dict[str, Any] | None = None
    # Opt-in ground-truth completion check, parsed downstream by
    # ``CommandVerification.from_value``. When the master switch is on AND this
    # child can act (capability_mode ∈ {execute, all}), its claimed completion
    # is gated behind this command passing. Default off / inactive leaves the
    # completion ungated; a supplied-but-inactive spec is surfaced in the
    # spawn response + event, never silently dropped.
    verification: dict[str, Any] | None = None
    # Task-typed Communication Unit shape for
    # this sub-agent's final summary. ``None`` (default) is the untyped path —
    # ``_spawn_one`` stamps ``AgentResult.summary_kind`` with its own "generic"
    # default in that case, so an unset field changes nothing about a spawn's
    # behaviour or output.
    summary_kind: SummaryKind | None = None
    # Coarse delegation privilege ceiling — privilege attenuation.
    # A pre-filter LAYERED UNDER
    # ``allowed_tools``/``denied_tools`` — it can only remove more tools, never
    # add. Gates BOTH surfaces (the two-surface law): file/registry tools
    # via ``filter_specs`` AND per-agent session action tools via
    # ``SessionToolRegistry.build_for`` (session tools default to tier
    # ``execute``, so ``read_only`` admits none unless declared ``read``).
    # ``"all"`` (default) = no capability filtering.
    # Narrowed monotonically against the parent's effective
    # mode at spawn time, so a child can only ever restrict further. See
    # ``CapabilityMode`` for the tier law.
    capability_mode: CapabilityMode = "all"
    # Filesystem-containment ceiling — the SECOND privilege axis,
    # orthogonal to ``capability_mode``. Defaults to ``workspace_write`` (the
    # sensible sub-agent default: reads + writes confined to the workspace), yet
    # because narrowing is min-wins against the parent's own tier AND the ROOT
    # default is ``full_access``, a child spawned off a default root resolves to
    # ``workspace_write`` — so THIS seam, not the root's tier, is where the
    # filesystem firebreak is actually drawn. Narrowed monotonically, so a child
    # can only ever restrict further. See ``WorkspaceContainment`` for the tier
    # law and ``agent.workspace_enforcement`` for the switch that arms it.
    workspace_mode: WorkspaceMode = "workspace_write"
    # Delegation approval policy — PARSED + validated here but
    # ENFORCED by nothing in v1 (the approval-gate wiring is a later wave). Kept
    # on the schema now so the wire contract is forward-stable; a spawn that sets
    # it today behaves exactly as ``on_failure`` (i.e. no gate).
    approval_policy: ApprovalPolicy = "on_failure"

    # The ONE mapping from a declared
    # ``summary_kind`` to the one-line directive appended to the child's task
    # text. A `Literal`-keyed class constant, not a per-call `if kind ==`
    # chain: ``_spawn_one`` does a plain dict lookup. ``"generic"`` has no
    # entry — it is the untyped default and appends nothing.
    SUMMARY_KIND_DIRECTIVES: ClassVar[dict[SummaryKind, str]] = {
        "evidence": (
            "Your final summary must be an evidence package: the "
            "facts/quotes/paths the parent needs, not narrative."
        ),
        "running_summary": (
            "Your final summary must be a running summary: the task's "
            "cumulative state so far, written so a fresh reader needs no "
            "prior turns to pick it up."
        ),
        "code_signature": (
            "Your final summary must be a function/class-signature catalog: "
            "the names, signatures, and one-line purpose of every "
            "function/class you touched or introduced."
        ),
    }

    def to_args(self) -> dict[str, Any]:
        """Project to the ``args`` dict the single-spawn path consumes.

        Unset (``None``) fields are dropped so the downstream ``args.get(...)``
        defaults apply exactly as they do for an ad-hoc ``spawn_agent`` call —
        keeping the batch a thin reuse of ``_spawn_one`` rather than a fork.
        """
        return {k: v for k, v in self.model_dump().items() if v is not None}

to_args() -> dict[str, Any]

Project to the args dict the single-spawn path consumes.

Unset (None) fields are dropped so the downstream args.get(...) defaults apply exactly as they do for an ad-hoc spawn_agent call — keeping the batch a thin reuse of _spawn_one rather than a fork.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
189
190
191
192
193
194
195
196
def to_args(self) -> dict[str, Any]:
    """Project to the ``args`` dict the single-spawn path consumes.

    Unset (``None``) fields are dropped so the downstream ``args.get(...)``
    defaults apply exactly as they do for an ad-hoc ``spawn_agent`` call —
    keeping the batch a thin reuse of ``_spawn_one`` rather than a fork.
    """
    return {k: v for k, v in self.model_dump().items() if v is not None}

SpawnAgentTool

Spawns a child ToolUseLoop as a sub-agent.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
2028
2029
2030
2031
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
2100
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
2115
2116
2117
2118
2119
2120
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
2143
2144
2145
2146
2147
2148
2149
2150
2151
2152
2153
2154
2155
2156
2157
2158
2159
2160
2161
2162
2163
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
2178
2179
2180
2181
2182
2183
2184
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
2204
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
2226
2227
2228
2229
2230
2231
2232
2233
2234
2235
2236
2237
2238
2239
2240
2241
2242
2243
2244
2245
2246
2247
2248
2249
2250
2251
2252
2253
2254
2255
2256
2257
2258
2259
2260
2261
2262
2263
2264
2265
2266
2267
2268
2269
2270
2271
2272
2273
2274
2275
2276
2277
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
2291
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
2308
2309
2310
2311
2312
2313
2314
2315
2316
2317
2318
2319
2320
2321
2322
2323
2324
2325
2326
2327
2328
2329
2330
2331
2332
2333
2334
2335
2336
2337
2338
2339
2340
2341
2342
2343
2344
2345
2346
2347
2348
2349
2350
2351
2352
2353
2354
2355
2356
2357
2358
2359
2360
2361
2362
2363
2364
2365
2366
2367
2368
2369
2370
2371
2372
2373
2374
2375
2376
2377
2378
2379
2380
2381
2382
2383
2384
2385
2386
2387
2388
2389
2390
2391
2392
2393
2394
2395
2396
2397
2398
2399
2400
2401
2402
2403
2404
2405
2406
2407
2408
2409
2410
2411
2412
2413
2414
2415
2416
2417
2418
2419
2420
2421
2422
2423
2424
2425
2426
2427
2428
2429
2430
2431
2432
2433
2434
2435
2436
2437
2438
2439
2440
2441
2442
2443
2444
2445
2446
2447
2448
2449
2450
2451
2452
2453
2454
2455
2456
2457
2458
2459
2460
2461
2462
2463
2464
2465
2466
2467
2468
2469
2470
2471
2472
2473
2474
2475
2476
2477
2478
2479
2480
2481
2482
2483
2484
2485
2486
2487
2488
2489
2490
2491
2492
2493
2494
2495
2496
2497
2498
2499
2500
2501
2502
2503
2504
2505
2506
2507
2508
2509
2510
2511
2512
2513
2514
2515
2516
2517
2518
2519
2520
2521
2522
2523
2524
2525
2526
2527
2528
2529
2530
2531
2532
2533
2534
2535
2536
2537
2538
2539
2540
2541
2542
2543
2544
2545
2546
2547
2548
2549
2550
2551
2552
2553
2554
2555
2556
2557
2558
2559
2560
2561
2562
2563
2564
2565
2566
2567
2568
2569
2570
2571
2572
2573
2574
2575
2576
2577
2578
2579
2580
2581
2582
2583
2584
2585
2586
2587
2588
2589
2590
2591
2592
2593
2594
2595
2596
2597
2598
2599
2600
2601
2602
2603
2604
2605
2606
2607
2608
2609
2610
2611
2612
2613
2614
2615
2616
2617
2618
2619
2620
2621
2622
2623
2624
2625
2626
2627
2628
2629
2630
2631
2632
2633
2634
2635
2636
2637
2638
2639
2640
2641
2642
2643
2644
2645
2646
2647
2648
2649
2650
2651
2652
2653
2654
2655
2656
2657
2658
2659
2660
2661
2662
2663
2664
2665
2666
2667
2668
2669
2670
2671
2672
2673
2674
2675
2676
2677
2678
2679
2680
2681
2682
2683
2684
2685
2686
2687
2688
2689
2690
2691
2692
2693
2694
2695
2696
2697
2698
2699
2700
2701
2702
2703
2704
2705
2706
2707
2708
2709
2710
2711
2712
2713
2714
2715
2716
2717
2718
2719
2720
2721
2722
2723
2724
2725
2726
2727
2728
2729
2730
2731
2732
2733
2734
2735
2736
2737
2738
2739
2740
2741
2742
2743
2744
2745
2746
2747
2748
2749
2750
2751
2752
2753
2754
2755
2756
2757
2758
2759
2760
2761
2762
2763
2764
2765
2766
2767
2768
2769
2770
2771
2772
2773
2774
2775
2776
2777
2778
2779
2780
2781
2782
2783
2784
2785
2786
2787
2788
2789
2790
2791
2792
2793
2794
2795
2796
2797
2798
2799
2800
2801
2802
2803
2804
2805
2806
2807
2808
2809
2810
2811
2812
2813
2814
2815
2816
2817
2818
2819
2820
2821
2822
2823
2824
2825
2826
2827
2828
2829
2830
2831
2832
2833
2834
2835
2836
2837
2838
2839
2840
2841
2842
2843
2844
2845
2846
2847
2848
2849
2850
2851
2852
2853
2854
2855
2856
2857
2858
2859
2860
2861
2862
2863
2864
2865
2866
2867
2868
2869
2870
2871
2872
2873
2874
2875
2876
2877
2878
2879
2880
2881
2882
2883
2884
2885
2886
2887
2888
2889
2890
2891
2892
2893
2894
2895
2896
2897
2898
2899
2900
2901
2902
2903
2904
2905
2906
2907
2908
2909
2910
2911
2912
2913
2914
2915
2916
class SpawnAgentTool:
    """Spawns a child ToolUseLoop as a sub-agent."""

    def __init__(
        self,
        *,
        agent_context: AgentContext,
        tool_registry: ToolRegistry,
        permission_policy: PermissionPolicy,
        approval_callback: Callable[[ActionStep], bool] | None = None,
        hook_manager: HookManager,
        safety_plane: SafetyPlane | None = None,
        project_instructions: str | None = None,
        user_instructions: str | None = None,
        cwd: str | None = None,
        agent_registry: Any = None,
        session_tool_registry: SessionToolRegistry | None = None,
        session_capabilities: tuple[str, ...] = (),
        enable_skills: bool = True,
        catalog: ProjectCatalog | None = None,
    ) -> None:
        """Initialize with parent context and shared registries."""
        self._agent_context = agent_context
        self._tool_registry = tool_registry
        self._permission_policy = permission_policy
        self._approval_callback = approval_callback
        self._hook_manager = hook_manager
        self._safety_plane = safety_plane
        self._project_instructions = project_instructions
        # Operator-authored custom instructions are inherited by every child,
        # exactly like the project instructions — they describe the deployment,
        # not one agent's task, so a sub-agent that lost them would be running
        # under different rules than its parent.
        self._user_instructions = user_instructions
        self._cwd = cwd
        # The catalog a per-spawn ``project`` key is resolved through. Injected
        # rather than reached for: a CLI drive has no app-side stores to build
        # one from, and ``None`` there must refuse a ``project`` outright rather
        # than resolve it to something plausible.
        self._catalog = catalog
        self._agent_registry = agent_registry
        self._session_tool_registry = session_tool_registry
        self._session_capabilities = session_capabilities
        # Children inherit the parent drive's skill policy: a headless search
        # run disables auto-skill injection for the ROOT *and* every probe it
        # spawns (the audit found every server-side agent burning step 1 on
        # ``activate_skill``).
        self._enable_skills = enable_skills
        # Plan-mode context — set by ToolUseLoop.run() so children
        # inherit the session's plan path and mode.
        self.session_id: str | None = None
        self.parent_mode: str = "act"
        # The parent's EFFECTIVE spec set, stamped by ToolUseLoop.run() alongside
        # the plan context (both are run()-time state — the loop only learns its
        # own specs when it is handed them, long after this tool is constructed).
        # Containment must be monotone: a child narrows, never widens. Deriving a
        # child from the raw registry instead would hand it tools the parent never
        # held, since a parent's set can be narrowed by scoping this tool cannot
        # reconstruct (a console session's ``allowed_tools`` over MCP tools, an
        # ancestor's own filtering). ``None`` means unstamped — fall back to the
        # registry so a directly-constructed tool keeps the full set.
        self.parent_tool_specs: list[ToolSpec] | None = None
        # Track lifecycle manager tasks for deterministic cleanup.
        self._lifecycle_tasks: list[asyncio.Task[None]] = []

    def rebind_active_model(self, model_name: str) -> None:
        """Re-seat this tool's parent context onto an escalated model.

        The public seam for the loop's sticky model escalation. A parent that
        healed itself onto a rescue model must not keep spawning children onto
        the dead one: :meth:`_resolve_model` falls back to the parent context's
        ``model_name``, and each child's own context is derived from it, so a
        stale value here re-infects the whole subtree.

        ``AgentContext`` is frozen, so this REPLACES rather than mutates — and
        that is exactly why the seam is a method and not an attribute write.
        Knowing the context is a frozen dataclass is this class's business, not
        the loop's; reaching in to do the ``replace`` from outside couples the
        caller to a representation it should never have to know. Idempotent, so
        the loop may call it on every turn.
        """
        if not model_name or model_name == self._agent_context.model_name:
            return
        self._agent_context = replace(self._agent_context, model_name=model_name)

    def rebind_cwd(self, cwd: str, *, project_instructions: str | None) -> None:
        """Re-point the workspace every FUTURE child inherits.

        The seam a session-level project switch drives: from here on, a spawn
        that names no ``project`` of its own lands in *cwd* and reads
        *project_instructions* instead of the ones this tool was built with.

        Already-running children are deliberately untouched. A live child's loop
        captured its own directory when it was built and has been resolving
        paths against it ever since; moving that out from under it would be a
        race no one owns — its containment root, its verifier's working
        directory and its half-finished edits would disagree about where it is.
        A child finishes where it started; the switch applies to the next one.
        That is also why a spawn resolves its :class:`ChildWorkspace` once at
        admission rather than re-reading these fields on each retry attempt.
        """
        self._cwd = cwd
        self._project_instructions = project_instructions

    def _child_workspace(self, project: object) -> ChildWorkspace:
        """Resolve the workspace ONE child runs in, or refuse to spawn it.

        An absent ``project`` inherits this agent's own directory and
        instructions — the overwhelmingly common case.

        A named project is resolved through the injected catalog, and a failure
        REFUSES rather than falling back: a fleet quietly working in the wrong
        repository produces plausible-looking work no one thinks to check, which
        is a far worse outcome than a spawn that never started. The caller named
        a key deliberately, so the parent's directory is not a safe substitute
        for it.

        Raises:
            ProjectResolutionError: the key names no runnable directory, or
                there is no catalog to resolve it with.
        """
        if not isinstance(project, str) or not project.strip():
            return ChildWorkspace(
                cwd=self._cwd, project_instructions=self._project_instructions
            )
        key = project.strip()
        if self._catalog is None:
            raise ProjectResolutionError(
                "no_catalog",
                f"Project '{key}' cannot be resolved: this session has no project "
                "catalog. Omit 'project' to run the sub-agent in this agent's own "
                "directory.",
            )
        cwd = self._catalog.resolve(key)
        # A child in another repository reads THAT repository's rules. Handing
        # it the parent's is the same class of miss as handing it the parent's
        # directory, and the quieter of the two.
        return ChildWorkspace(
            cwd=cwd, project_instructions=discover_project_instructions(cwd)
        )

    async def run_async(self, action_step: ActionStep) -> MockSpeaker:
        """Execute a single sub-agent. Returns the result as a MockSpeaker.

        Thin wrapper over :meth:`_spawn_one` — the batch path
        (:meth:`run_batch_async`) shares the same core. The outcome is
        projected through :meth:`_SpawnOutcome.report`, not read off
        ``content``: reading the string would throw away the refusal's status
        and typed cause, handing the model a bare sentence for five different
        failures.
        """
        args = (
            action_step.tool_input
            if isinstance(action_step.tool_input, dict)
            else {"task": str(action_step.tool_input)}
        )
        outcome = await self._spawn_one(args, blocking_admit=True)
        return MockSpeaker(content=outcome.report())

    async def run_batch_async(self, action_step: ActionStep) -> MockSpeaker:
        """Fan out a batch of independent sub-agents from ONE tool call.

        Pure composition over :meth:`_spawn_one`, in two phases. Each entry is
        RESOLVED first (its workspace, agent type and model settled, its handle
        registered, its id minted), then the whole batch is handed to the
        scheduler in ONE atomic acceptance: every entry that resolved is
        accepted, and the scheduler decides only which start now and which
        wait. A batch is therefore throttled by concurrency, never truncated by
        it. Marking the surplus ``rejected`` and discarding it would read to
        the caller as N agents when only ``max_concurrent`` existed.

        Returns the ordered ``agent_id``s; the model collects results via the
        existing ``check_agents``. The orchestration loop is untouched.
        """
        raw = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
        raw_tasks = raw.get("tasks")
        if not isinstance(raw_tasks, list) or not raw_tasks:
            return MockSpeaker(
                content="ERROR: spawn_agents requires a non-empty 'tasks' array."
            )
        try:
            tasks = [SpawnAgentTask.model_validate(entry) for entry in raw_tasks]
        except ValidationError as exc:
            return MockSpeaker(content=f"ERROR: invalid spawn_agents task: {exc}")

        agents: list[dict[str, Any]] = []
        agent_ids: list[str | None] = []
        units: list[SpawnUnit] = []
        accepted = 0
        for idx, task in enumerate(tasks):
            outcome = await self._spawn_one(
                task.to_args(), blocking_admit=False, batch=units, batch_index=idx
            )
            agent_ids.append(outcome.agent_id)
            if outcome.agent_id is not None:
                accepted += 1
            entry: dict[str, Any] = {
                "index": idx,
                "agent_id": outcome.agent_id,
                "status": outcome.status,
                "task": task.task[:200],
            }
            if outcome.agent_id is None:
                # ADDITIVE, and load-bearing: a refused slot is otherwise the
                # only place a refusal's cause is discarded. ``code`` is the
                # machine-readable half — "no free concurrency slot" would be a
                # wrong answer that sends the caller retrying a project key
                # that will never resolve.
                entry["reason"] = outcome.content
                if outcome.code is not None:
                    entry["code"] = outcome.code
            agents.append(entry)

        # ONE atomic acceptance for the whole fan-out. Nothing here can refuse;
        # the flags say only which entries got a slot immediately. Every
        # accepted entry stays ``submitted`` either way — starting now versus
        # waiting is a SCHEDULING fact, reported as a count, never as a
        # per-agent status the whole client tree would have to learn.
        decisions = await self._agent_context.registry.accept_batch(units)
        started = sum(1 for dispatched in decisions if dispatched)

        rejected = len(tasks) - accepted
        deferred = len(units) - started
        summary = f"Accepted {accepted}/{len(tasks)} agent(s)"
        if units:
            summary += (
                f" — {started} started now, {deferred} waiting for a free slot "
                "(they start automatically)"
            )
        if rejected:
            summary += f"; {rejected} refused — see each entry's 'code' and 'reason'"
        summary += ". Use check_agents to monitor progress and collect results."
        return MockSpeaker(
            content=json.dumps(
                {
                    "kind": "agent_batch",
                    "text": summary,
                    "agents": agents,
                    "agent_ids": agent_ids,
                    # ``accepted`` is the name the count now deserves and the
                    # one consumers read FIRST; ``spawned`` is retained at the
                    # same value so an older reader is unaffected and so the
                    # consumer-side ``accepted ?? spawned`` fallback keeps
                    # serving events already in the store. Emitting only
                    # ``spawned`` would leave every consumer permanently on its
                    # fallback branch with the preferred key dead.
                    "accepted": accepted,
                    "spawned": accepted,
                    "dispatched": started,
                    "deferred": deferred,
                    "rejected": rejected,
                }
            )
        )

    def _mark_dispatched(
        self,
        child_ctx: AgentContext,
        handle: AgentHandle,
        *,
        task_desc: str,
        agent_type: str | None,
        model: str,
        contract: DelegationContract,
        verification_note: str | None,
        model_fallback_note: str | None,
    ) -> None:
        """Flip an ACCEPTED child to running and write its start provenance.

        The ONE seam for "this agent is actually beginning now", shared by the
        deferred root path (called from the scheduler's launcher) and the
        nested path (called inline, since a nested child runs on its parent's
        slot and is never deferred). Everything here is deliberately downstream
        of acceptance: a child that waits for a slot and is cancelled first
        must leave no start hook, no ``start`` event and no spawn attestation
        behind, because none of them would ever be closed.

        ``started_at`` is re-stamped HERE rather than trusted from
        construction: it is what a wall-clock contract burns down against, and
        a child that waited must not arrive with part of its deadline spent.
        The ``start`` event is emitted while the handle still reads
        ``submitted``; the flip to ``running`` follows.
        """
        handle.started_at = time.monotonic()
        self._hook_manager.run_on_agent_start(handle)
        self._emit_event(
            child_ctx,
            "start",
            task_desc,
            handle=handle,
            verification=verification_note,
            model_fallback=model_fallback_note,
        )
        handle.attestation_spawn_hash = self._record_spawn_attestation(
            child_ctx,
            agent_type=agent_type,
            model=model,
            capability_mode=child_ctx.capability_mode,
            contract=contract,
        )
        # Transition to "running" when execution begins.
        handle.status = "running"

    @staticmethod
    def _refusal(*, code: SpawnRefusalCode, reason: str) -> _SpawnOutcome:
        """Build the ONE refusal shape — a typed cause beside the human reason.

        Every refusal reachable from here is PERMANENT: the caller must change
        an argument, because re-issuing the same call fails identically.
        Capacity is deliberately absent — an over-subscribed spawn waits in the
        :class:`~mewbo_core.agents.hypervisor.AgentQueue`, it is never refused.
        """
        return _SpawnOutcome(
            content=f"ERROR: {reason}", agent_id=None, status="rejected", code=code
        )

    async def _spawn_one(
        self,
        args: dict[str, Any],
        *,
        blocking_admit: bool = True,
        batch: list[SpawnUnit] | None = None,
        batch_index: int = 0,
    ) -> _SpawnOutcome:
        """Admit and launch ONE sub-agent — shared core of single + batch spawn.

        ``blocking_admit`` names the CALLER SHAPE, not an admission mode:
        admission never blocks or refuses for capacity, so the flag only
        distinguishes a single, deliberate ``spawn_agent`` (``True`` →
        ``critical`` queue precedence) from one entry of a wide ``spawn_agents``
        fan-out (``False`` → ``normal``), keeping a targeted delegation from
        starving behind a 26-way batch.

        ``batch`` defers acceptance: when supplied, the resolved unit is
        appended to it instead of being handed to the scheduler here, so
        :meth:`run_batch_async` can accept the whole fan-out in one atomic
        step. ``batch_index`` is that entry's declared position.

        Returns a :class:`_SpawnOutcome` carrying the human-readable
        ``content`` plus the structured ``agent_id``/``status``/``code``.
        """
        from mewbo_core.llm.prompt_registry import get_prompt_registry

        registry_prompts = get_prompt_registry()
        task_desc = str(args.get("task", ""))
        acceptance_criteria = str(args.get("acceptance_criteria", "") or "")
        if acceptance_criteria:
            # Contract-first decomposition: delegation is contingent upon the
            # outcome having precise verification.
            task_desc += registry_prompts.render(
                "spawn.acceptance_criteria", acceptance_criteria=acceptance_criteria
            )
        # Task-typed Communication Unit: an
        # explicit `summary_kind` appends its one-line format directive to the
        # child's task text and is stamped onto every `AgentResult` this spawn
        # eventually returns. Unset/unrecognised -> "generic", the untyped
        # summary, which appends nothing to the task text. The directive
        # lookup is a plain dict read on `SpawnAgentTask` — never a branch here.
        raw_summary_kind = args.get("summary_kind")
        summary_kind: SummaryKind = "generic"
        if (
            isinstance(raw_summary_kind, str)
            and raw_summary_kind in SpawnAgentTask.SUMMARY_KIND_DIRECTIVES
        ):
            summary_kind = cast(SummaryKind, raw_summary_kind)
            task_desc += f"\n\n{SpawnAgentTask.SUMMARY_KIND_DIRECTIVES[summary_kind]}"
        # Coarse delegation privilege ceiling. The raw
        # REQUESTED mode; the child context narrows it against this agent's own
        # effective ceiling in ``child()``, so it can only ever restrict.
        raw_capability_mode = args.get("capability_mode", "all")
        requested_capability_mode = (
            raw_capability_mode if isinstance(raw_capability_mode, str) else "all"
        )
        # Filesystem-containment ceiling. The raw REQUESTED tier;
        # ``child()`` narrows it min-wins against this agent's own effective tier,
        # so it can only ever restrict. Default ``workspace_write`` mirrors the
        # ``SpawnAgentTask`` field default (inert until enforcement is flipped on).
        raw_workspace_mode = args.get("workspace_mode", "workspace_write")
        requested_workspace_mode = (
            raw_workspace_mode if isinstance(raw_workspace_mode, str) else "workspace_write"
        )
        # Per-agent delegation bounds — layered UNDER the shared
        # session step budget, never above it. Parsed total (never raises); an
        # absent/malformed ``contract`` is the disabled default: an unbounded
        # child.
        contract = DelegationContract.from_value(args.get("contract"))
        # Ground-truth completion check — parsed total (never raises); an
        # absent/malformed ``verification`` is ``None`` (the ungated child).
        # Threaded into the child loop, which owns the authoritative two-gate
        # activeness decision; the inactive-note below is only for reporting.
        verification = CommandVerification.from_value(args.get("verification"))
        # Per-spawn workspace, resolved BEFORE anything is registered so a
        # refusal costs nothing. An absent ``project`` inherits the parent's
        # directory + instructions; an unresolvable one refuses with the
        # catalog's own reason, and NEVER falls back.
        try:
            workspace = self._child_workspace(args.get("project"))
        except ProjectResolutionError as exc:
            return self._refusal(
                code="unresolvable_project",
                reason=f"sub-agent not spawned — {exc.message}",
            )
        model_override = args.get("model")

        # agent_type: look up registered agent definition and apply its config.
        agent_type = args.get("agent_type")
        if agent_type and self._agent_registry:
            agent_def = self._agent_registry.get(
                agent_type, self._session_capabilities
            )
            if agent_def is None:
                return self._refusal(
                    code="unknown_agent_type",
                    reason=f"Unknown agent type '{agent_type}'",
                )
            # Prepend agent system prompt to task, running the plugin-generic
            # body substitution first so ``${SESSION_ID}``,
            # ``${CLAUDE_PLUGIN_ROOT}``, and bash-style ``${VAR:-default}``
            # expansions resolve before the body hits the child LLM.
            body_subs = {
                "SESSION_ID": self.session_id or "",
                "CLAUDE_PLUGIN_ROOT": agent_def.plugin_root,
            }
            rendered_body = self.substitute_agent_body(agent_def.body, body_subs)
            task_desc = registry_prompts.render(
                "spawn.task_body", body=rendered_body, task=task_desc
            )
            # Apply agent's tool scope if specified and not overridden by caller.
            # ``is not None`` so an AgentDef declaring ``tools: []`` (grant
            # nothing) is applied rather than skipped as if it declared nothing.
            if agent_def.allowed_tools is not None and "allowed_tools" not in args:
                args["allowed_tools"] = agent_def.allowed_tools
            if agent_def.denied_tools and "denied_tools" not in args:
                args["denied_tools"] = agent_def.denied_tools
            # Apply agent's model if specified.
            # Registered agent types with a configured model are authoritative.
            # LLM's model arg on spawn_agent is ignored — config has already made this decision.
            # (Ad-hoc spawns without agent_type continue to honor the LLM's model arg.)
            if agent_def.model:
                model_override = agent_def.model

        registry = self._agent_context.registry

        # A declared ``model_tier`` is the THIRD, LOWEST-priority model source
        # — applied only when neither an agent_type's configured model nor the
        # caller's explicit ``model`` arg already set one (explicit always wins).
        if contract.model_tier is not None:
            raw_tier_map = get_config_value("agent", "model_tiers", default={})
            tier_map = raw_tier_map if isinstance(raw_tier_map, dict) else {}
            allowed_models = self._coerce_list(
                get_config_value("agent", "allowed_models", default=[])
            )
            tier_override = contract.resolve_model_override(
                model_override, tier_map, allowed_models or None
            )
            if tier_override is not None:
                model_override = tier_override

        # 1. Resolve and validate model. A model the gateway will refuse kills
        # the child at step 0 while the parent runs on healthily, and an
        # AgentDef-pinned model is the common way in: it overrides the caller
        # entirely, so nothing upstream ever checked it against what this
        # deployment can actually serve. Under the curated-foundry assumption
        # the declared allowlist IS that check.
        #
        # Falling back beats refusing. The parent's own model is proven — it is
        # what this agent is running on right now — so a child whose declared
        # model is unavailable runs on a model that works instead of dying
        # before its first step. Surfaced in the response + start event, never
        # silently swapped: a caller that pinned a model is owed the fact that
        # it did not get it.
        model_fallback_note: str | None = None
        resolved_model = self._resolve_model(model_override)
        if resolved_model.startswith("ERROR:"):
            parent_model = self._agent_context.model_name
            if not parent_model:
                return self._refusal(
                    code="model_unavailable",
                    reason=resolved_model[len("ERROR: ") :],
                )
            model_fallback_note = (
                f"requested model unavailable ({resolved_model[len('ERROR: '):]}); "
                f"running on the parent's model '{parent_model}'"
            )
            logging.warning("Sub-agent model fallback: {}", model_fallback_note)
            resolved_model = parent_model

        # 2. No admission gate here any more. A depth>=1 spawn drives its child
        # INLINE, so its own slot covers the child for the whole run: the
        # parent is blocked, not running, and only RUNNING consumes a slot.
        # Taking a second slot for a subtree that adds no concurrency is what
        # would deadlock a nested fan-out — every depth-1 parent holding one
        # slot while waiting for another. Root spawns are the ones that genuinely
        # add concurrency, and they go through the queue below.

        child_ctx: AgentContext | None = None
        handle: AgentHandle | None = None
        tq = None  # Initialized early so error handlers can read partial results
        # Set once the root path hands this child to the scheduler + background
        # lifecycle manager: they own its registration, its terminal event and
        # its slot from that point, so the finally must not settle it here.
        lifecycle_owned = False
        try:
            # 3. Create child context. The effective capability_mode is the
            # narrower of the request and this agent's own ceiling; the
            # effective ``atomic`` bit is likewise this contract's OR'd
            # with whatever the parent already carries (see ``child()``).
            child_ctx = self._agent_context.child(
                model_name=resolved_model,
                capability_mode=requested_capability_mode,
                workspace_mode=requested_workspace_mode,
                atomic=contract.atomic,
            )

            # NO-SILENT-DROP: a supplied verification spec whose gate is inert
            # for THIS child (master switch off, or a below-execute
            # capability_mode) is surfaced verbatim in the spawn response +
            # event, never dropped to an invisible null. Non-``None`` only when
            # a spec was supplied AND it will not run — the child loop owns the
            # authoritative decision via the SAME predicate.
            verification_inactive_note: str | None = None
            if verification is not None:
                verification_inactive_note = CommandVerification.inactive_reason(
                    enabled=bool(
                        get_config_value("agent", "verification_enabled", default=False)
                    ),
                    capability_mode=child_ctx.capability_mode,
                )

            # 4. Register in registry.
            # Agent starts as "submitted", transitions to "running"
            handle = AgentHandle(
                agent_id=child_ctx.agent_id,
                parent_id=child_ctx.parent_id,
                depth=child_ctx.depth,
                model_name=child_ctx.model_name,
                task_description=task_desc[:200],
                # The AgentDef name this child was spawned as (``None`` for an
                # ad-hoc spawn) — carried on the handle so every lifecycle event
                # (incl. the background ``stop`` in ``_run_child_lifecycle``,
                # which only holds the handle) can stamp the lane identity.
                agent_type=agent_type if isinstance(agent_type, str) else None,
                status="submitted",
                message_queue=child_ctx.message_queue,
                contract=contract,
            )
            await registry.register(handle)
            # The ``start`` event, the start hook and the spawn attestation all
            # fire at DISPATCH (``_mark_dispatched``), not here. Registration is
            # only acceptance, and a capacity-deferred child may be cancelled
            # or fail to launch before it ever runs — emitting a start for it
            # would leave a span no ``stop`` ever closes, which is precisely
            # the "a start with no stop pins that agent live forever" trap.

            # 5. Filter tool specs (the "filter before binding" pattern). The
            # capability_mode is the child's EFFECTIVE (already-narrowed) mode.
            child_specs = self._filter_tool_specs(
                args, capability_mode=child_ctx.capability_mode
            )

            # 6. Resolve child tool scope + opt-in bounded-retry policy.
            # Privilege attenuation — sub-agents
            # inherit parent's approval policy (not None, which blocks all writes).
            # Three-state, and both collapses matter: ``_coerce_list`` cannot
            # tell absent from empty (it returns ``[]`` for either), so the
            # ``is None`` test happens BEFORE coercion. A parent that spawns a
            # child with ``allowed_tools: []`` means zero tools; reading that as
            # "unrestricted" handed the child the parent's entire spec set.
            _requested_allowed = args.get("allowed_tools")
            child_allowed_tools = (
                None
                if _requested_allowed is None
                else self._coerce_list(_requested_allowed)
            )
            retry = RetryPolicy.from_value(args.get("retry"))

            # Root agent delegates non-blockingly
            # to maintain continuous monitoring capability (epoll model).
            if self._agent_context.depth == 0:
                # Non-blocking: the lifecycle manager drives the child (with
                # bounded retry) and stores the result in the background. Bound
                # to non-optional locals so the launcher closes over values the
                # type checker can see are present.
                launch_ctx, launch_handle = child_ctx, handle
                child_id = launch_ctx.agent_id
                # Captured HERE, not inside ``_launch``: a deferred unit is
                # dispatched by whichever task frees a slot, so by the time the
                # launcher runs the spawning span is neither ambient nor even in
                # the same task. The link is what lets the child's first span
                # name a parent that was actually exported.
                spawn_link = LangfuseTraceLink.capture()

                async def _launch() -> None:
                    """Start this child on the slot the scheduler just handed it."""
                    self._mark_dispatched(
                        launch_ctx,
                        launch_handle,
                        task_desc=task_desc,
                        agent_type=agent_type if isinstance(agent_type, str) else None,
                        model=resolved_model,
                        contract=contract,
                        verification_note=verification_inactive_note,
                        model_fallback_note=model_fallback_note,
                    )
                    with langfuse_child_task_link(spawn_link):
                        lm_task = asyncio.create_task(
                            self._run_child_lifecycle(
                                launch_ctx,
                                launch_handle,
                                child_specs,
                                child_allowed_tools,
                                task_desc,
                                retry,
                                summary_kind,
                                contract,
                                verification,
                                workspace,
                            )
                        )
                    self._lifecycle_tasks.append(lm_task)

                unit = SpawnUnit(
                    spawn=ScheduledSpawn(
                        agent_id=child_id,
                        batch_index=batch_index,
                        priority="critical" if blocking_admit else "normal",
                        enqueued_at=time.monotonic(),
                    ),
                    launch=_launch,
                )
                # Scheduler + lifecycle manager own this child from here.
                lifecycle_owned = True
                # ``submitted`` whether it starts now or waits for a slot: it
                # is accepted and registered either way, and A2A's SUBMITTED is
                # exactly "acknowledged and accepted". The launcher flips it to
                # ``running`` when a slot is actually handed over.
                handle.status = "submitted"
                started_now = False
                if batch is not None:
                    batch.append(unit)
                else:
                    started_now = await registry.accept(unit)
                submitted_body: dict[str, Any] = {
                    "agent_id": child_id,
                    "status": "submitted",
                    "task": task_desc[:200],
                    "message": (
                        "Agent spawned. Use check_agents to monitor "
                        "progress and collect results."
                    )
                    if started_now
                    else (
                        "Agent accepted; the fleet is at its concurrency limit, "
                        "so it starts automatically as soon as a slot frees. "
                        "Use check_agents to monitor it."
                    ),
                }
                # NO-SILENT-DROP: surface a supplied-but-inert verification so the
                # caller never reads a dropped gate as an invisible null. Same
                # contract for a model this deployment cannot serve.
                if verification_inactive_note is not None:
                    submitted_body["verification"] = verification_inactive_note
                if model_fallback_note is not None:
                    submitted_body["model_fallback"] = model_fallback_note
                return _SpawnOutcome(
                    content=json.dumps(submitted_body),
                    agent_id=child_id,
                    status="submitted",
                )

            # A nested spawn dispatches immediately — it runs INSIDE the
            # parent's slot, so there is nothing to wait for and nothing to
            # defer.
            self._mark_dispatched(
                child_ctx,
                handle,
                task_desc=task_desc,
                agent_type=agent_type if isinstance(agent_type, str) else None,
                model=resolved_model,
                contract=contract,
                verification_note=verification_inactive_note,
                model_fallback_note=model_fallback_note,
            )

            # Blocking: current behavior for non-root agents. The retry driver
            # re-delegates the SAME task on a retryable terminal failure (a
            # single attempt when retry is off), raising the last error once the
            # attempt budget is spent.
            tq, state = await self._drive_with_retry(
                child_ctx=child_ctx,
                handle=handle,
                child_specs=child_specs,
                child_allowed_tools=child_allowed_tools,
                task_desc=task_desc,
                retry=retry,
                contract=contract,
                verification=verification,
                workspace=workspace,
            )

            # 7. Settle, with the terminal the child actually reached — the
            # loop returning rather than raising is not evidence of success.
            settled = _ChildSettled(
                summary_kind=summary_kind,
                tq=tq,
                state=state,
                summary=self._child_summary(tq, state),
            )
            result = await self._settle_child(child_ctx, handle, settled)
            result_body = asdict(result)
            if verification_inactive_note is not None:
                result_body["verification"] = verification_inactive_note
            if model_fallback_note is not None:
                result_body["model_fallback"] = model_fallback_note
            return _SpawnOutcome(
                content=json.dumps(result_body),
                agent_id=child_ctx.agent_id,
                # The terminal this settle REACHED, not the result's wider
                # declared type — see ``_SpawnOutcome``. Same value either way.
                status=settled.status,
            )

        except AgentDepthExceeded as exc:
            # child() raises this before `handle` is built or registered, so
            # there is nothing registered to mark done — just surface the result.
            result = AgentResult(
                content=f"Depth exceeded: {exc}",
                status="cannot_solve",
                steps_used=0,
                warnings=[str(exc)],
                summary_kind=summary_kind,
            )
            # A depth refusal answers with a real ``AgentResult`` rather than
            # the bare refusal envelope, so its typed cause rides INSIDE that
            # body — the parent still gets one shape to parse, and the cause
            # is typed rather than only inferable from the prose.
            depth_body = asdict(result)
            depth_body["code"] = "depth_exceeded"
            return _SpawnOutcome(
                content=json.dumps(depth_body),
                agent_id=None,
                status="cannot_solve",
                code="depth_exceeded",
            )

        except asyncio.CancelledError:
            await self._settle_child(
                child_ctx,
                handle,
                _ChildCancelled(
                    summary_kind=summary_kind,
                    partial=(tq.task_result or "") if tq is not None else "",
                ),
            )
            raise  # Re-raise for TaskGroup propagation.

        except Exception as exc:
            logging.error("Sub-agent failed: {}", exc)
            result = await self._settle_child(
                child_ctx,
                handle,
                _ChildFailed.of(
                    exc,
                    child_ctx=child_ctx,
                    summary_kind=summary_kind,
                    task_desc=task_desc,
                    partial=(tq.task_result or "")[:500] if tq is not None else "",
                    fallback_model=resolved_model,
                ),
            )
            return _SpawnOutcome(
                content=json.dumps(asdict(result)),
                agent_id=child_ctx.agent_id if child_ctx else None,
                status="failed",
            )

        finally:
            # Nothing to release: a depth>=1 child ran on its parent's slot,
            # and a root child's slot belongs to the scheduler, which hands it
            # to the lifecycle manager and takes it back exactly once on settle.
            if child_ctx is not None and not lifecycle_owned:
                # Cancel any children spawned by this sub-agent, settling each
                # one in the transcript — see _cascade_cancel_children.
                await self._cascade_cancel_children(child_ctx)

                await registry.unregister(child_ctx.agent_id)

    # ------------------------------------------------------------------
    # Terminal settle — the ONE way a child reaches its terminal state
    # ------------------------------------------------------------------

    async def _settle_child(
        self,
        child_ctx: AgentContext | None,
        handle: AgentHandle | None,
        terminal: _ChildTerminal,
    ) -> AgentResult:
        """Settle ONE child into its terminal — the ONE sequence, for every end.

        Registry mark-done, the stop hook, the single terminal ``stop`` event
        and the terminal attestation happen here, in this order, however the
        child ended. Everything that DIFFERS between a completion, a
        cancellation and a failure is data *terminal* owns, so there is no
        branch on the kind here.

        Both spawn paths ran their own copy of this sequence — the blocking
        nested spawn in :meth:`_spawn_one` and the background manager in
        :meth:`_run_child_lifecycle` — and the copies had drifted. One owner is
        what makes a change to how a child settles reach both callers.

        *child_ctx* and *handle* are optional because a failure can land before
        either exists: the settle then performs the parts it still can rather
        than raising a second exception over the first. The attestation hash is
        ``""`` in that case, exactly as it is when no chain is wired.
        """
        registry = self._agent_context.registry
        if child_ctx is not None:
            await registry.mark_done(
                child_ctx.agent_id,
                terminal.status,
                error=terminal.registry_error(child_ctx, handle),
            )
        if handle is not None:
            self._hook_manager.run_on_agent_stop(handle)
        attestation_hash = ""
        if child_ctx is not None and handle is not None:
            self._emit_terminal_stop(
                child_ctx.event_logger,
                handle,
                terminal.stop_detail,
                summary=terminal.stop_summary,
                summary_kind=terminal.stop_summary_kind,
            )
            attestation_hash = self._record_terminal_attestation(
                child_ctx,
                handle,
                terminal_state=terminal.status,
                summary_kind=terminal.summary_kind,
                summary_text=terminal.attested_summary,
                done_reason=terminal.attested_done_reason,
            )
        return terminal.build_result(handle, attestation_hash)

    @staticmethod
    def _child_summary(tq: TaskQueue, state: OrchestrationState) -> str:
        """Return the child's Communication Unit, falling back to its summary.

        The heaviest children reach their parent with an EMPTY ``task_result``:
        their work landed as store side-effects and their final turn produced no
        text, so everything they learned was discarded at the collection seam.
        A compaction summary is the one compressed record of that work the child
        already produced, so it is used rather than handing the parent nothing.

        This narrows the hole; it does not close it. A child that neither
        answered nor compacted still has no CU, which needs a forced closing
        summary turn inside the child loop.
        """
        if isinstance(tq.task_result, str) and tq.task_result.strip():
            return tq.task_result
        if isinstance(state.summary, str) and state.summary.strip():
            return state.summary
        return ""

    @staticmethod
    def _coerce_timeout(value: object, default: float = 30.0) -> float:
        """Coerce a model-supplied wait timeout, degrading to *default*.

        ``tool_input`` is authored by the model, so this is a trust boundary and
        the read is where it has to be validated. A bare ``float(...)`` over an
        untyped value raises ``ValueError``/``TypeError`` straight out of a
        monitoring call — a malformed argument would kill the run instead of
        merely being ignored. Total, mirroring ``RetryPolicy.from_value``.

        ``bool`` is excluded explicitly: it is an ``int`` subclass, so
        ``timeout: true`` would otherwise silently mean one second.
        """
        if isinstance(value, bool) or value is None:
            return default
        if isinstance(value, (int, float)):
            return float(value)
        if isinstance(value, str):
            try:
                return float(value)
            except ValueError:
                return default
        return default

    @staticmethod
    def _coerce_list(value: object) -> list[str]:
        """Coerce a loose config or tool-argument value to a list of strings.

        Sibling of :meth:`_coerce_timeout` and total for the same reason: it
        reads BOTH an operator's config value (``allowed_models``, which a
        deployment may spell as a comma-separated string) and a model-supplied
        ``allowed_tools``/``denied_tools``, so neither may raise into a spawn.

        It cannot tell ABSENT from EMPTY — both answer ``[]`` — which is why
        every three-state ``allowed_tools`` read tests ``is None`` BEFORE
        coercing.
        """
        if isinstance(value, list):
            return [str(v) for v in value if v]
        if isinstance(value, str):
            return [s.strip() for s in value.split(",") if s.strip()]
        return []

    @staticmethod
    def substitute_agent_body(
        body: str,
        subs: Mapping[str, str],
        env: Mapping[str, str] | None = None,
    ) -> str:
        """Render an agent's body with plugin-generic variable substitution.

        Three passes, in order:

        1. ``${KEY}`` literal substitution from *subs*. Core passes
           ``SESSION_ID`` and ``CLAUDE_PLUGIN_ROOT``; plugins author their
           prompts against these names.
        2. Bash-style ``${VAR:-default}`` — if ``VAR`` is unset in *env*,
           the text expands to ``default``. If ``VAR`` is set, it expands
           to the env value. This keeps plugin prompts self-documenting
           (operator override path is obvious in the source).
        3. Plain ``$VAR`` expansion as a final pass, matching
           :func:`os.path.expandvars` semantics. Unset variables remain
           literal so authors can spot typos at glance.

        A ``staticmethod`` on this class rather than a loose module function:
        this tool is the only production caller, and *subs*/*env* arrive as
        ARGS so the renderer stays testable with a fake environment. The
        module-level ``substitute_agent_body`` name is a thin alias over it, so
        the import path plugins and tests already use is unchanged.
        """
        if env is None:
            env = os.environ
        # Pass 1: direct substitutions.
        for key, value in subs.items():
            body = body.replace(f"${{{key}}}", value)

        # Pass 2: bash-style ${VAR:-default}. Read from env; fall back to default.
        def _bash_default(match: re.Match[str]) -> str:
            var_name, default = match.group(1), match.group(2)
            return env.get(var_name, default)

        body = _BASH_DEFAULT_RE.sub(_bash_default, body)

        # Pass 3: plain ``$VAR`` expansion for anything still referencing env.
        # Matches ``os.path.expandvars`` semantics without touching the real
        # ``os.environ`` when a test supplies a fake *env* mapping.
        def _plain_var(match: re.Match[str]) -> str:
            var_name = match.group(1)
            return env.get(var_name, match.group(0))

        return re.sub(r"\$(\w+)", _plain_var, body)

    async def has_live_owned_runs(self) -> bool:
        """True while any agent this session owns is still non-terminal.

        The ownership index behind the promise-as-completion gate: a clean
        terminal declared while owned work is live is a promise, not a
        completion. Exposed here because the hypervisor is the only thing that
        knows, and the seam that must ask is the one accepting a completion
        claim.
        """
        try:
            agents = await self._agent_context.registry.list_all()
        except Exception:  # pragma: no cover - defensive; never fail a run here
            return False
        return any(handle.status in ACTIVE_STATUSES for handle in agents)

    # ------------------------------------------------------------------
    # Model resolution
    # ------------------------------------------------------------------

    def _resolve_model(self, model_override: object) -> str:
        """Resolve model for the child agent.

        Returns model name, or ``"ERROR: ..."`` string on validation failure.
        """
        allowed_models = self._coerce_list(
            get_config_value("agent", "allowed_models", default=[])
        )
        default_sub = str(get_config_value("agent", "default_sub_model", default="") or "").strip()

        if model_override and isinstance(model_override, str):
            model = model_override.strip()
            if allowed_models and model not in allowed_models:
                return (
                    f"ERROR: Model '{model}' not in allowed_models. "
                    f"Available: {', '.join(allowed_models)}"
                )
            return model

        if default_sub:
            return default_sub

        return self._agent_context.model_name

    # ------------------------------------------------------------------
    # Tool spec filtering (the "filter before binding" pattern)
    # ------------------------------------------------------------------

    def _filter_tool_specs(
        self, args: dict[str, Any], *, capability_mode: str = "all"
    ) -> list[ToolSpec]:
        """Filter tool specs for a child agent.

        Filtering starts from the PARENT's effective set (the registry only when
        this tool was never stamped), so the child can only ever narrow it —
        the same monotone containment ``capability_mode`` already enforces at
        :meth:`AgentContext.child`, applied to the spec set itself.

        Denied tools are removed from the child's ``bind_tools()`` list —
        the child LLM never sees them. ``capability_mode`` is the
        child's effective privilege ceiling, applied as a coarse pre-filter
        LAYERED UNDER the allow/deny gates (it only removes more, never adds).
        """
        parent_specs = self.parent_tool_specs
        if parent_specs is None:
            parent_specs = self._tool_registry.list_specs()
        # ``allowed_tools`` is three-state — see ``_spawn_one``'s resolution of
        # ``child_allowed_tools``. Absent stays unrestricted; an empty list is
        # forwarded as an empty list so ``filter_specs`` grants nothing.
        requested_allowed = args.get("allowed_tools")
        return filter_specs(
            parent_specs,
            allowed=(
                None if requested_allowed is None else self._coerce_list(requested_allowed)
            ),
            denied=self._coerce_list(args.get("denied_tools") or []),
            capability_mode=capability_mode,
        )

    # ------------------------------------------------------------------
    # Agent management handlers (root-only)
    # ------------------------------------------------------------------

    async def handle_check_agents(self, action_step: ActionStep) -> MockSpeaker:
        """Return agent tree state with completed results and progress.

        Emits a JSON payload with ``kind: "agent_tree"``. The ``text`` field
        carries the rendered ASCII tree the LLM consumes; the ``agents`` list
        is the structured snapshot the console uses to render CheckAgentsCard.
        """
        args = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
        wait = bool(args.get("wait", False))
        timeout = self._coerce_timeout(args.get("timeout"))
        registry = self._agent_context.registry
        parent_id = self._agent_context.agent_id

        if wait:
            running = await registry.collect_running(parent_id)
            if running:
                waiters = [asyncio.create_task(h.done_event.wait()) for h in running]
                # asyncio.wait() does not raise on timeout — it returns
                # (done, pending) with pending non-empty when the deadline hits.
                _done, pending = await asyncio.wait(
                    waiters,
                    timeout=timeout,
                    return_when=asyncio.FIRST_COMPLETED,
                )
                for t in pending:
                    t.cancel()

        tree = await registry.render_agent_tree(
            exclude_agent_id=self._agent_context.agent_id,
        )
        completed = await registry.collect_completed(parent_id)
        running = await registry.collect_running(parent_id)

        parts: list[str] = []
        if tree:
            parts.append(f"Agent tree:\n{tree}")

        if completed:
            parts.append("\nCompleted results:")
            for h in completed:
                r = h.result
                if r:
                    parts.append(f"  [{h.agent_id[:8]}] {r.status}: {r.summary or r.content[:300]}")

        if running:
            parts.append(f"\n{len(running)} agent(s) still running.")
        elif not completed:
            parts.append("No agents spawned.")

        text = "\n".join(parts) or "No agents."

        agents_payload: list[dict[str, Any]] = []
        for h in await registry.list_visible(exclude_agent_id=self._agent_context.agent_id):
            entry: dict[str, Any] = {
                "id": h.agent_id,
                "parent_id": h.parent_id,
                "depth": h.depth,
                "task": h.task_description,
                "status": h.status,
                "steps_completed": h.steps_completed,
                "last_tool_id": h.last_tool_id,
                "progress_note": h.progress_note,
                "compaction_count": h.compaction_count,
                "attempts": h.attempts,  # Retry provenance for the console
                "result": (
                    {
                        "status": h.result.status,
                        "summary": h.result.summary,
                        "content": h.result.content,
                    }
                    if h.result is not None
                    else None
                ),
            }
            # Additive, ONLY when a real bound was declared, so a
            # contract-less child's payload carries no key. Carries LIVE
            # progress (steps/elapsed) alongside the bounded-scalar
            # ``snapshot()`` shape the attestation record reuses.
            if h.contract.enabled:
                entry["contract"] = {
                    **h.contract.snapshot(),
                    "steps_completed": h.steps_completed,
                    "elapsed_s": time.monotonic() - h.started_at,
                }
            agents_payload.append(entry)

        payload = {
            "kind": "agent_tree",
            "text": text,
            "agents": agents_payload,
            "parent_id": parent_id,
            "wait": wait,
        }
        return MockSpeaker(content=json.dumps(payload))

    async def handle_steer_agent(self, action_step: ActionStep) -> MockSpeaker:
        """Send a steering message to or cancel a running agent."""
        args = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
        agent_id = str(args.get("agent_id", ""))
        action = str(args.get("action", ""))
        message = str(args.get("message", ""))
        registry = self._agent_context.registry

        # Resolve short prefix to full agent_id.
        handle = await registry.get(agent_id)
        if handle is None:
            all_agents = await registry.list_all()
            matches = [h for h in all_agents if h.agent_id.startswith(agent_id)]
            if len(matches) == 1:
                handle = matches[0]
                agent_id = handle.agent_id
            elif len(matches) > 1:
                return MockSpeaker(
                    content=f"ERROR: Ambiguous prefix '{agent_id}' matches {len(matches)} agents.",
                )
            else:
                return MockSpeaker(
                    content=f"ERROR: Agent '{agent_id}' not found.",
                )

        if action == "cancel":
            reason = await registry.cancel_agent(agent_id)
            if reason is None:
                return MockSpeaker(
                    content=f"Agent {agent_id[:8]} cancelled.",
                )
            return MockSpeaker(
                content=f"Agent {agent_id[:8]} cannot cancel: {reason}",
            )
        if action == "message":
            if not message:
                return MockSpeaker(
                    content="ERROR: 'message' is required when action='message'.",
                )
            reason = await registry.send_message(
                agent_id,
                f"[From parent] {message}",
            )
            if reason is None:
                return MockSpeaker(content="Message sent.")
            return MockSpeaker(content=f"Message failed: {reason}")

        return MockSpeaker(
            content=f"ERROR: Unknown action '{action}'. Use 'message' or 'cancel'.",
        )

    # ------------------------------------------------------------------
    # Event emission
    # ------------------------------------------------------------------

    def _emit_event(
        self,
        ctx: AgentContext,
        action: str,
        detail: str,
        handle: AgentHandle | None = None,
        summary: str | None = None,
        summary_kind: SummaryKind | None = None,
        verification: str | None = None,
        model_fallback: str | None = None,
    ) -> None:
        """Emit a sub_agent lifecycle event.

        ``summary`` (additive, set only on the terminal ``stop``) carries the
        child's compressed result — the Communication Unit a downstream consumer
        can project as the sub-agent's actual response (e.g. the agentic-search
        trace's per-lane evidence block, where the lifecycle ``detail`` is just
        the ``done_reason``). Omitted on every other phase, so a consumer
        reading only the base keys is unaffected.
        ``summary_kind`` rides alongside ``summary`` and is
        likewise omitted when it's the untyped ``"generic"`` default, so a
        consumer that never declared a kind sees no new key.
        ``verification`` (additive, set only on ``start`` when a supplied spec's
        gate is inert) is the no-silent-drop note — omitted otherwise, so a
        spawn with no verification, or with an active one, carries no key.
        ``model_fallback`` is the same contract for a declared model this
        deployment cannot serve: present only when the child was moved off it.
        """
        if ctx.event_logger is None:
            return
        self._write_lifecycle(
            ctx.event_logger,
            action=action,
            agent_id=ctx.agent_id,
            parent_id=ctx.parent_id,
            depth=ctx.depth,
            model=ctx.model_name,
            detail=detail,
            handle=handle,
            summary=summary,
            summary_kind=summary_kind,
            verification=verification,
            model_fallback=model_fallback,
        )

    def _write_lifecycle(
        self,
        logger: Callable[[Event], None],
        *,
        action: str,
        agent_id: str,
        parent_id: str | None,
        depth: int,
        model: str,
        detail: str,
        handle: AgentHandle | None = None,
        summary: str | None = None,
        summary_kind: SummaryKind | None = None,
        verification: str | None = None,
        model_fallback: str | None = None,
    ) -> None:
        """Write ONE ``sub_agent`` lifecycle event — the single payload shape.

        Identity is passed explicitly rather than read off a context because a
        parent settling a child during teardown cascade holds the child's
        HANDLE, not its context. Consoles parse this shape, so every lifecycle
        phase must build it here and nowhere else.
        """
        payload: dict[str, Any] = {
            "action": action,
            "agent_id": agent_id,
            "parent_id": parent_id,
            "depth": depth,
            "model": model,
            "detail": detail,
            "status": handle.status if handle else action,
            "steps_completed": handle.steps_completed if handle else 0,
            "input_tokens": handle.input_tokens if handle else 0,
            "output_tokens": handle.output_tokens if handle else 0,
        }
        # ``agent_type`` (additive) carries the spawned AgentDef name so a
        # consumer can label the lane by its DEFINITION (e.g.
        # ``scg-path-probe``) instead of falling back to the model name —
        # the agentic-search trace projection's lane identity. Read off the
        # handle so the terminal ``stop`` (emitted from the background
        # lifecycle manager, which holds only the handle) carries it too.
        # Omitted for an ad-hoc spawn (no ``agent_type``) so a consumer
        # reading only the base keys is untouched.
        if handle is not None and handle.agent_type:
            payload["agent_type"] = handle.agent_type
        if summary:
            payload["summary"] = summary
        if summary_kind and summary_kind != "generic":
            payload["summary_kind"] = summary_kind
        if verification:
            payload["verification"] = verification
        if model_fallback:
            payload["model_fallback"] = model_fallback
        event: Event = {"type": "sub_agent", "payload": payload}
        logger(event)

    def _emit_terminal_stop(
        self,
        logger: Callable[[Event], None] | None,
        handle: AgentHandle,
        detail: str,
        *,
        summary: str | None = None,
        summary_kind: SummaryKind | None = None,
    ) -> None:
        """Write this agent's ONE terminal ``stop``, at most once.

        Every way an agent can settle — success, failure, cancellation, or a
        parent's teardown cascade — routes through here, because a consumer
        derives liveness from the last payload it sees and a start with no stop
        pins that agent live forever. Several of those paths can fire for the
        SAME agent (a cascade cancel lands, then the cancelled child's own
        handler unwinds), so the at-most-once guarantee is latched on the handle
        rather than assumed from the paths being mutually exclusive.

        Called AFTER the registry has been marked done, so ``handle.status``
        already carries the terminal state the payload reports.
        """
        if handle.terminal_emitted:
            return
        handle.terminal_emitted = True
        if logger is None:
            return
        self._write_lifecycle(
            logger,
            action="stop",
            agent_id=handle.agent_id,
            parent_id=handle.parent_id,
            depth=handle.depth,
            model=handle.model_name,
            detail=detail,
            handle=handle,
            summary=summary,
            summary_kind=summary_kind,
        )

    async def _cascade_cancel_children(self, ctx: AgentContext) -> None:
        """Cancel this agent's own children and settle each one in the log.

        A child of a settling agent has no one left to drive it, and one still
        at ``submitted`` never had a loop task whose cancellation could raise
        into its own handler — so the terminal is written HERE.

        The skip condition is the emission latch, NOT the child's status: a
        terminal status is no evidence a terminal EVENT was ever written. Other
        teardown paths (the child loop's own end-of-run cascade, the
        hypervisor's shutdown force-mark) settle handles in memory only, and
        they run first — so a child arriving here already marked ``cancelled``
        is precisely the one whose span would otherwise stay open forever.

        ``mark_done`` follows the cancel so the emitted payload always reports a
        terminal status: ``cancel_agent`` is a no-op for a child that never got
        an asyncio task, which would otherwise emit a ``stop`` still reading
        ``submitted``.
        """
        registry = self._agent_context.registry
        for child in await registry.list_children(ctx.agent_id):
            if child.terminal_emitted:
                continue
            if child.status in ACTIVE_STATUSES:
                await registry.cancel_agent(child.agent_id)
                await registry.mark_done(child.agent_id, "cancelled")
                detail = "parent agent settled; child cancelled with it"
            else:
                # Already terminal in memory but never written: report how it
                # actually settled rather than claiming this cascade ended it.
                detail = f"settled as {child.status} with no terminal event recorded"
            self._emit_terminal_stop(ctx.event_logger, child, detail)

    # ------------------------------------------------------------------
    # Attestation provenance
    # ------------------------------------------------------------------

    def _record_spawn_attestation(
        self,
        child_ctx: AgentContext,
        *,
        agent_type: str | None,
        model: str,
        capability_mode: str,
        contract: DelegationContract,
    ) -> str:
        """Best-effort spawn provenance record for ``child_ctx``.

        ``""`` when no ``AttestationChain`` is wired for this session (the
        common case today — the feature is orchestrator-injected and gated on
        config) or the chain's own append failed; never raises.
        """
        chain = getattr(child_ctx.registry, "attestation", None)
        if chain is None:
            return ""
        return chain.record_spawn(
            child_ctx.event_logger,
            agent_id=child_ctx.agent_id,
            parent_id=child_ctx.parent_id,
            depth=child_ctx.depth,
            agent_type=agent_type,
            model=model,
            capability_mode=capability_mode,
            contract=contract,
        )

    def _record_terminal_attestation(
        self,
        child_ctx: AgentContext,
        handle: AgentHandle,
        *,
        terminal_state: AgentStatus,
        summary_kind: SummaryKind,
        summary_text: str,
        done_reason: str | None,
    ) -> str:
        """Best-effort terminal provenance record for ``child_ctx``.

        Mirrors :meth:`_record_spawn_attestation`'s no-chain/no-raise contract.
        ``summary_text`` is hashed inside the chain, never persisted verbatim.
        """
        chain = getattr(child_ctx.registry, "attestation", None)
        if chain is None:
            return ""
        return chain.record_terminal(
            child_ctx.event_logger,
            agent_id=child_ctx.agent_id,
            parent_id=child_ctx.parent_id,
            depth=child_ctx.depth,
            terminal_state=terminal_state,
            spawn_hash=handle.attestation_spawn_hash,
            attempts=handle.attempts,
            steps_completed=handle.steps_completed,
            input_tokens=handle.input_tokens,
            output_tokens=handle.output_tokens,
            summary_kind=summary_kind,
            summary_text=summary_text,
            done_reason=done_reason,
        )

    # ------------------------------------------------------------------
    # Bounded retry driver
    # ------------------------------------------------------------------

    def _build_child_loop(
        self,
        child_ctx: AgentContext,
        child_allowed_tools: list[str] | None,
        contract: DelegationContract | None = None,
        verification: CommandVerification | None = None,
        workspace: ChildWorkspace | None = None,
    ) -> Any:
        """Construct a fresh child ``ToolUseLoop`` for one attempt.

        A fresh loop per attempt means each retry gets its own
        ``RetryStrategy`` (the model-fallback ladder) — so model-level recovery
        is reused, never reinvented at this layer. ``contract``,
        ``verification`` and ``workspace`` ride along unchanged across retries —
        they are the SPAWNER's declared bound/check/directory for this child,
        not per-attempt state. All three default to off/absent so existing
        callers are unaffected; the child loop owns the authoritative two-gate
        decision on whether verification runs.
        """
        # Import here to avoid a circular import at module load time.
        from mewbo_core.loop.tool_use_loop import ToolUseLoop

        # An absent workspace is a caller that resolved none, which inherits
        # this agent's current directory — the pre-``project`` behaviour.
        ws = workspace or ChildWorkspace(
            cwd=self._cwd, project_instructions=self._project_instructions
        )
        return ToolUseLoop(
            agent_context=child_ctx,
            tool_registry=self._tool_registry,
            permission_policy=self._permission_policy,
            approval_callback=self._approval_callback,
            hook_manager=self._hook_manager,
            # Forwarded unchanged — the SAME object as the parent, never
            # re-loaded from disk or config. A child inherits its parent's
            # plane exactly as it inherits ``hook_manager``; there is no path
            # by which a descendant runs under a weaker (or no) plane than the
            # session that spawned it.
            safety_plane=self._safety_plane,
            project_instructions=ws.project_instructions,
            user_instructions=self._user_instructions,
            session_tool_registry=self._session_tool_registry,
            allowed_tools=child_allowed_tools,
            # A spawned sub-agent's allowlist is AUTHORITATIVE — its specs are
            # already strictly filtered (``_filter_tool_specs``, no built-in
            # exemption), and the spawn_agent gate must honour it so a leaf
            # scoped without spawn_agent cannot recurse into copies of itself.
            strict_tool_scope=True,
            cwd=ws.cwd,
            session_id=self.session_id,
            session_capabilities=self._session_capabilities,
            enable_skills=self._enable_skills,
            contract=contract or DelegationContract(),
            verification=verification,
        )

    async def _drive_with_retry(
        self,
        *,
        child_ctx: AgentContext,
        handle: AgentHandle,
        child_specs: list[ToolSpec],
        child_allowed_tools: list[str] | None,
        task_desc: str,
        retry: RetryPolicy,
        contract: DelegationContract | None = None,
        verification: CommandVerification | None = None,
        workspace: ChildWorkspace | None = None,
    ) -> tuple[Any, Any]:
        """Run the child loop, re-delegating the SAME task on a retryable failure.

        Returns ``(tq, state)`` from the first attempt that reached a real
        completion. Re-raises the LAST exception once the attempt budget is
        spent or the failure cause is not in ``retry.on`` (default off ⇒
        exactly one attempt).
        ``CancelledError`` is never retried — parent cancellation is terminal
        and bubbles straight up.

        BOTH child-death shapes are retryable: an exception out of the loop, and
        a loop that RETURNS having stopped short (a halt, a spent budget, a
        failed ground-truth check). The second shape is the common one, so
        watching only the first is why this contract had never once fired in
        production. A stopped-short return is still handed back after the
        budget is spent — the caller reports it honestly rather than raising.

        Slot discipline: the one semaphore slot already acquired in
        ``run_async`` is *held across all attempts* — re-admission re-uses that
        slot rather than releasing and racing for a new one, so concurrency stays
        bounded exactly as on the no-retry path. ``handle.attempts`` is bumped per
        attempt so the agent tree / ``check_agents`` surface the re-delegation.
        """
        attempt = 0
        while True:
            attempt += 1
            handle.attempts = attempt
            # Each attempt is a fresh run on the SAME handle/agent_id: reset the
            # transient running state (the prior attempt's loop marked it failed
            # in its own finally) so the tree reflects the live attempt.
            handle.status = "running"
            handle.error = None
            child_loop = self._build_child_loop(
                child_ctx, child_allowed_tools, contract, verification, workspace
            )
            # The child's whole span subtree hangs off whatever is ambient at
            # THIS create_task; a blocking (depth ≥ 1) spawn still sits inside
            # the spawning span here, while the root path already carries the
            # link its launcher bound. Capturing covers both without a branch.
            with langfuse_child_task_link():
                child_task = asyncio.create_task(
                    child_loop.run(task_desc, tool_specs=child_specs, mode=self.parent_mode)
                )
            # Populate asyncio_task so cancel_agent() and 3-phase cleanup target
            # the live attempt.
            handle.asyncio_task = child_task
            try:
                tq, state = await child_task
            except asyncio.CancelledError:
                raise  # parent cancellation is terminal — never retried
            except Exception as exc:  # noqa: BLE001 — child-loop failure is opaque
                cause = RetryPolicy.classify_cause(exc)
                if not retry.should_retry(cause, attempt):
                    raise
                await self._back_off_before_retry(child_ctx, handle, retry, attempt, cause)
                continue

            # A child that stops short RETURNS; it does not raise. Doom-loop
            # halts, spent budgets and failed ground-truth checks are the
            # dominant child-death shapes and every one of them arrives here as
            # an ordinary return — so a policy watching only the exception path
            # could never fire for them, which is why the contract had never
            # fired at all. Bucketed as the generic ``failed`` cause: nothing
            # about a halt names a provider, so ``timeout`` would be a lie.
            # Keyed on ``failed`` rather than "anything but completed": a child
            # that returns having been CANCELLED also stops short of completed,
            # and re-delegating it would restart work somebody deliberately
            # stopped. Cancellation is structurally unretryable on the raising
            # path directly above; a cooperative cancel returns instead of
            # raising, so it needs the same guarantee stated here.
            if state.terminal_status() == "failed" and retry.should_retry(
                "failed", attempt
            ):
                await self._back_off_before_retry(
                    child_ctx, handle, retry, attempt, state.done_reason or "failed"
                )
                continue
            return tq, state

    async def _back_off_before_retry(
        self,
        child_ctx: AgentContext,
        handle: AgentHandle,
        retry: RetryPolicy,
        attempt: int,
        cause: str,
    ) -> None:
        """Announce a re-delegation and sleep its backoff.

        Shared by both retryable paths so a raising failure and a stopped-short
        return are announced identically — a consumer counting re-delegations
        must not have to know which shape produced one.
        """
        delay = retry.backoff_for(attempt)
        self._emit_event(
            child_ctx,
            "retry",
            f"attempt {attempt} {cause}; re-delegating (max {retry.max})",
            handle=handle,
        )
        logging.warning(
            "Sub-agent {} attempt {} failed ({}); retrying in {:.1f}s",
            child_ctx.agent_id[:8],
            attempt,
            cause,
            delay,
        )
        if delay > 0:
            await asyncio.sleep(delay)

    # ------------------------------------------------------------------
    # Non-blocking lifecycle manager
    # ------------------------------------------------------------------

    async def _run_child_lifecycle(
        self,
        child_ctx: AgentContext,
        handle: AgentHandle,
        child_specs: list[ToolSpec],
        child_allowed_tools: list[str] | None,
        task_desc: str,
        retry: RetryPolicy,
        summary_kind: SummaryKind,
        contract: DelegationContract | None = None,
        verification: CommandVerification | None = None,
        workspace: ChildWorkspace | None = None,
    ) -> None:
        """Background lifecycle manager for non-blocking child execution.

        Drives the child (with bounded retry), stores the ``AgentResult`` on the
        handle, and notifies the parent via ``send_to_parent``.

        Emits a lifecycle event at each phase transition and stores the CU on
        the handle for async retrieval. ``summary_kind``
        is the caller's declared CU shape, resolved once in
        ``_spawn_one`` and stamped onto every ``AgentResult`` this manager builds.
        """
        registry = self._agent_context.registry
        tq = None
        try:
            tq, state = await self._drive_with_retry(
                child_ctx=child_ctx,
                handle=handle,
                child_specs=child_specs,
                child_allowed_tools=child_allowed_tools,
                task_desc=task_desc,
                retry=retry,
                contract=contract,
                verification=verification,
                workspace=workspace,
            )

            # The child's REAL terminal, not the fact that its loop returned —
            # see ``OrchestrationState.terminal_status``.
            handle.result = await self._settle_child(
                child_ctx,
                handle,
                _ChildSettled(
                    summary_kind=summary_kind,
                    tq=tq,
                    state=state,
                    summary=self._child_summary(tq, state),
                ),
            )

        except asyncio.CancelledError:
            handle.result = await self._settle_child(
                child_ctx,
                handle,
                _ChildCancelled(
                    summary_kind=summary_kind,
                    partial=(tq.task_result or "") if tq is not None else "",
                ),
            )

        except Exception as exc:
            logging.error("Sub-agent lifecycle failed: {}", exc)
            handle.result = await self._settle_child(
                child_ctx,
                handle,
                _ChildFailed.of(
                    exc,
                    child_ctx=child_ctx,
                    summary_kind=summary_kind,
                    task_desc=task_desc,
                    partial=(tq.task_result or "")[:500] if tq is not None else "",
                ),
            )

        finally:
            # Cascade cleanup to children of this child, settling each one in
            # the transcript — see _cascade_cancel_children.
            await self._cascade_cancel_children(child_ctx)

            # Notify parent before releasing the semaphore slot.
            # Result and status are set in the try/except blocks above.
            # The handle stays in the registry so check_agents and
            # render_agent_tree can surface the result; session cleanup()
            # clears it at session end.
            if handle and handle.result:
                notification = (
                    f"[Agent {child_ctx.agent_id[:8]} {handle.result.status}] "
                    f"Task: {task_desc} | "
                    f"{handle.result.summary or handle.result.content[:300]}"
                )
            else:
                _status = handle.status if handle else "unknown"
                notification = f"[Agent {child_ctx.agent_id[:8]} {_status}] Task: {task_desc}"
            await registry.send_to_parent(child_ctx.agent_id, notification)

            # The dispatch pump: this hands the freed slot straight to whatever
            # spawn has been waiting for one, so a queued child starts on the
            # completion path that every settling agent already runs.
            await registry.release()

    async def await_lifecycle_managers(self, timeout: float = 3.0) -> None:
        """Settle every background lifecycle manager before the run tears down.

        Called from ``ToolUseLoop.run()``'s finally block. Collect-or-cancel:
        managers get *timeout* to finish on their own, then the stragglers are
        cancelled — and, crucially, AWAITED.

        Cancelling without awaiting is what leaked children. ``cancel()`` only
        schedules the ``CancelledError``; the handler that marks the child done
        and writes its ONE terminal ``stop`` runs on a later turn of the event
        loop, which never comes if the loop tears down first. The child was then
        settled by nothing in this process — its span stayed open until a boot
        sweep reaped it days later, or forever. Awaiting here is what makes
        settlement in-process rather than next-boot.

        Exceptions are swallowed by ``return_exceptions``: these tasks own their
        own terminal reporting, and a manager that fails while being torn down
        must not take down the run that is already ending.

        **The set to await is not fixed at entry.** Every manager that settles
        hands its slot to a QUEUED child through the dispatch pump, which
        creates a new manager — so this DRAINS in rounds against one deadline
        rather than awaiting a snapshot. Cancelling waiting units up front
        would be the simpler code and the wrong behaviour: it would discard, at
        teardown, exactly the work this scheduler exists to stop discarding.
        Only once the budget is spent are the still-waiting units dropped —
        after that nothing will ever dispatch them, so they must be settled
        rather than left reading as live.
        """
        deadline = time.monotonic() + max(0.0, timeout)
        while True:
            pending = [t for t in self._lifecycle_tasks if not t.done()]
            remaining = deadline - time.monotonic()
            if not pending or remaining <= 0:
                break
            await asyncio.wait(
                pending,
                timeout=remaining,
                return_when=asyncio.ALL_COMPLETED,
            )

        # ROOT ONLY. The queue is SESSION-global while this method runs at the
        # end of EVERY agent's loop, root or child — so an unguarded clear here
        # let the first sub-agent to finish discard the root's still-waiting
        # fan-out, which is precisely the loss this scheduler exists to stop.
        # A depth>=1 tool owns nothing in the queue by construction: a nested
        # spawn runs inline on its parent's slot and is never enqueued.
        if self._agent_context.depth == 0:
            await self._agent_context.registry.cancel_pending()

        still_pending = [t for t in self._lifecycle_tasks if not t.done()]
        for task in still_pending:
            task.cancel()
        if still_pending:
            await asyncio.gather(*still_pending, return_exceptions=True)
        self._lifecycle_tasks.clear()

__init__(*, agent_context: AgentContext, tool_registry: ToolRegistry, permission_policy: PermissionPolicy, approval_callback: Callable[[ActionStep], bool] | None = None, hook_manager: HookManager, safety_plane: SafetyPlane | None = None, project_instructions: str | None = None, user_instructions: str | None = None, cwd: str | None = None, agent_registry: Any = None, session_tool_registry: SessionToolRegistry | None = None, session_capabilities: tuple[str, ...] = (), enable_skills: bool = True, catalog: ProjectCatalog | None = None) -> None

Initialize with parent context and shared registries.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
def __init__(
    self,
    *,
    agent_context: AgentContext,
    tool_registry: ToolRegistry,
    permission_policy: PermissionPolicy,
    approval_callback: Callable[[ActionStep], bool] | None = None,
    hook_manager: HookManager,
    safety_plane: SafetyPlane | None = None,
    project_instructions: str | None = None,
    user_instructions: str | None = None,
    cwd: str | None = None,
    agent_registry: Any = None,
    session_tool_registry: SessionToolRegistry | None = None,
    session_capabilities: tuple[str, ...] = (),
    enable_skills: bool = True,
    catalog: ProjectCatalog | None = None,
) -> None:
    """Initialize with parent context and shared registries."""
    self._agent_context = agent_context
    self._tool_registry = tool_registry
    self._permission_policy = permission_policy
    self._approval_callback = approval_callback
    self._hook_manager = hook_manager
    self._safety_plane = safety_plane
    self._project_instructions = project_instructions
    # Operator-authored custom instructions are inherited by every child,
    # exactly like the project instructions — they describe the deployment,
    # not one agent's task, so a sub-agent that lost them would be running
    # under different rules than its parent.
    self._user_instructions = user_instructions
    self._cwd = cwd
    # The catalog a per-spawn ``project`` key is resolved through. Injected
    # rather than reached for: a CLI drive has no app-side stores to build
    # one from, and ``None`` there must refuse a ``project`` outright rather
    # than resolve it to something plausible.
    self._catalog = catalog
    self._agent_registry = agent_registry
    self._session_tool_registry = session_tool_registry
    self._session_capabilities = session_capabilities
    # Children inherit the parent drive's skill policy: a headless search
    # run disables auto-skill injection for the ROOT *and* every probe it
    # spawns (the audit found every server-side agent burning step 1 on
    # ``activate_skill``).
    self._enable_skills = enable_skills
    # Plan-mode context — set by ToolUseLoop.run() so children
    # inherit the session's plan path and mode.
    self.session_id: str | None = None
    self.parent_mode: str = "act"
    # The parent's EFFECTIVE spec set, stamped by ToolUseLoop.run() alongside
    # the plan context (both are run()-time state — the loop only learns its
    # own specs when it is handed them, long after this tool is constructed).
    # Containment must be monotone: a child narrows, never widens. Deriving a
    # child from the raw registry instead would hand it tools the parent never
    # held, since a parent's set can be narrowed by scoping this tool cannot
    # reconstruct (a console session's ``allowed_tools`` over MCP tools, an
    # ancestor's own filtering). ``None`` means unstamped — fall back to the
    # registry so a directly-constructed tool keeps the full set.
    self.parent_tool_specs: list[ToolSpec] | None = None
    # Track lifecycle manager tasks for deterministic cleanup.
    self._lifecycle_tasks: list[asyncio.Task[None]] = []

await_lifecycle_managers(timeout: float = 3.0) -> None async

Settle every background lifecycle manager before the run tears down.

Called from ToolUseLoop.run()'s finally block. Collect-or-cancel: managers get timeout to finish on their own, then the stragglers are cancelled — and, crucially, AWAITED.

Cancelling without awaiting is what leaked children. cancel() only schedules the CancelledError; the handler that marks the child done and writes its ONE terminal stop runs on a later turn of the event loop, which never comes if the loop tears down first. The child was then settled by nothing in this process — its span stayed open until a boot sweep reaped it days later, or forever. Awaiting here is what makes settlement in-process rather than next-boot.

Exceptions are swallowed by return_exceptions: these tasks own their own terminal reporting, and a manager that fails while being torn down must not take down the run that is already ending.

The set to await is not fixed at entry. Every manager that settles hands its slot to a QUEUED child through the dispatch pump, which creates a new manager — so this DRAINS in rounds against one deadline rather than awaiting a snapshot. Cancelling waiting units up front would be the simpler code and the wrong behaviour: it would discard, at teardown, exactly the work this scheduler exists to stop discarding. Only once the budget is spent are the still-waiting units dropped — after that nothing will ever dispatch them, so they must be settled rather than left reading as live.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
2861
2862
2863
2864
2865
2866
2867
2868
2869
2870
2871
2872
2873
2874
2875
2876
2877
2878
2879
2880
2881
2882
2883
2884
2885
2886
2887
2888
2889
2890
2891
2892
2893
2894
2895
2896
2897
2898
2899
2900
2901
2902
2903
2904
2905
2906
2907
2908
2909
2910
2911
2912
2913
2914
2915
2916
async def await_lifecycle_managers(self, timeout: float = 3.0) -> None:
    """Settle every background lifecycle manager before the run tears down.

    Called from ``ToolUseLoop.run()``'s finally block. Collect-or-cancel:
    managers get *timeout* to finish on their own, then the stragglers are
    cancelled — and, crucially, AWAITED.

    Cancelling without awaiting is what leaked children. ``cancel()`` only
    schedules the ``CancelledError``; the handler that marks the child done
    and writes its ONE terminal ``stop`` runs on a later turn of the event
    loop, which never comes if the loop tears down first. The child was then
    settled by nothing in this process — its span stayed open until a boot
    sweep reaped it days later, or forever. Awaiting here is what makes
    settlement in-process rather than next-boot.

    Exceptions are swallowed by ``return_exceptions``: these tasks own their
    own terminal reporting, and a manager that fails while being torn down
    must not take down the run that is already ending.

    **The set to await is not fixed at entry.** Every manager that settles
    hands its slot to a QUEUED child through the dispatch pump, which
    creates a new manager — so this DRAINS in rounds against one deadline
    rather than awaiting a snapshot. Cancelling waiting units up front
    would be the simpler code and the wrong behaviour: it would discard, at
    teardown, exactly the work this scheduler exists to stop discarding.
    Only once the budget is spent are the still-waiting units dropped —
    after that nothing will ever dispatch them, so they must be settled
    rather than left reading as live.
    """
    deadline = time.monotonic() + max(0.0, timeout)
    while True:
        pending = [t for t in self._lifecycle_tasks if not t.done()]
        remaining = deadline - time.monotonic()
        if not pending or remaining <= 0:
            break
        await asyncio.wait(
            pending,
            timeout=remaining,
            return_when=asyncio.ALL_COMPLETED,
        )

    # ROOT ONLY. The queue is SESSION-global while this method runs at the
    # end of EVERY agent's loop, root or child — so an unguarded clear here
    # let the first sub-agent to finish discard the root's still-waiting
    # fan-out, which is precisely the loss this scheduler exists to stop.
    # A depth>=1 tool owns nothing in the queue by construction: a nested
    # spawn runs inline on its parent's slot and is never enqueued.
    if self._agent_context.depth == 0:
        await self._agent_context.registry.cancel_pending()

    still_pending = [t for t in self._lifecycle_tasks if not t.done()]
    for task in still_pending:
        task.cancel()
    if still_pending:
        await asyncio.gather(*still_pending, return_exceptions=True)
    self._lifecycle_tasks.clear()

handle_check_agents(action_step: ActionStep) -> MockSpeaker async

Return agent tree state with completed results and progress.

Emits a JSON payload with kind: "agent_tree". The text field carries the rendered ASCII tree the LLM consumes; the agents list is the structured snapshot the console uses to render CheckAgentsCard.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
2183
2184
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
2204
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
2226
2227
2228
2229
2230
2231
2232
2233
2234
2235
2236
2237
2238
2239
2240
2241
2242
2243
2244
2245
2246
2247
2248
2249
2250
2251
2252
2253
2254
2255
2256
2257
2258
2259
2260
2261
2262
2263
2264
2265
2266
2267
2268
2269
2270
2271
2272
2273
2274
2275
2276
async def handle_check_agents(self, action_step: ActionStep) -> MockSpeaker:
    """Return agent tree state with completed results and progress.

    Emits a JSON payload with ``kind: "agent_tree"``. The ``text`` field
    carries the rendered ASCII tree the LLM consumes; the ``agents`` list
    is the structured snapshot the console uses to render CheckAgentsCard.
    """
    args = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
    wait = bool(args.get("wait", False))
    timeout = self._coerce_timeout(args.get("timeout"))
    registry = self._agent_context.registry
    parent_id = self._agent_context.agent_id

    if wait:
        running = await registry.collect_running(parent_id)
        if running:
            waiters = [asyncio.create_task(h.done_event.wait()) for h in running]
            # asyncio.wait() does not raise on timeout — it returns
            # (done, pending) with pending non-empty when the deadline hits.
            _done, pending = await asyncio.wait(
                waiters,
                timeout=timeout,
                return_when=asyncio.FIRST_COMPLETED,
            )
            for t in pending:
                t.cancel()

    tree = await registry.render_agent_tree(
        exclude_agent_id=self._agent_context.agent_id,
    )
    completed = await registry.collect_completed(parent_id)
    running = await registry.collect_running(parent_id)

    parts: list[str] = []
    if tree:
        parts.append(f"Agent tree:\n{tree}")

    if completed:
        parts.append("\nCompleted results:")
        for h in completed:
            r = h.result
            if r:
                parts.append(f"  [{h.agent_id[:8]}] {r.status}: {r.summary or r.content[:300]}")

    if running:
        parts.append(f"\n{len(running)} agent(s) still running.")
    elif not completed:
        parts.append("No agents spawned.")

    text = "\n".join(parts) or "No agents."

    agents_payload: list[dict[str, Any]] = []
    for h in await registry.list_visible(exclude_agent_id=self._agent_context.agent_id):
        entry: dict[str, Any] = {
            "id": h.agent_id,
            "parent_id": h.parent_id,
            "depth": h.depth,
            "task": h.task_description,
            "status": h.status,
            "steps_completed": h.steps_completed,
            "last_tool_id": h.last_tool_id,
            "progress_note": h.progress_note,
            "compaction_count": h.compaction_count,
            "attempts": h.attempts,  # Retry provenance for the console
            "result": (
                {
                    "status": h.result.status,
                    "summary": h.result.summary,
                    "content": h.result.content,
                }
                if h.result is not None
                else None
            ),
        }
        # Additive, ONLY when a real bound was declared, so a
        # contract-less child's payload carries no key. Carries LIVE
        # progress (steps/elapsed) alongside the bounded-scalar
        # ``snapshot()`` shape the attestation record reuses.
        if h.contract.enabled:
            entry["contract"] = {
                **h.contract.snapshot(),
                "steps_completed": h.steps_completed,
                "elapsed_s": time.monotonic() - h.started_at,
            }
        agents_payload.append(entry)

    payload = {
        "kind": "agent_tree",
        "text": text,
        "agents": agents_payload,
        "parent_id": parent_id,
        "wait": wait,
    }
    return MockSpeaker(content=json.dumps(payload))

handle_steer_agent(action_step: ActionStep) -> MockSpeaker async

Send a steering message to or cancel a running agent.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
2291
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
2308
2309
2310
2311
2312
2313
2314
2315
2316
2317
2318
2319
2320
2321
2322
2323
2324
2325
2326
2327
async def handle_steer_agent(self, action_step: ActionStep) -> MockSpeaker:
    """Send a steering message to or cancel a running agent."""
    args = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
    agent_id = str(args.get("agent_id", ""))
    action = str(args.get("action", ""))
    message = str(args.get("message", ""))
    registry = self._agent_context.registry

    # Resolve short prefix to full agent_id.
    handle = await registry.get(agent_id)
    if handle is None:
        all_agents = await registry.list_all()
        matches = [h for h in all_agents if h.agent_id.startswith(agent_id)]
        if len(matches) == 1:
            handle = matches[0]
            agent_id = handle.agent_id
        elif len(matches) > 1:
            return MockSpeaker(
                content=f"ERROR: Ambiguous prefix '{agent_id}' matches {len(matches)} agents.",
            )
        else:
            return MockSpeaker(
                content=f"ERROR: Agent '{agent_id}' not found.",
            )

    if action == "cancel":
        reason = await registry.cancel_agent(agent_id)
        if reason is None:
            return MockSpeaker(
                content=f"Agent {agent_id[:8]} cancelled.",
            )
        return MockSpeaker(
            content=f"Agent {agent_id[:8]} cannot cancel: {reason}",
        )
    if action == "message":
        if not message:
            return MockSpeaker(
                content="ERROR: 'message' is required when action='message'.",
            )
        reason = await registry.send_message(
            agent_id,
            f"[From parent] {message}",
        )
        if reason is None:
            return MockSpeaker(content="Message sent.")
        return MockSpeaker(content=f"Message failed: {reason}")

    return MockSpeaker(
        content=f"ERROR: Unknown action '{action}'. Use 'message' or 'cancel'.",
    )

has_live_owned_runs() -> bool async

True while any agent this session owns is still non-terminal.

The ownership index behind the promise-as-completion gate: a clean terminal declared while owned work is live is a promise, not a completion. Exposed here because the hypervisor is the only thing that knows, and the seam that must ask is the one accepting a completion claim.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
async def has_live_owned_runs(self) -> bool:
    """True while any agent this session owns is still non-terminal.

    The ownership index behind the promise-as-completion gate: a clean
    terminal declared while owned work is live is a promise, not a
    completion. Exposed here because the hypervisor is the only thing that
    knows, and the seam that must ask is the one accepting a completion
    claim.
    """
    try:
        agents = await self._agent_context.registry.list_all()
    except Exception:  # pragma: no cover - defensive; never fail a run here
        return False
    return any(handle.status in ACTIVE_STATUSES for handle in agents)

rebind_active_model(model_name: str) -> None

Re-seat this tool's parent context onto an escalated model.

The public seam for the loop's sticky model escalation. A parent that healed itself onto a rescue model must not keep spawning children onto the dead one: :meth:_resolve_model falls back to the parent context's model_name, and each child's own context is derived from it, so a stale value here re-infects the whole subtree.

AgentContext is frozen, so this REPLACES rather than mutates — and that is exactly why the seam is a method and not an attribute write. Knowing the context is a frozen dataclass is this class's business, not the loop's; reaching in to do the replace from outside couples the caller to a representation it should never have to know. Idempotent, so the loop may call it on every turn.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
def rebind_active_model(self, model_name: str) -> None:
    """Re-seat this tool's parent context onto an escalated model.

    The public seam for the loop's sticky model escalation. A parent that
    healed itself onto a rescue model must not keep spawning children onto
    the dead one: :meth:`_resolve_model` falls back to the parent context's
    ``model_name``, and each child's own context is derived from it, so a
    stale value here re-infects the whole subtree.

    ``AgentContext`` is frozen, so this REPLACES rather than mutates — and
    that is exactly why the seam is a method and not an attribute write.
    Knowing the context is a frozen dataclass is this class's business, not
    the loop's; reaching in to do the ``replace`` from outside couples the
    caller to a representation it should never have to know. Idempotent, so
    the loop may call it on every turn.
    """
    if not model_name or model_name == self._agent_context.model_name:
        return
    self._agent_context = replace(self._agent_context, model_name=model_name)

rebind_cwd(cwd: str, *, project_instructions: str | None) -> None

Re-point the workspace every FUTURE child inherits.

The seam a session-level project switch drives: from here on, a spawn that names no project of its own lands in cwd and reads project_instructions instead of the ones this tool was built with.

Already-running children are deliberately untouched. A live child's loop captured its own directory when it was built and has been resolving paths against it ever since; moving that out from under it would be a race no one owns — its containment root, its verifier's working directory and its half-finished edits would disagree about where it is. A child finishes where it started; the switch applies to the next one. That is also why a spawn resolves its :class:ChildWorkspace once at admission rather than re-reading these fields on each retry attempt.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
def rebind_cwd(self, cwd: str, *, project_instructions: str | None) -> None:
    """Re-point the workspace every FUTURE child inherits.

    The seam a session-level project switch drives: from here on, a spawn
    that names no ``project`` of its own lands in *cwd* and reads
    *project_instructions* instead of the ones this tool was built with.

    Already-running children are deliberately untouched. A live child's loop
    captured its own directory when it was built and has been resolving
    paths against it ever since; moving that out from under it would be a
    race no one owns — its containment root, its verifier's working
    directory and its half-finished edits would disagree about where it is.
    A child finishes where it started; the switch applies to the next one.
    That is also why a spawn resolves its :class:`ChildWorkspace` once at
    admission rather than re-reading these fields on each retry attempt.
    """
    self._cwd = cwd
    self._project_instructions = project_instructions

run_async(action_step: ActionStep) -> MockSpeaker async

Execute a single sub-agent. Returns the result as a MockSpeaker.

Thin wrapper over :meth:_spawn_one — the batch path (:meth:run_batch_async) shares the same core. The outcome is projected through :meth:_SpawnOutcome.report, not read off content: reading the string would throw away the refusal's status and typed cause, handing the model a bare sentence for five different failures.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
async def run_async(self, action_step: ActionStep) -> MockSpeaker:
    """Execute a single sub-agent. Returns the result as a MockSpeaker.

    Thin wrapper over :meth:`_spawn_one` — the batch path
    (:meth:`run_batch_async`) shares the same core. The outcome is
    projected through :meth:`_SpawnOutcome.report`, not read off
    ``content``: reading the string would throw away the refusal's status
    and typed cause, handing the model a bare sentence for five different
    failures.
    """
    args = (
        action_step.tool_input
        if isinstance(action_step.tool_input, dict)
        else {"task": str(action_step.tool_input)}
    )
    outcome = await self._spawn_one(args, blocking_admit=True)
    return MockSpeaker(content=outcome.report())

run_batch_async(action_step: ActionStep) -> MockSpeaker async

Fan out a batch of independent sub-agents from ONE tool call.

Pure composition over :meth:_spawn_one, in two phases. Each entry is RESOLVED first (its workspace, agent type and model settled, its handle registered, its id minted), then the whole batch is handed to the scheduler in ONE atomic acceptance: every entry that resolved is accepted, and the scheduler decides only which start now and which wait. A batch is therefore throttled by concurrency, never truncated by it. Marking the surplus rejected and discarding it would read to the caller as N agents when only max_concurrent existed.

Returns the ordered agent_ids; the model collects results via the existing check_agents. The orchestration loop is untouched.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
async def run_batch_async(self, action_step: ActionStep) -> MockSpeaker:
    """Fan out a batch of independent sub-agents from ONE tool call.

    Pure composition over :meth:`_spawn_one`, in two phases. Each entry is
    RESOLVED first (its workspace, agent type and model settled, its handle
    registered, its id minted), then the whole batch is handed to the
    scheduler in ONE atomic acceptance: every entry that resolved is
    accepted, and the scheduler decides only which start now and which
    wait. A batch is therefore throttled by concurrency, never truncated by
    it. Marking the surplus ``rejected`` and discarding it would read to
    the caller as N agents when only ``max_concurrent`` existed.

    Returns the ordered ``agent_id``s; the model collects results via the
    existing ``check_agents``. The orchestration loop is untouched.
    """
    raw = action_step.tool_input if isinstance(action_step.tool_input, dict) else {}
    raw_tasks = raw.get("tasks")
    if not isinstance(raw_tasks, list) or not raw_tasks:
        return MockSpeaker(
            content="ERROR: spawn_agents requires a non-empty 'tasks' array."
        )
    try:
        tasks = [SpawnAgentTask.model_validate(entry) for entry in raw_tasks]
    except ValidationError as exc:
        return MockSpeaker(content=f"ERROR: invalid spawn_agents task: {exc}")

    agents: list[dict[str, Any]] = []
    agent_ids: list[str | None] = []
    units: list[SpawnUnit] = []
    accepted = 0
    for idx, task in enumerate(tasks):
        outcome = await self._spawn_one(
            task.to_args(), blocking_admit=False, batch=units, batch_index=idx
        )
        agent_ids.append(outcome.agent_id)
        if outcome.agent_id is not None:
            accepted += 1
        entry: dict[str, Any] = {
            "index": idx,
            "agent_id": outcome.agent_id,
            "status": outcome.status,
            "task": task.task[:200],
        }
        if outcome.agent_id is None:
            # ADDITIVE, and load-bearing: a refused slot is otherwise the
            # only place a refusal's cause is discarded. ``code`` is the
            # machine-readable half — "no free concurrency slot" would be a
            # wrong answer that sends the caller retrying a project key
            # that will never resolve.
            entry["reason"] = outcome.content
            if outcome.code is not None:
                entry["code"] = outcome.code
        agents.append(entry)

    # ONE atomic acceptance for the whole fan-out. Nothing here can refuse;
    # the flags say only which entries got a slot immediately. Every
    # accepted entry stays ``submitted`` either way — starting now versus
    # waiting is a SCHEDULING fact, reported as a count, never as a
    # per-agent status the whole client tree would have to learn.
    decisions = await self._agent_context.registry.accept_batch(units)
    started = sum(1 for dispatched in decisions if dispatched)

    rejected = len(tasks) - accepted
    deferred = len(units) - started
    summary = f"Accepted {accepted}/{len(tasks)} agent(s)"
    if units:
        summary += (
            f" — {started} started now, {deferred} waiting for a free slot "
            "(they start automatically)"
        )
    if rejected:
        summary += f"; {rejected} refused — see each entry's 'code' and 'reason'"
    summary += ". Use check_agents to monitor progress and collect results."
    return MockSpeaker(
        content=json.dumps(
            {
                "kind": "agent_batch",
                "text": summary,
                "agents": agents,
                "agent_ids": agent_ids,
                # ``accepted`` is the name the count now deserves and the
                # one consumers read FIRST; ``spawned`` is retained at the
                # same value so an older reader is unaffected and so the
                # consumer-side ``accepted ?? spawned`` fallback keeps
                # serving events already in the store. Emitting only
                # ``spawned`` would leave every consumer permanently on its
                # fallback branch with the preferred key dead.
                "accepted": accepted,
                "spawned": accepted,
                "dispatched": started,
                "deferred": deferred,
                "rejected": rejected,
            }
        )
    )

substitute_agent_body(body: str, subs: Mapping[str, str], env: Mapping[str, str] | None = None) -> str staticmethod

Render an agent's body with plugin-generic variable substitution.

Three passes, in order:

  1. ${KEY} literal substitution from subs. Core passes SESSION_ID and CLAUDE_PLUGIN_ROOT; plugins author their prompts against these names.
  2. Bash-style ${VAR:-default} — if VAR is unset in env, the text expands to default. If VAR is set, it expands to the env value. This keeps plugin prompts self-documenting (operator override path is obvious in the source).
  3. Plain $VAR expansion as a final pass, matching :func:os.path.expandvars semantics. Unset variables remain literal so authors can spot typos at glance.

A staticmethod on this class rather than a loose module function: this tool is the only production caller, and subs/env arrive as ARGS so the renderer stays testable with a fake environment. The module-level substitute_agent_body name is a thin alias over it, so the import path plugins and tests already use is unchanged.

Source code in packages/mewbo_core/src/mewbo_core/agents/spawn_agent.py
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
@staticmethod
def substitute_agent_body(
    body: str,
    subs: Mapping[str, str],
    env: Mapping[str, str] | None = None,
) -> str:
    """Render an agent's body with plugin-generic variable substitution.

    Three passes, in order:

    1. ``${KEY}`` literal substitution from *subs*. Core passes
       ``SESSION_ID`` and ``CLAUDE_PLUGIN_ROOT``; plugins author their
       prompts against these names.
    2. Bash-style ``${VAR:-default}`` — if ``VAR`` is unset in *env*,
       the text expands to ``default``. If ``VAR`` is set, it expands
       to the env value. This keeps plugin prompts self-documenting
       (operator override path is obvious in the source).
    3. Plain ``$VAR`` expansion as a final pass, matching
       :func:`os.path.expandvars` semantics. Unset variables remain
       literal so authors can spot typos at glance.

    A ``staticmethod`` on this class rather than a loose module function:
    this tool is the only production caller, and *subs*/*env* arrive as
    ARGS so the renderer stays testable with a fake environment. The
    module-level ``substitute_agent_body`` name is a thin alias over it, so
    the import path plugins and tests already use is unchanged.
    """
    if env is None:
        env = os.environ
    # Pass 1: direct substitutions.
    for key, value in subs.items():
        body = body.replace(f"${{{key}}}", value)

    # Pass 2: bash-style ${VAR:-default}. Read from env; fall back to default.
    def _bash_default(match: re.Match[str]) -> str:
        var_name, default = match.group(1), match.group(2)
        return env.get(var_name, default)

    body = _BASH_DEFAULT_RE.sub(_bash_default, body)

    # Pass 3: plain ``$VAR`` expansion for anything still referencing env.
    # Matches ``os.path.expandvars`` semantics without touching the real
    # ``os.environ`` when a test supplies a fake *env* mapping.
    def _plain_var(match: re.Match[str]) -> str:
        var_name = match.group(1)
        return env.get(var_name, match.group(0))

    return re.sub(r"\$(\w+)", _plain_var, body)

mewbo_core.loop.planning

Prompt construction and planning helpers.

Planner

Generate action plans via LLM.

Source code in packages/mewbo_core/src/mewbo_core/loop/planning.py
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
class Planner:
    """Generate action plans via LLM."""

    def __init__(self, tool_registry: ToolRegistry | None) -> None:
        """Initialize the planner."""
        self._tool_registry = tool_registry
        self._prompt_builder = PromptBuilder(tool_registry)

    @staticmethod
    def _build_example_messages(available_tool_ids: list[str], *, mode: str) -> list[BaseMessage]:
        if mode != "plan":
            return []

        def wrap(text: str) -> str:
            return f"{EXAMPLE_TAG_OPEN}{text}{EXAMPLE_TAG_CLOSE}"

        return [
            HumanMessage(content=wrap("Turn on strip lights and heater.")),
            AIMessage(
                content=wrap(
                    get_task_master_examples(example_id=0, available_tools=available_tool_ids)
                )
            ),
            HumanMessage(content=wrap("What is the weather today?")),
            AIMessage(
                content=wrap(
                    get_task_master_examples(example_id=1, available_tools=available_tool_ids)
                )
            ),
        ]

    def generate(
        self,
        user_query: str,
        model_name: str,
        context: ContextSnapshot | None = None,
        *,
        tool_specs: list[ToolSpec] | None = None,
        mode: str = "act",
        feedback: str | None = None,
        project_instructions: str | None = None,
    ) -> Plan:
        """Generate a plan from the user query."""
        if self._tool_registry is None:
            raise ValueError("Tool registry is required for planning.")
        langfuse_handler = build_langfuse_handler(
            user_id="mewbo-task-master",
            session_id=f"planning-{os.getpid()}-{os.urandom(4).hex()}",
            trace_name="planning",
            version=get_version(),
            release=get_config_value("runtime", "envmode", default="Not Specified"),
        )
        model = build_chat_model(model_name=model_name)
        parser = PydanticOutputParser(pydantic_object=Plan)
        component_status = self._resolve_component_status()
        specs = tool_specs if tool_specs is not None else self._tool_registry.list_specs()
        if mode == "act" and tool_specs is None:
            specs = self._filter_specs_by_intent(specs, user_query)
        available_tool_ids = [spec.tool_id for spec in specs]
        system_prompt = self._prompt_builder.build(
            get_system_prompt(),
            context,
            component_status=component_status if mode == "act" else None,
            mode=mode,
            tool_specs=specs,
            include_tool_guidance=False,
            project_instructions=project_instructions,
        )
        example_messages = self._build_example_messages(available_tool_ids, mode=mode)
        if mode == "act":
            instruction = "## Generate the minimal plan for the user query"
        else:
            instruction = "## Generate a plan for the user query"

        user_prompt = "{user_query}"
        if feedback:
            user_prompt += f"\n\nPrevious plan was rejected. Feedback: {feedback}"

        prompt = ChatPromptTemplate(
            messages=[
                SystemMessage(content=system_prompt),
                *example_messages,
                HumanMessagePromptTemplate.from_template(
                    f"## Format Instructions\n{{format_instructions}}\n{instruction}\n{user_prompt}"
                ),
            ],
            partial_variables={"format_instructions": parser.get_format_instructions()},
            input_variables=["user_query"],
        )
        logging.info(
            "Generating action plan <model='{}'; user_query='{}'>",
            model_name,
            user_query,
        )
        logging.info("Input prompt token length is `{}`.", num_tokens_from_string(str(prompt)))
        config: dict[str, object] = {}
        if langfuse_handler is not None:
            config["callbacks"] = [langfuse_handler]
            metadata = getattr(langfuse_handler, "langfuse_metadata", None)
            if isinstance(metadata, dict) and metadata:
                config["metadata"] = metadata
        with langfuse_trace_span(
            "planning",
            metadata={"model": model_name, "mode": mode},
            input_data={"user_query": user_query.strip()[:200]},
        ) as span:
            action_plan = (prompt | model | parser).invoke(
                {"user_query": user_query.strip()},
                config=config or None,
            )
            if span is not None:
                try:
                    span.update_trace(output={"step_count": len(action_plan.steps or [])})
                except Exception:
                    pass
        action_plan.human_message = user_query
        return action_plan

    @staticmethod
    def _infer_intent_capabilities(user_query: str) -> set[str]:
        lowered = user_query.lower()
        requested: set[str] = set()
        for intent, keywords in INTENT_KEYWORDS.items():
            if any(keyword in lowered for keyword in keywords):
                requested |= INTENT_CAPABILITIES[intent]
        return requested

    @staticmethod
    def _spec_capabilities(spec) -> set[str]:
        metadata = spec.metadata or {}
        capabilities = metadata.get("capabilities")
        if isinstance(capabilities, list):
            return {str(item) for item in capabilities if isinstance(item, str)}

        tool_id = spec.tool_id.lower()
        inferred: set[str] = set()
        if "internet_search" in tool_id or "web_search" in tool_id or "searxng" in tool_id:
            inferred.add("web_search")
        if "web_url_read" in tool_id or "web_url" in tool_id:
            inferred.add("web_read")
        if "aider_read_file" in tool_id or "aider_list_dir" in tool_id:
            inferred.add("file_read")
        if "edit" in tool_id and ("file" in tool_id or "aider" in tool_id or "block" in tool_id):
            inferred.add("file_write")
        if "shell" in tool_id:
            inferred.add("shell_exec")
        if "home_assistant" in tool_id:
            inferred.add("home_assistant")
        return inferred

    def _filter_specs_by_intent(self, specs, user_query: str):
        requested = self._infer_intent_capabilities(user_query)
        if not requested:
            return specs
        filtered = [spec for spec in specs if self._spec_capabilities(spec).intersection(requested)]
        return filtered or specs

    def _resolve_component_status(self) -> list[ComponentStatus]:
        return [resolve_langfuse_status()]

__init__(tool_registry: ToolRegistry | None) -> None

Initialize the planner.

Source code in packages/mewbo_core/src/mewbo_core/loop/planning.py
181
182
183
184
def __init__(self, tool_registry: ToolRegistry | None) -> None:
    """Initialize the planner."""
    self._tool_registry = tool_registry
    self._prompt_builder = PromptBuilder(tool_registry)

generate(user_query: str, model_name: str, context: ContextSnapshot | None = None, *, tool_specs: list[ToolSpec] | None = None, mode: str = 'act', feedback: str | None = None, project_instructions: str | None = None) -> Plan

Generate a plan from the user query.

Source code in packages/mewbo_core/src/mewbo_core/loop/planning.py
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
def generate(
    self,
    user_query: str,
    model_name: str,
    context: ContextSnapshot | None = None,
    *,
    tool_specs: list[ToolSpec] | None = None,
    mode: str = "act",
    feedback: str | None = None,
    project_instructions: str | None = None,
) -> Plan:
    """Generate a plan from the user query."""
    if self._tool_registry is None:
        raise ValueError("Tool registry is required for planning.")
    langfuse_handler = build_langfuse_handler(
        user_id="mewbo-task-master",
        session_id=f"planning-{os.getpid()}-{os.urandom(4).hex()}",
        trace_name="planning",
        version=get_version(),
        release=get_config_value("runtime", "envmode", default="Not Specified"),
    )
    model = build_chat_model(model_name=model_name)
    parser = PydanticOutputParser(pydantic_object=Plan)
    component_status = self._resolve_component_status()
    specs = tool_specs if tool_specs is not None else self._tool_registry.list_specs()
    if mode == "act" and tool_specs is None:
        specs = self._filter_specs_by_intent(specs, user_query)
    available_tool_ids = [spec.tool_id for spec in specs]
    system_prompt = self._prompt_builder.build(
        get_system_prompt(),
        context,
        component_status=component_status if mode == "act" else None,
        mode=mode,
        tool_specs=specs,
        include_tool_guidance=False,
        project_instructions=project_instructions,
    )
    example_messages = self._build_example_messages(available_tool_ids, mode=mode)
    if mode == "act":
        instruction = "## Generate the minimal plan for the user query"
    else:
        instruction = "## Generate a plan for the user query"

    user_prompt = "{user_query}"
    if feedback:
        user_prompt += f"\n\nPrevious plan was rejected. Feedback: {feedback}"

    prompt = ChatPromptTemplate(
        messages=[
            SystemMessage(content=system_prompt),
            *example_messages,
            HumanMessagePromptTemplate.from_template(
                f"## Format Instructions\n{{format_instructions}}\n{instruction}\n{user_prompt}"
            ),
        ],
        partial_variables={"format_instructions": parser.get_format_instructions()},
        input_variables=["user_query"],
    )
    logging.info(
        "Generating action plan <model='{}'; user_query='{}'>",
        model_name,
        user_query,
    )
    logging.info("Input prompt token length is `{}`.", num_tokens_from_string(str(prompt)))
    config: dict[str, object] = {}
    if langfuse_handler is not None:
        config["callbacks"] = [langfuse_handler]
        metadata = getattr(langfuse_handler, "langfuse_metadata", None)
        if isinstance(metadata, dict) and metadata:
            config["metadata"] = metadata
    with langfuse_trace_span(
        "planning",
        metadata={"model": model_name, "mode": mode},
        input_data={"user_query": user_query.strip()[:200]},
    ) as span:
        action_plan = (prompt | model | parser).invoke(
            {"user_query": user_query.strip()},
            config=config or None,
        )
        if span is not None:
            try:
                span.update_trace(output={"step_count": len(action_plan.steps or [])})
            except Exception:
                pass
    action_plan.human_message = user_query
    return action_plan

PromptBuilder

Build system prompts with contextual sections.

Source code in packages/mewbo_core/src/mewbo_core/loop/planning.py
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
class PromptBuilder:
    """Build system prompts with contextual sections."""

    def __init__(self, tool_registry: ToolRegistry | None) -> None:
        """Initialize prompt builder dependencies."""
        self._tool_registry = tool_registry

    def build(
        self,
        base_prompt: str,
        context: ContextSnapshot | None,
        component_status: Iterable[ComponentStatus] | None = None,
        *,
        mode: str = "act",
        tool_specs=None,
        include_tool_guidance: bool = True,
        project_instructions: str | None = None,
    ) -> str:
        """Build an augmented system prompt string."""
        registry = get_prompt_registry()
        sections = [base_prompt]
        if project_instructions:
            sections.append(
                registry.render(
                    "planning.section.project_instructions",
                    project_instructions=project_instructions,
                )
            )
        if context and context.summary:
            sections.append(
                registry.render("planning.section.session_summary", summary=context.summary)
            )
        if context and context.selected_events:
            rendered = render_event_lines(context.selected_events)
            if rendered:
                sections.append(
                    registry.render(
                        "planning.section.relevant_earlier_context", rendered=rendered
                    )
                )
        if context and context.recent_events:
            rendered = render_event_lines(context.recent_events)
            if rendered:
                sections.append(
                    registry.render("planning.section.recent_conversation", rendered=rendered)
                )
        if self._tool_registry is not None:
            specs = tool_specs or self._tool_registry.list_specs()
            if specs:
                tool_lines = "\n".join(f"- {spec.tool_id}: {spec.description}" for spec in specs)
                sections.append(
                    registry.render("planning.section.available_tools", tool_lines=tool_lines)
                )
            if mode == "act":
                if include_tool_guidance:
                    tool_prompts = self._render_tool_prompts(specs, local_only=True)
                    if tool_prompts:
                        sections.append(
                            registry.render(
                                "planning.section.tool_guidance",
                                tool_prompts_joined="\n\n".join(tool_prompts),
                            )
                        )
        if component_status:
            sections.append(
                registry.render(
                    "planning.section.component_status",
                    status=format_component_status(component_status),
                )
            )
        return "\n\n".join(sections)

    @staticmethod
    def _render_tool_prompts(specs, *, local_only: bool = False) -> list[str]:
        prompts: list[str] = []
        for spec in specs:
            if not spec.prompt_path:
                continue
            if local_only and spec.kind != "local":
                continue
            try:
                tool_prompt = get_system_prompt(spec.prompt_path)
            except OSError as exc:
                logging.warning("Failed to load tool prompt for {}: {}", spec.tool_id, exc)
                continue
            if tool_prompt:
                prompts.append(tool_prompt)
        return prompts

__init__(tool_registry: ToolRegistry | None) -> None

Initialize prompt builder dependencies.

Source code in packages/mewbo_core/src/mewbo_core/loop/planning.py
91
92
93
def __init__(self, tool_registry: ToolRegistry | None) -> None:
    """Initialize prompt builder dependencies."""
    self._tool_registry = tool_registry

build(base_prompt: str, context: ContextSnapshot | None, component_status: Iterable[ComponentStatus] | None = None, *, mode: str = 'act', tool_specs=None, include_tool_guidance: bool = True, project_instructions: str | None = None) -> str

Build an augmented system prompt string.

Source code in packages/mewbo_core/src/mewbo_core/loop/planning.py
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
def build(
    self,
    base_prompt: str,
    context: ContextSnapshot | None,
    component_status: Iterable[ComponentStatus] | None = None,
    *,
    mode: str = "act",
    tool_specs=None,
    include_tool_guidance: bool = True,
    project_instructions: str | None = None,
) -> str:
    """Build an augmented system prompt string."""
    registry = get_prompt_registry()
    sections = [base_prompt]
    if project_instructions:
        sections.append(
            registry.render(
                "planning.section.project_instructions",
                project_instructions=project_instructions,
            )
        )
    if context and context.summary:
        sections.append(
            registry.render("planning.section.session_summary", summary=context.summary)
        )
    if context and context.selected_events:
        rendered = render_event_lines(context.selected_events)
        if rendered:
            sections.append(
                registry.render(
                    "planning.section.relevant_earlier_context", rendered=rendered
                )
            )
    if context and context.recent_events:
        rendered = render_event_lines(context.recent_events)
        if rendered:
            sections.append(
                registry.render("planning.section.recent_conversation", rendered=rendered)
            )
    if self._tool_registry is not None:
        specs = tool_specs or self._tool_registry.list_specs()
        if specs:
            tool_lines = "\n".join(f"- {spec.tool_id}: {spec.description}" for spec in specs)
            sections.append(
                registry.render("planning.section.available_tools", tool_lines=tool_lines)
            )
        if mode == "act":
            if include_tool_guidance:
                tool_prompts = self._render_tool_prompts(specs, local_only=True)
                if tool_prompts:
                    sections.append(
                        registry.render(
                            "planning.section.tool_guidance",
                            tool_prompts_joined="\n\n".join(tool_prompts),
                        )
                    )
    if component_status:
        sections.append(
            registry.render(
                "planning.section.component_status",
                status=format_component_status(component_status),
            )
        )
    return "\n\n".join(sections)

mewbo_core.loop.session_runtime

Shared session runtime utilities for CLI and API.

GoalRetryGate

Re-invoke a session ONCE when it ended without meeting its stated goal.

The escalation that begins where the in-band gates stop. A run that claims a clean terminal while a declared obligation is unmet is nudged inside the turn, with the context still warm and bounded to three attempts; that is strictly the cheaper correction and it stays the first line. This gate is what remains when the nudges are spent and the session has already ended unmet_goal — the case a human otherwise fixes by clicking Continue, which routinely produces the missing result in ONE further step.

Exactly one automatic attempt, ever. The ledger is the recovery transcript event the recovery path already writes: this gate stamps it with trigger="auto_goal_unmet", and refuses whenever a marker carrying that trigger is already present. Durable, restart-proof, and no new store field.

Cost ceiling — state it plainly, because nothing else in the system bounds a session's token spend (TokenBudget is compaction accounting, not a wallet): one additional run per session, whose own cost is bounded only by that run's step/wall budget. Worst case a session costs twice what it otherwise would. That is the entire price, and it cannot compound: the second run's terminal reaches this gate too, finds its own marker, and refuses.

action="continue" is load-bearing, not a preference. retry truncates the transcript back past the failed turn, which would delete the marker and make the one-shot guard structurally impossible to enforce.

Cost class: O(one record) — one digest fold plus one transcript read for the marker scan, on a terminal path, once per run.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
class GoalRetryGate:
    """Re-invoke a session ONCE when it ended without meeting its stated goal.

    The escalation that begins where the in-band gates stop. A run that claims a
    clean terminal while a declared obligation is unmet is nudged inside the
    turn, with the context still warm and bounded to three attempts; that is
    strictly the cheaper correction and it stays the first line. This gate is
    what remains when the nudges are spent and the session has already ended
    ``unmet_goal`` — the case a human otherwise fixes by clicking Continue,
    which routinely produces the missing result in ONE further step.

    **Exactly one automatic attempt, ever.** The ledger is the ``recovery``
    transcript event the recovery path already writes: this gate stamps it with
    ``trigger="auto_goal_unmet"``, and refuses whenever a marker carrying that
    trigger is already present. Durable, restart-proof, and no new store field.

    Cost ceiling — state it plainly, because nothing else in the system bounds a
    session's token spend (``TokenBudget`` is compaction accounting, not a wallet):
    **one additional run per session, whose own cost is bounded only by that
    run's step/wall budget.** Worst case a session costs twice what it otherwise
    would. That is the entire price, and it cannot compound: the second run's
    terminal reaches this gate too, finds its own marker, and refuses.

    ``action="continue"`` is load-bearing, not a preference. ``retry`` truncates
    the transcript back past the failed turn, which would delete the marker and
    make the one-shot guard structurally impossible to enforce.

    Cost class: ``O(one record)`` — one digest fold plus one transcript read for
    the marker scan, on a terminal path, once per run.
    """

    #: Stamped on the marker this gate writes and matched when refusing.
    TRIGGER: ClassVar[RecoveryTrigger] = "auto_goal_unmet"

    #: Session types this gate REFUSES, by :class:`SessionTag` session type.
    #:
    #: A wiki INDEXING job already carries three independent correction layers —
    #: finalize fails a job that produced no pages and again one that produced an
    #: empty graph, and boot-time job recovery re-runs a failed job up to its own
    #: cap. A generic re-invocation would stack a fourth attempt on the most
    #: expensive workload in the product, and it would be invisible to that cap,
    #: which counts only its own retries. The uncovered case this gate exists for
    #: is wiki QA, which shares the ``wiki`` origin but not the session type —
    #: hence the exclusion keys off the type, and is a named table so a second
    #: self-correcting job kind is one row rather than a new branch.
    #:
    #: ``wiki_act`` (the scoped-refresh act session, ``session_provenance.py``'s
    #: ``wiki:act:`` sub-kind) joins it for the identical reason: a scoped
    #: refresh already checkpoints its own stage-2 retry — ``ScopedRefreshRunner``
    #: resumes from ``stage="act"`` under the same slug-keyed ``JobRecovery`` cap
    #: wiki_index uses — and ``WikiIndexingSessionEndHook`` deliberately leaves a
    #: scoped job's non-terminal status alone (``_is_scoped_refresh`` in
    #: ``apps/mewbo_api/src/mewbo_api/wiki/jobs.py``) precisely so the RUNNER
    #: owns that job's terminality. A generic re-drive here would race that
    #: runner with a second, uncapped correction attempt on the same session.
    EXCLUDED_SESSION_TYPES: ClassVar[frozenset[str]] = frozenset({"wiki_index", "wiki_act"})

    def __init__(self, *, runtime: SessionRuntime) -> None:
        """Bind the gate to the runtime that owns the session it re-drives."""
        self._runtime = runtime

    def maybe_retry(self, session_id: str, relaunch: Callable[[str], str]) -> bool:
        """Re-drive *session_id* once if its run ended with an unmet goal.

        *relaunch* starts a fresh run on the session with the given query,
        carrying the ORIGINAL run's collaborators (tool registry, hooks,
        capability mode, cwd …) unchanged, and returns its run id — ``""`` when
        the registry refused. Returns True only when a retry was actually
        started.

        Total by construction: a best-effort correction must never break the
        completion of the run that triggered it, so every failure is logged and
        swallowed. It is called from the run's own thread after that run's
        registry slot has been released.
        """
        try:
            return self._retry(session_id, relaunch)
        except Exception as exc:  # noqa: BLE001 — a side effect at the edge
            logging.warning(
                "Automatic goal retry skipped for session {}: {}: {}",
                session_id,
                type(exc).__name__,
                exc,
            )
            return False

    def _retry(self, session_id: str, relaunch: Callable[[str], str]) -> bool:
        """Run the four checks, then drive the one re-invocation."""
        store = self._runtime.session_store
        record = store.load_session_records([session_id])[session_id]
        summary = self._runtime.summarize_session(session_id, record=record)
        if summary.get("status") != "unmet_goal" or not summary.get("recoverable"):
            return False
        excluded = self._excluded_type(record.tags)
        if excluded is not None:
            logging.info(
                "Automatic goal retry declined for session {}: {} sessions are "
                "excluded — that workload owns its own retry policy",
                session_id,
                excluded,
            )
            return False
        if self._already_triggered(session_id):
            logging.info(
                "Automatic goal retry already spent for session {}; not retrying",
                session_id,
            )
            return False

        goal = self._goal_text(summary)
        query = get_prompt_registry().render("loop.goal_unmet_retry", goal=goal)
        # Writes the marker AND yields the query in one call, so the ledger this
        # gate reads back is the same event every other recovery writes.
        query = self._runtime.resolve_recovery_query(
            session_id,
            "continue",
            replacement_text=query,
            trigger=self.TRIGGER,
        )
        self._runtime.reinject_recovery_context(session_id)
        run_id = relaunch(query)
        if not run_id:
            logging.warning(
                "Automatic goal retry for session {} was refused by the run "
                "registry; the one-shot budget is spent",
                session_id,
            )
            return False
        logging.info(
            "Automatic goal retry started for session {} as run {} (goal: {})",
            session_id,
            run_id,
            goal,
        )
        return True

    @classmethod
    def _excluded_type(cls, tags: list[str]) -> str | None:
        """Return the excluded session type *tags* names, or ``None``."""
        for tag in tags:
            parsed = SessionTag.parse(tag)
            if parsed is not None and parsed.session_type in cls.EXCLUDED_SESSION_TYPES:
                return parsed.session_type
        return None

    def _already_triggered(self, session_id: str) -> bool:
        """True when this gate's marker is already in the transcript.

        Reads the FULL transcript rather than the digest: the digest projects the
        events a summary folds, and this one is not among them, so a shorter read
        would report "never retried" for a session that had been.
        """
        for event in self._runtime.load_events(session_id):
            if event.get("type") != "recovery":
                continue
            payload = event.get("payload")
            if isinstance(payload, dict) and payload.get("trigger") == self.TRIGGER:
                return True
        return False

    @staticmethod
    def _goal_text(summary: Mapping[str, object]) -> str:
        """Name the goal that went unmet, in the most specific terms available.

        An outcome assertion's one-line ``detail`` is written by the component
        that OWNS the goal, so it is the most specific statement of it; its
        ``reason`` token is the same fact, coarser. A loop-derived terminal
        (a halt, a failed verification) has neither, and the ``done_reason``
        token is then the only honest thing to name.
        """
        for key in ("unmet_goal_detail", "unmet_goal_reason", "done_reason"):
            value = summary.get(key)
            if isinstance(value, str) and value:
                return value
        return "the task this session was started to complete"

__init__(*, runtime: SessionRuntime) -> None

Bind the gate to the runtime that owns the session it re-drives.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
413
414
415
def __init__(self, *, runtime: SessionRuntime) -> None:
    """Bind the gate to the runtime that owns the session it re-drives."""
    self._runtime = runtime

maybe_retry(session_id: str, relaunch: Callable[[str], str]) -> bool

Re-drive session_id once if its run ended with an unmet goal.

relaunch starts a fresh run on the session with the given query, carrying the ORIGINAL run's collaborators (tool registry, hooks, capability mode, cwd …) unchanged, and returns its run id — "" when the registry refused. Returns True only when a retry was actually started.

Total by construction: a best-effort correction must never break the completion of the run that triggered it, so every failure is logged and swallowed. It is called from the run's own thread after that run's registry slot has been released.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
def maybe_retry(self, session_id: str, relaunch: Callable[[str], str]) -> bool:
    """Re-drive *session_id* once if its run ended with an unmet goal.

    *relaunch* starts a fresh run on the session with the given query,
    carrying the ORIGINAL run's collaborators (tool registry, hooks,
    capability mode, cwd …) unchanged, and returns its run id — ``""`` when
    the registry refused. Returns True only when a retry was actually
    started.

    Total by construction: a best-effort correction must never break the
    completion of the run that triggered it, so every failure is logged and
    swallowed. It is called from the run's own thread after that run's
    registry slot has been released.
    """
    try:
        return self._retry(session_id, relaunch)
    except Exception as exc:  # noqa: BLE001 — a side effect at the edge
        logging.warning(
            "Automatic goal retry skipped for session {}: {}: {}",
            session_id,
            type(exc).__name__,
            exc,
        )
        return False

RunHandle dataclass

Active orchestration tracking.

A run is backed by EITHER a daemon thread (CPython sync-wrapper path) or an event-loop task (Pyodide WebLoop / async harnesses). Exactly one of thread / loop_active reflects liveness; every other field applies uniformly. is_alive() is the backend-agnostic liveness check.

task holds the loop-backed run's asyncio.Task (mirrors AgentHandle.asyncio_task in hypervisor.py / _lifecycle_tasks in spawn_agent.py): register_loop_run mints the handle before the coroutine exists, so task starts None and is attached via :meth:RunRegistry.attach_loop_task right after loop.create_task(...). Holding this strong reference on the registry-owned handle is what keeps asyncio from garbage-collecting the fire-and-forget task mid-run — a bare local task = loop.create_task(...) with no held reference is eligible for GC as soon as the enclosing function returns.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
@dataclass
class RunHandle:
    """Active orchestration tracking.

    A run is backed by EITHER a daemon thread (CPython sync-wrapper path) or an
    event-loop task (Pyodide WebLoop / async harnesses). Exactly one of
    ``thread`` / ``loop_active`` reflects liveness; every other field applies
    uniformly. ``is_alive()`` is the backend-agnostic liveness check.

    ``task`` holds the loop-backed run's ``asyncio.Task`` (mirrors
    ``AgentHandle.asyncio_task`` in ``hypervisor.py`` / ``_lifecycle_tasks`` in
    ``spawn_agent.py``): ``register_loop_run`` mints the handle before the
    coroutine exists, so ``task`` starts ``None`` and is attached via
    :meth:`RunRegistry.attach_loop_task` right after ``loop.create_task(...)``.
    Holding this strong reference on the registry-owned handle is what keeps
    asyncio from garbage-collecting the fire-and-forget task mid-run — a bare
    local ``task = loop.create_task(...)`` with no held reference is eligible
    for GC as soon as the enclosing function returns.
    """

    cancel_event: threading.Event
    started_at: str
    thread: threading.Thread | None = field(default=None)
    loop_active: bool = field(default=False)
    message_queue: queue.Queue[str] | None = field(default=None)
    interrupt_step: threading.Event | None = field(default=None)
    task: asyncio.Task | None = field(default=None)

    def is_alive(self) -> bool:
        """Return whether the underlying run is still in progress."""
        if self.thread is not None:
            return self.thread.is_alive()
        return self.loop_active

is_alive() -> bool

Return whether the underlying run is still in progress.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
196
197
198
199
200
def is_alive(self) -> bool:
    """Return whether the underlying run is still in progress."""
    if self.thread is not None:
        return self.thread.is_alive()
    return self.loop_active

RunRegistry

Track active orchestration runs (thread- or loop-backed) per session.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
class RunRegistry:
    """Track active orchestration runs (thread- or loop-backed) per session."""

    def __init__(self) -> None:
        """Initialize the run registry."""
        self._lock = threading.Lock()
        self._runs: dict[str, RunHandle] = {}

    def start(
        self,
        session_id: str,
        target: Callable[[threading.Event], None],
        *,
        message_queue: queue.Queue[str] | None = None,
        interrupt_step: threading.Event | None = None,
        on_release: Callable[[], None] | None = None,
    ) -> bool:
        """Start a new thread-backed run for the session if not already active.

        *on_release* fires once, on the run's own thread, AFTER the handle has
        been dropped — so the session's slot is genuinely free and a callback
        may start a follow-up run on it. Firing it any earlier (from the run's
        own ``finally``, say) cannot work: this registry still holds the slot
        there, so a same-session ``start_async`` refuses and returns ``""``.
        """
        with self._lock:
            existing = self._runs.get(session_id)
            if existing and existing.is_alive():
                return False
            cancel_event = threading.Event()
            thread = threading.Thread(
                target=self._wrap_run,
                args=(session_id, cancel_event, target, on_release),
                daemon=True,
            )
            self._runs[session_id] = RunHandle(
                thread=thread,
                cancel_event=cancel_event,
                started_at=_utc_now(),
                message_queue=message_queue,
                interrupt_step=interrupt_step,
            )
            thread.start()
            return True

    def register_loop_run(
        self,
        session_id: str,
        *,
        cancel_event: threading.Event,
        message_queue: queue.Queue[str] | None = None,
        interrupt_step: threading.Event | None = None,
    ) -> bool:
        """Register a loop-backed run (no thread). Mirror of :meth:`start`.

        Used when the caller is already inside a running event loop (Pyodide's
        WebLoop) and drives the orchestration as an ``asyncio`` task rather than
        a daemon thread. Returns ``False`` if a live run already exists for the
        session. The caller must invoke :meth:`finalize_loop_run` from the
        task's done-callback so the registry stays consistent.
        """
        with self._lock:
            existing = self._runs.get(session_id)
            if existing and existing.is_alive():
                return False
            self._runs[session_id] = RunHandle(
                cancel_event=cancel_event,
                started_at=_utc_now(),
                loop_active=True,
                message_queue=message_queue,
                interrupt_step=interrupt_step,
            )
            return True

    def attach_loop_task(self, session_id: str, task: asyncio.Task) -> None:
        """Attach a loop-backed run's ``asyncio.Task`` to its handle.

        ``register_loop_run`` reserves the run's slot before the coroutine —
        and thus the task — exists, so the task is attached here right after
        the caller's ``loop.create_task(...)``. From that point on the
        registry (via the handle) holds a strong reference for the run's
        lifetime, preventing the fire-and-forget task from being garbage
        collected mid-run. A no-op if the handle is already gone (e.g. the
        task finished and its done-callback finalized the run before this
        call — not expected in practice but harmless).
        """
        with self._lock:
            handle = self._runs.get(session_id)
            if handle is not None:
                handle.task = task

    def finalize_loop_run(self, session_id: str) -> None:
        """Mark a loop-backed run as completed and drop its handle."""
        with self._lock:
            handle = self._runs.get(session_id)
            if handle is not None and handle.thread is None:
                self._runs.pop(session_id, None)

    def _wrap_run(
        self,
        session_id: str,
        cancel_event: threading.Event,
        target: Callable[[threading.Event], None],
        on_release: Callable[[], None] | None = None,
    ) -> None:
        try:
            target(cancel_event)
        finally:
            with self._lock:
                handle = self._runs.get(session_id)
                current_ident = threading.current_thread().ident
                if (
                    handle is not None
                    and handle.thread is not None
                    and handle.thread.ident == current_ident
                ):
                    self._runs.pop(session_id, None)
            # Post-release, and INSIDE the finally so a run that raised still
            # reaches it — the callback's whole job is to look at the terminal
            # the run left behind, and a crash is a terminal too. It owns its
            # own failure isolation; a raise here would replace the run's
            # exception with the callback's, which is why nothing above depends
            # on it returning.
            if on_release is not None:
                on_release()

    def cancel(self, session_id: str) -> bool:
        """Request cancellation for an active session run."""
        with self._lock:
            handle = self._runs.get(session_id)
            if not handle:
                return False
            handle.cancel_event.set()
            return True

    def is_running(self, session_id: str) -> bool:
        """Return True if the session has an active run."""
        with self._lock:
            handle = self._runs.get(session_id)
            return bool(handle and handle.is_alive())

    def get_cancel_event(self, session_id: str) -> threading.Event | None:
        """Return the cancel event for a session, if present."""
        with self._lock:
            handle = self._runs.get(session_id)
            return handle.cancel_event if handle else None

    def get_handle(self, session_id: str) -> RunHandle | None:
        """Return the run handle for a session, if present."""
        with self._lock:
            return self._runs.get(session_id)

__init__() -> None

Initialize the run registry.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
206
207
208
209
def __init__(self) -> None:
    """Initialize the run registry."""
    self._lock = threading.Lock()
    self._runs: dict[str, RunHandle] = {}

attach_loop_task(session_id: str, task: asyncio.Task) -> None

Attach a loop-backed run's asyncio.Task to its handle.

register_loop_run reserves the run's slot before the coroutine — and thus the task — exists, so the task is attached here right after the caller's loop.create_task(...). From that point on the registry (via the handle) holds a strong reference for the run's lifetime, preventing the fire-and-forget task from being garbage collected mid-run. A no-op if the handle is already gone (e.g. the task finished and its done-callback finalized the run before this call — not expected in practice but harmless).

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
def attach_loop_task(self, session_id: str, task: asyncio.Task) -> None:
    """Attach a loop-backed run's ``asyncio.Task`` to its handle.

    ``register_loop_run`` reserves the run's slot before the coroutine —
    and thus the task — exists, so the task is attached here right after
    the caller's ``loop.create_task(...)``. From that point on the
    registry (via the handle) holds a strong reference for the run's
    lifetime, preventing the fire-and-forget task from being garbage
    collected mid-run. A no-op if the handle is already gone (e.g. the
    task finished and its done-callback finalized the run before this
    call — not expected in practice but harmless).
    """
    with self._lock:
        handle = self._runs.get(session_id)
        if handle is not None:
            handle.task = task

cancel(session_id: str) -> bool

Request cancellation for an active session run.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
329
330
331
332
333
334
335
336
def cancel(self, session_id: str) -> bool:
    """Request cancellation for an active session run."""
    with self._lock:
        handle = self._runs.get(session_id)
        if not handle:
            return False
        handle.cancel_event.set()
        return True

finalize_loop_run(session_id: str) -> None

Mark a loop-backed run as completed and drop its handle.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
294
295
296
297
298
299
def finalize_loop_run(self, session_id: str) -> None:
    """Mark a loop-backed run as completed and drop its handle."""
    with self._lock:
        handle = self._runs.get(session_id)
        if handle is not None and handle.thread is None:
            self._runs.pop(session_id, None)

get_cancel_event(session_id: str) -> threading.Event | None

Return the cancel event for a session, if present.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
344
345
346
347
348
def get_cancel_event(self, session_id: str) -> threading.Event | None:
    """Return the cancel event for a session, if present."""
    with self._lock:
        handle = self._runs.get(session_id)
        return handle.cancel_event if handle else None

get_handle(session_id: str) -> RunHandle | None

Return the run handle for a session, if present.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
350
351
352
353
def get_handle(self, session_id: str) -> RunHandle | None:
    """Return the run handle for a session, if present."""
    with self._lock:
        return self._runs.get(session_id)

is_running(session_id: str) -> bool

Return True if the session has an active run.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
338
339
340
341
342
def is_running(self, session_id: str) -> bool:
    """Return True if the session has an active run."""
    with self._lock:
        handle = self._runs.get(session_id)
        return bool(handle and handle.is_alive())

register_loop_run(session_id: str, *, cancel_event: threading.Event, message_queue: queue.Queue[str] | None = None, interrupt_step: threading.Event | None = None) -> bool

Register a loop-backed run (no thread). Mirror of :meth:start.

Used when the caller is already inside a running event loop (Pyodide's WebLoop) and drives the orchestration as an asyncio task rather than a daemon thread. Returns False if a live run already exists for the session. The caller must invoke :meth:finalize_loop_run from the task's done-callback so the registry stays consistent.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
def register_loop_run(
    self,
    session_id: str,
    *,
    cancel_event: threading.Event,
    message_queue: queue.Queue[str] | None = None,
    interrupt_step: threading.Event | None = None,
) -> bool:
    """Register a loop-backed run (no thread). Mirror of :meth:`start`.

    Used when the caller is already inside a running event loop (Pyodide's
    WebLoop) and drives the orchestration as an ``asyncio`` task rather than
    a daemon thread. Returns ``False`` if a live run already exists for the
    session. The caller must invoke :meth:`finalize_loop_run` from the
    task's done-callback so the registry stays consistent.
    """
    with self._lock:
        existing = self._runs.get(session_id)
        if existing and existing.is_alive():
            return False
        self._runs[session_id] = RunHandle(
            cancel_event=cancel_event,
            started_at=_utc_now(),
            loop_active=True,
            message_queue=message_queue,
            interrupt_step=interrupt_step,
        )
        return True

start(session_id: str, target: Callable[[threading.Event], None], *, message_queue: queue.Queue[str] | None = None, interrupt_step: threading.Event | None = None, on_release: Callable[[], None] | None = None) -> bool

Start a new thread-backed run for the session if not already active.

on_release fires once, on the run's own thread, AFTER the handle has been dropped — so the session's slot is genuinely free and a callback may start a follow-up run on it. Firing it any earlier (from the run's own finally, say) cannot work: this registry still holds the slot there, so a same-session start_async refuses and returns "".

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
def start(
    self,
    session_id: str,
    target: Callable[[threading.Event], None],
    *,
    message_queue: queue.Queue[str] | None = None,
    interrupt_step: threading.Event | None = None,
    on_release: Callable[[], None] | None = None,
) -> bool:
    """Start a new thread-backed run for the session if not already active.

    *on_release* fires once, on the run's own thread, AFTER the handle has
    been dropped — so the session's slot is genuinely free and a callback
    may start a follow-up run on it. Firing it any earlier (from the run's
    own ``finally``, say) cannot work: this registry still holds the slot
    there, so a same-session ``start_async`` refuses and returns ``""``.
    """
    with self._lock:
        existing = self._runs.get(session_id)
        if existing and existing.is_alive():
            return False
        cancel_event = threading.Event()
        thread = threading.Thread(
            target=self._wrap_run,
            args=(session_id, cancel_event, target, on_release),
            daemon=True,
        )
        self._runs[session_id] = RunHandle(
            thread=thread,
            cancel_event=cancel_event,
            started_at=_utc_now(),
            message_queue=message_queue,
            interrupt_step=interrupt_step,
        )
        thread.start()
        return True

SessionRuntime

Shared orchestration runtime surface for CLI and API.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
 533
 534
 535
 536
 537
 538
 539
 540
 541
 542
 543
 544
 545
 546
 547
 548
 549
 550
 551
 552
 553
 554
 555
 556
 557
 558
 559
 560
 561
 562
 563
 564
 565
 566
 567
 568
 569
 570
 571
 572
 573
 574
 575
 576
 577
 578
 579
 580
 581
 582
 583
 584
 585
 586
 587
 588
 589
 590
 591
 592
 593
 594
 595
 596
 597
 598
 599
 600
 601
 602
 603
 604
 605
 606
 607
 608
 609
 610
 611
 612
 613
 614
 615
 616
 617
 618
 619
 620
 621
 622
 623
 624
 625
 626
 627
 628
 629
 630
 631
 632
 633
 634
 635
 636
 637
 638
 639
 640
 641
 642
 643
 644
 645
 646
 647
 648
 649
 650
 651
 652
 653
 654
 655
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
class SessionRuntime:
    """Shared orchestration runtime surface for CLI and API."""

    def __init__(
        self,
        *,
        session_store: SessionStoreBase | None = None,
        run_registry: RunRegistry | None = None,
        goal_retry_gate: GoalRetryGate | None = None,
    ) -> None:
        """Initialize the runtime with session storage and optional run registry."""
        self._session_store = session_store or create_session_store()
        self._run_registry = run_registry or RunRegistry()
        # The one-shot re-invocation of a session that ended without meeting its
        # goal. Injected so a caller can substitute or disable it; the default
        # binds to this runtime, which owns every seam the gate needs.
        self._goal_retry_gate = goal_retry_gate or GoalRetryGate(runtime=self)
        # Callbacks fired once when a session is permanently terminated. The
        # seam a Wave-2 trigger service registers on to cascade-cancel a
        # session's scheduled triggers; empty by default so termination has no
        # extra side effects until something opts in.
        self._on_terminate: list[Callable[[str], int | None]] = []

    @property
    def session_store(self) -> SessionStoreBase:
        """Expose the underlying session store."""
        return self._session_store

    def resolve_session(
        self,
        *,
        session_id: str | None = None,
        session_tag: str | None = None,
        fork_from: str | None = None,
        fork_at_ts: str | None = None,
        owner: str | None = None,
    ) -> str:
        """Resolve session identifiers, tags, and forks to a session id.

        When *fork_at_ts* is provided alongside *fork_from*, only events up to
        (and including) that timestamp are copied into the new session.

        *owner* stamps whichever NEW session this call mints — a fresh one or a
        fork. It is an opaque subject string; core never learns what a principal
        is (see ``SessionStoreBase.create_session``). Resolving to an EXISTING
        session ignores it: ownership is established once, at creation, so
        re-engaging someone else's session can never quietly re-stamp it.
        """
        if fork_from:
            source_session_id = self._session_store.resolve_tag(fork_from) or fork_from
            if self.is_terminated(source_session_id):
                raise SessionTerminatedError(
                    f"session {source_session_id} is permanently terminated; "
                    "a terminated session cannot be forked"
                )
            if fork_at_ts:
                session_id = self._session_store.fork_session_at(
                    source_session_id, fork_at_ts, owner
                )
            else:
                session_id = self._session_store.fork_session(source_session_id, owner)
        if session_tag and not session_id:
            resolved = self._session_store.resolve_tag(session_tag)
            session_id = resolved if resolved else None
        if not session_id:
            session_id = self._session_store.create_session(owner)
        if session_tag:
            self._session_store.tag_session(session_id, session_tag)
        assert session_id is not None
        return session_id

    def ensure_session(self, session_id: str) -> None:
        """Idempotently materialise a session record for a pre-minted id.

        Thin delegate to the store's ``ensure_session``. A caller that minted a
        ``session_id`` outside ``resolve_session`` (e.g. the realtime recorder,
        which pre-mints to open a Langfuse trace before any store write) calls
        this so the session is a real RECORD — visible to ``list_sessions`` and
        every read surface — not an orphan transcript.
        """
        self._session_store.ensure_session(session_id)

    def append_context_event(self, session_id: str, context: dict[str, object]) -> None:
        """Append a context event to the session transcript."""
        if not context:
            return
        self._session_store.append_event(session_id, {"type": "context", "payload": context})

    def tag_session(self, session_id: str, tag: str) -> None:
        """Associate a provenance/lookup tag with a session.

        Thin delegate to the store so callers that already hold a resolved
        ``session_id`` (e.g. a structured/realtime run stamping its origin tag)
        don't reach into ``session_store`` directly. ``resolve_session`` remains
        the seam for tag-keyed *resolution*; this is the write-only sibling for
        tagging a session you've already created.
        """
        self._session_store.tag_session(session_id, tag)

    def append_event(self, session_id: str, event: dict[str, object]) -> None:
        """Append a raw transcript event verbatim.

        Unlike :meth:`append_context_event` (which wraps payloads as
        ``{"type": "context", ...}``), this writes the event as-is, so a
        ``completion`` event reaches :meth:`summarize_session` — the single
        status authority — instead of being hidden inside a context payload.
        """
        self._session_store.append_event(session_id, event)

    def summarize_session(
        self,
        session_id: str,
        *,
        events: list[EventRecord] | None = None,
        record: SessionRecord | None = None,
    ) -> dict[str, object]:
        """Return a summarized view of a session.

        *record* supplies the session's stored metadata (title, archived,
        terminated, owner, tags) when the caller has already batch-loaded it for
        a whole page — see :meth:`list_sessions`. Omitting it reads the same
        five facts one at a time, which is what every single-session caller
        does and what this method has always done, so the derived summary is
        identical either way; only the number of store reads differs.

        With no *events*, the fold's input is the store's DIGEST of the session
        rather than its whole transcript — the same projection a listing row
        gets, and identical in result. It
        matters most on the poll path, which re-derives status once a second per
        open client: on the largest live session that read was 0.211 s of a
        0.470 s poll, for a status that had not changed.
        """
        if events is None:
            events = self._session_store.session_digest(session_id).events
        if record is None:
            record = self._session_store.load_session_records([session_id])[session_id]
        created_at = events[0]["ts"] if events else None
        stored_title = record.title
        title = stored_title
        status: SessionStatus = "idle"
        done_reason = None
        blocked_code: str | None = None
        has_user_event = False
        # Evidence for the model-attributable-failure projection, gathered in
        # this SAME pass — ``list_sessions`` calls this per session, so a second
        # walk of every transcript is a cost with no new information.
        failure_reason: str | None = None
        models_tried: list[str] = []
        outcome_assertion: OutcomeAssertion | None = None
        # Two more facts gathered in that same single pass, for the same reason:
        # how many lines the session changed, and which gated capabilities it
        # actually exercised (as opposed to which ones its client advertised).
        diff_stat = DiffStat()
        evidence = CapabilityEvidence()
        for event in events:
            evidence.observe(event)
            if event.get("type") == "tool_result":
                result_payload = event.get("payload", {})
                if isinstance(result_payload, dict):
                    diff_stat = diff_stat + DiffStat.from_tool_result(result_payload)
            if event.get("type") == "user":
                has_user_event = True
                if title is None:
                    payload = event.get("payload", {})
                    if isinstance(payload, dict):
                        raw = payload.get("text")
                        if isinstance(raw, str):
                            title = raw[:120]
            if event.get("type") in RESILIENCE_EVENT_TYPES:
                failure_reason = self._read_resilience_event(
                    event, models_tried, failure_reason
                )
            if event.get("type") == "outcome_assertion":
                assertion = self._read_outcome_assertion(event)
                if assertion is not None:
                    outcome_assertion = assertion
            if event.get("type") == "completion":
                payload = event.get("payload", {})
                if isinstance(payload, dict):
                    # A NEW terminal supersedes any assertion made about the
                    # PREVIOUS one. A session hosts many turns, and turn 5's
                    # status must not inherit turn 1's unmet purpose — the
                    # assertion is appended right after the terminal it
                    # describes, so append order alone scopes it correctly.
                    outcome_assertion = None
                    # EXACTLY ONE terminal event governs the derived state, and
                    # every field below is re-read from THIS payload. A session
                    # can carry two contradictory terminals — a boot sweep
                    # stamps a synthetic "interrupted" completion for a run it
                    # judged dead, and the run, still alive, appends its real
                    # one minutes later — so deriving field-by-field across
                    # events would blend the sweep's reason into the live run's
                    # status. Last terminal wins, whole.
                    done_reason = payload.get("done_reason")
                    status, blocked_code = self._completion_status(payload)
                    self._read_models_tried(payload, models_tried)
        # An outcome assertion is the ONLY signal that can contradict a terminal
        # the loop itself considers clean, so it is applied against the derived
        # status rather than folded into the reason table. It promotes ONLY a
        # claim of success: `failed`/`canceled` are already honest, and `blocked`
        # is both more specific and more actionable, so none of them is
        # overwritten by the coarser "purpose not met".
        if outcome_assertion is not None and status == "completed":
            status = "unmet_goal"
        running = self.is_running(session_id)
        if running:
            status = "running"
        # Permanent termination is the terminal-most state — it wins over
        # EVERYTHING, including a still-unwinding live run (cancellation is
        # cooperative, so ``is_running`` can briefly lag a terminate). Checked
        # AFTER the running override so ``terminated`` is never masked.
        terminated_at = record.terminated_at
        terminated = terminated_at is not None
        if terminated:
            status = "terminated"
        if not has_user_event and not running:
            created_at = None
        if not title:
            title = f"Session {session_id[:8]}"
        merged_context = self._session_store.merge_context_events(events)
        origin = SessionOrigin.classify(
            record.tags, merged_context
        )
        # ``recoverable`` = the FE/CLI may offer a Continue/Restart affordance.
        # True when the session is not running, did not complete successfully,
        # and has a prior user turn (so ``resolve_recovery_query`` won't raise).
        # The crucial case is a session that died mid-call with NO ``completion``
        # event at all (process killed) — status stays ``idle`` but a user turn
        # exists, so it must be recoverable.
        # ``awaiting_approval`` is recoverable: a plan-mode proposal (or a wiki QA
        # answer) parks here with no active run and no auto-exit. Recovery
        # ``continue`` re-engages it through the one loop, mirroring send_followup —
        # the only other way out. Only ``completed`` is genuinely terminal.
        # This is why the status vocabulary is where honesty has to be fixed
        # rather than the recovery rule: ``unmet_goal`` and ``blocked`` earn the
        # affordance purely by not being ``completed``, so every run the old
        # table laundered into success had lost it silently.
        recoverable = not running and status != "completed" and has_user_event
        # A terminated session is a hard dead-end — never offer Continue/Restart.
        if terminated:
            recoverable = False
        # Surface the EXERCISED capabilities + the workspace so the landing page
        # can show what a session actually did without re-reading the transcript.
        # Copying ``context.client_capabilities`` verbatim would be wrong: that
        # field is an ADVERTISEMENT, and the console sends the same fixed header
        # on every request, so every chat row would be chipped with four
        # capabilities it never touched.
        # ``CapabilityEvidence`` (fed above, in the one pass) holds each earnable
        # capability to a durable artifact event or a successful invocation of a
        # tool it gates, and passes the server-written scopes (``wiki``/``scg``)
        # through untouched — see its docstring for why the split falls there. A
        # runtime-granted capability is still intentionally NOT probed: it's a
        # live predicate, not a durable signal, already surfaced where it matters
        # (the run's Langfuse trace facet), and a per-row store probe in a session
        # LIST would be the wrong tradeoff.
        capabilities = evidence.resolve(merged_context.get("client_capabilities"))
        workspace = merged_context.get("structured_workspace") or merged_context.get(
            "workspace"
        )
        summary: dict[str, object] = {
            "session_id": session_id,
            "title": title,
            "created_at": created_at,
            "status": status,
            "done_reason": done_reason,
            "running": running,
            "recoverable": recoverable,
            "context": merged_context,
            "origin": origin.value,
            "capabilities": capabilities,
            "workspace": str(workspace) if workspace else None,
            "archived": record.archived,
            "terminated": terminated,
            "terminated_at": terminated_at,
        }
        # ``owner`` is APPENDED, and only when there IS one. An unowned session
        # is genuinely different from one owned by nobody, and reporting the
        # difference as an absent key rather than an explicit null is what keeps
        # this summary key-free for a deployment running without identity: no
        # principal exists there, so nothing is ever stamped, so the key never
        # appears. Appending (never inserting mid-dict) preserves key order too.
        owner = record.owner
        if owner is not None:
            summary["owner"] = owner
        # ``pinned`` follows the same append-when-present rule, and for a third
        # reason on top of the two above: pinning is a minority state, so an
        # unpinned session's summary carries neither key. Both keys travel
        # together — the boolean is what
        # a surface renders, the stamp is what it orders by — so a client never
        # has to infer one from the other. Read off *record* rather than a
        # direct store call — the whole point of batch-loading it is that a
        # listing's per-row cost stops scaling with row count; calling back into
        # the store here would reopen exactly that per-row round trip for these
        # two fields alone while every other field on the row stayed batched.
        if record.pinned_at is not None:
            summary["pinned"] = True
            summary["pinned_at"] = record.pinned_at
        # ``projects`` is the accumulated SET a project filter needs — never just
        # the current ``context.project`` — because an auto-select session may
        # have switched mid-task, and a filter reading only the latest binding
        # would miss every project it moved out of. Same append-when-present
        # rule: a session bound to nothing (a bare temp dir, or still on the
        # ``auto`` sentinel) carries no key at all.
        if record.projects:
            summary["projects"] = record.projects
        # The three failure facets follow the same append-when-present rule, for
        # the same reason: a session that never blocked and never re-tried a
        # model carries none of them.
        #
        # ``blocked_code`` says WHICH wall the run hit — the status alone says a
        # user can act, not what to act on. ``failure_reason`` + ``models_tried``
        # are what let a recovery affordance default AWAY from the model that
        # just failed instead of re-running the same one, which is the whole
        # point of recording them: retrying on a model whose failure was
        # model-specific amplifies the failure.
        if blocked_code is not None:
            summary["blocked_code"] = blocked_code
        if outcome_assertion is not None:
            # Surfaced whenever an assertion was made, even where it did not
            # change the status: a run that failed AND missed its purpose is
            # two facts, and dropping the second would hide the one a product
            # owner can act on.
            summary["unmet_goal_reason"] = outcome_assertion.reason
            if outcome_assertion.detail:
                summary["unmet_goal_detail"] = outcome_assertion.detail
        if failure_reason is not None:
            summary["failure_reason"] = failure_reason
        if models_tried:
            summary["models_tried"] = models_tried
        # Same append-when-present rule: a session that changed no lines carries
        # no ``diff_stat`` at all, so a consumer can treat the key's presence as
        # "this session edited something" without comparing against zero.
        if not diff_stat.is_empty:
            summary["diff_stat"] = {
                "additions": diff_stat.additions,
                "deletions": diff_stat.deletions,
            }
        return summary

    def _completion_status(
        self, payload: Mapping[str, object]
    ) -> tuple[SessionStatus, str | None]:
        """Derive ``(status, blocked_code)`` from ONE completion payload.

        Takes a ``Mapping`` because the caller holds an ``EventPayload`` — a
        union of the payload TypedDicts and the generic JSON dict. Every member
        satisfies ``Mapping[str, object]``, so the contract is accepted whole
        rather than the caller having to widen (or cast) at the call site.

        ``blocked`` outranks every reason-derived status. An unrecovered
        repo-access / network / permission / quota envelope is both the more
        SPECIFIC fact about why the run stopped and the only one a user can act
        on, so it must not be masked by the coarser reason underneath it.

        A reason absent from :data:`_STATUS_BY_DONE_REASON` falls back to
        trusting ``done``.
        """
        raw_code = payload.get("blocked_code")
        if isinstance(raw_code, str) and raw_code in BLOCKED_CODES:
            return "blocked", raw_code
        done_reason = payload.get("done_reason")
        if isinstance(done_reason, str):
            mapped = _STATUS_BY_DONE_REASON.get(done_reason)
            if mapped is not None:
                return mapped, None
            if done_reason.startswith("command_failed:"):
                return "failed", None
        return ("completed" if payload.get("done") else "incomplete"), None

    @staticmethod
    def _read_resilience_event(
        event: EventRecord, models_tried: list[str], current_reason: str | None
    ) -> str | None:
        """Fold one ``llm_retry``/``llm_fallback`` event into the failure facets.

        Appends any newly-named model to *models_tried* in attempt order and
        returns the reason token to carry forward.

        Precedence is deliberate: ``llm_fallback.reason`` IS the
        ``RetryStrategy.classify`` reason token (or ``retries_exhausted`` when a
        transient error simply spent the per-model cap), so it is preferred
        whenever a switch occurred. ``llm_retry`` carries only the exception
        CLASS — but a run that exhausted its retries on a single model never
        emits a fallback at all, which is precisely the shape that burned five
        attempts on one model, so dropping that case would blind the projection
        to the failure it most needs to describe.
        """
        payload = event.get("payload")
        if not isinstance(payload, dict):
            return current_reason
        if event.get("type") == "llm_fallback":
            for key in ("from_model", "to_model"):
                name = payload.get(key)
                if isinstance(name, str) and name and name not in models_tried:
                    models_tried.append(name)
            candidate = payload.get("reason")
        else:
            name = payload.get("model")
            if isinstance(name, str) and name and name not in models_tried:
                models_tried.append(name)
            candidate = payload.get("error_type")
        if isinstance(candidate, str) and candidate:
            return candidate
        return current_reason

    @staticmethod
    def _read_outcome_assertion(event: EventRecord) -> OutcomeAssertion | None:
        """Validate a persisted outcome-assertion event, or ignore it.

        A stored document is a trust boundary, so it is parsed through the model
        rather than read key-by-key. Total by design: a malformed record must not
        break the status derivation for the whole session — a status that cannot
        be computed is strictly worse than one missing an assertion.
        """
        payload = event.get("payload")
        if not isinstance(payload, dict):
            return None
        try:
            return OutcomeAssertion.model_validate(payload)
        except ValidationError:
            logging.warning("Ignoring malformed outcome_assertion event payload")
            return None

    @staticmethod
    def _read_models_tried(payload: Mapping[str, object], models_tried: list[str]) -> None:
        """Fold a completion's recorded provider list into *models_tried*.

        ``RunError.provider`` is populated from ``LlmResilienceExhausted``'s own
        ``models_tried`` (comma-joined), so this is the ONLY evidence for a run
        that died without ever emitting a retry or fallback event — a fatal
        first attempt names no model anywhere else.
        """
        detail = payload.get("error_detail")
        if not isinstance(detail, dict):
            return
        provider = detail.get("provider")
        if not isinstance(provider, str):
            return
        for name in provider.split(","):
            cleaned = name.strip()
            if cleaned and cleaned not in models_tried:
                models_tried.append(cleaned)

    def list_sessions(
        self,
        query: SessionQuery | None = None,
        *,
        limit: int | None = None,
        offset: int = 0,
    ) -> list[dict[str, object]]:
        """List sessions with summary metadata, narrowed by *query*.

        ``O(collection)`` in rows, and never ``O(all history)`` in reads. TWO
        batched store calls serve the whole page, and neither scales with how
        much any session has recorded:

        * :meth:`~SessionStoreBase.list_session_digests` returns each session's
          transcript reduced to the events a summary folds. A per-id
          ``load_transcript`` here would make a listing read every event ever
          stored, and degrade monotonically forever.
        * ``load_session_records`` returns the metadata that is stored ON the
          session rather than folded from events (title, archived, terminated,
          owner, tags, ``pinned_at``, ``projects``) in one call for the page
          instead of per-row reads.

        **`query` narrows the digest fetch itself, and that ORDER is what makes
        it compose with a page.** The whole query goes to
        ``list_session_digests`` (still called EXACTLY ONCE per listing, which a
        test pins), so every predicate answerable from the session record —
        ``owner``/``archived``/``pinned``/``projects`` — is applied by the store
        before it decides which candidates to page, and a session the query
        rejects is never opened. Filtering afterwards instead would be worse
        than merely slower once ``limit`` exists: paging first and filtering the
        page would make ``pinned=True`` return only the pinned sessions inside
        the newest N candidates, which is usually none of them.

        **The fold itself is unchanged and still single-sourced.**
        ``summarize_session`` sees a smaller event list, not a different
        derivation — re-deriving a row's fields with a store-side aggregation
        would be a second implementation, free to drift until a row disagrees
        with the session it names.

        **``limit``/``offset`` bound how many CANDIDATES this call examines, not
        how many rows it is guaranteed to return.** They page
        ``list_session_digests`` after *query* has narrowed the candidate set but
        before the two VISIBILITY rules below run, which is what lets a
        Mongo-backed store scope its expensive read to the page instead of the
        whole collection (see that method's docstring for the measured saving) —
        but a candidate dropped below (no visible event and not running, or no
        ``created_at``) shrinks the page rather than being backfilled from the
        next one. Those two rules stay here because they are the only ones that
        need the transcript's EVENTS, so no store query can decide them; they
        are hygiene rather than user filters, which is why shrinkage is
        acceptable for them and would not have been for ``pinned``.
        ``limit=None`` (the default) is an unpaginated call.

        Ordering is **pinned first, then newest first.** Two stable sorts
        rather than one composite key: the second pass only has to move pinned
        rows to the front, and stability preserves the recency order the first
        pass established within each group. Pinning is applied HERE, as an
        ordering over what the query already admitted — never as a way around
        it. A caller that hides an origin keeps hiding it when the row is
        pinned, because filtering a sorted list preserves relative order. That
        is what makes "a mobile surface shows only its own pinned sessions"
        true by construction, with no surface-specific branch anywhere in this
        method.
        """
        # Resolved HERE rather than left to the store: an omitted query means
        # "the default listing" (archived hidden), while the store's own
        # ``None`` means the unnarrowed read a self-filtering caller wants.
        query = query or SessionQuery()
        summaries: list[dict[str, object]] = []
        digests = self._session_store.list_session_digests(
            query, limit=limit, offset=offset
        )
        records = self._session_store.load_session_records(
            [digest.session_id for digest in digests]
        )
        for digest in digests:
            session_id = digest.session_id
            events = digest.events
            record = records[session_id]
            summary = self.summarize_session(session_id, events=events, record=record)
            # Evaluated over the DIGEST, and that is exact rather than merely
            # close: a session with a visible event has a ``user`` event, which
            # the digest always carries, so this can only read False for a
            # session the next check drops anyway (no user turn and not running
            # ⇒ ``created_at`` is None). The two filters are kept separate
            # regardless — the digest's contents are a store concern, and a
            # listing rule that silently depended on them would be one
            # projection change away from dropping rows.
            has_visible_event = any(
                event.get("type") not in {"session", "context"} for event in events
            )
            if not has_visible_event and not summary.get("running"):
                continue
            if summary.get("created_at") is None and not summary.get("running"):
                continue
            summaries.append(summary)
        summaries.sort(key=lambda s: str(s.get("created_at") or ""), reverse=True)
        summaries.sort(key=lambda s: str(s.get("pinned_at") or ""), reverse=True)
        return summaries

    def set_session_pinned(self, session_id: str, pinned: bool) -> str | None:
        """Pin or unpin a session, returning the resulting ``pinned_at`` stamp.

        Lives on the runtime for the same reason archive/rename/fork do: it is
        the ONE place a session's record is mutated, so every surface reaches the
        store through it rather than around it.
        """
        self._session_store.set_pinned(session_id, pinned)
        return self._session_store.get_pinned_at(session_id)

    def load_events(self, session_id: str, after: str | None = None) -> list[EventRecord]:
        """Load a session's events, narrowed to those newer than *after*.

        ``O(matched events)`` on a driver that can range-read, ``O(one session's
        events)`` on one that cannot — the store decides, which is the point.
        Materialising the whole transcript and filtering it in Python would
        make ``after`` narrow the RESPONSE while the read stayed the record's
        entire history: a cursor in name only. Measure the TIME, not the
        payload — a shrinking response hides constant work.

        An unparseable *after* returns everything, unchanged: an in-process
        caller has no channel to be refused on, and most of them pass no cursor
        at all. That widening is a DEFENSIVE default and nothing may rely on it
        — the refusal belongs to the surface that accepted the value from a
        client, and the ``/events`` route 400s before reaching here.
        """
        cursor = EventCursor.parse(after) if after else None
        return self._session_store.load_events_after(session_id, cursor)

    def start_async(
        self,
        *,
        session_id: str,
        user_query: str,
        model_name: str | None = None,
        fallback_models: tuple[str, ...] | None = None,
        max_iters: int = 3,
        initial_plan: Plan | None = None,
        tool_registry=None,
        permission_policy=None,
        approval_callback=None,
        hook_manager=None,
        mode: str | None = None,
        allowed_tools: list[str] | None = None,
        denied_tools: list[str] | None = None,
        strict_tool_scope: bool = False,
        capability_mode: str = "all",
        skill_instructions: str | None = None,
        cwd: str | None = None,
        session_step_budget: int = 0,
        user_id: str | None = None,
        source_platform: str | None = None,
        invocation_id: str | None = None,
        extra_session_tools: list[SessionTool] | None = None,
        enable_skills: bool = True,
        project_autoselect: bool = False,
        attachments: list[dict] | None = None,
        session_mcp_servers: dict[str, dict] | None = None,
    ) -> str:
        """Start an asynchronous orchestration run for the session.

        ``capability_mode`` is the ROOT delegation-privilege ceiling (default
        ``"all"`` — no filtering). A caller that resolved a
        principal's role into a narrower tier (``read_only`` for a viewer)
        passes it here so the session's own tools — not only spawned children —
        are capped; it travels unchanged into ``AgentContext.root`` alongside
        ``workspace_mode``.

        ``fallback_models`` (when provided) opts this run into cross-model
        fallback; ``None`` defers to the resolved config policy.

        Returns a storeless per-run handle ``run_id`` of the form
        ``"<session_id>:r<seq>"`` where *seq* is the 1-based count of runs
        started on this session so far (so one session can host many runs).
        The run_id is resolvable back to the session id by splitting on the
        first ``:`` — no new store or index is required. When the run
        registry refuses the start (a run is already active for this
        session), an empty string is returned so existing ``if not started``
        callers still detect the refusal.

        Accepting a turn PERSISTS it: the ``user`` event is written here, before
        the run is handed off, so a caller that got a run_id can rely on the
        transcript already holding the turn. See the block below for why the
        executor's own append is suppressed rather than deduplicated.

        When the run ends without meeting an explicit goal, :class:`GoalRetryGate`
        re-drives the session ONCE from the post-release seam below. That retry
        replays THIS call's arguments unchanged, so it inherits the same tool
        registry, hooks, capability ceiling and workspace the operator's run had.
        """
        # Snapshot the accepted arguments BEFORE any local is bound — at this
        # point ``locals()`` is exactly this call's parameters. A one-shot goal
        # retry has to replay all of them, and re-listing thirty names here
        # would be a second home for the same contract, drifting silently the
        # day a parameter is added.
        run_arguments: dict[str, Any] = dict(locals())
        run_arguments.pop("self", None)
        run_id = self._mint_run_id(session_id)
        # A caller-supplied invocation_id wins; otherwise seed the Langfuse
        # trace id from the run id we just minted, so each run gets its own
        # trace instead of every run of a session collapsing onto one. The
        # snapshot above already captured the unresolved (possibly None)
        # value, so a goal-retry replay mints its OWN distinct run/trace id
        # rather than inheriting this one.
        if invocation_id is None:
            invocation_id = run_id
        msg_queue: queue.Queue[str] = queue.Queue()
        interrupt_event = threading.Event()

        # Emit an instant ``run_accepted`` lifecycle marker BEFORE the background
        # run begins its heavy synchronous setup (``Orchestrator.__init__``
        # → tool-registry build + project-instruction discovery, which precede the
        # first run-phase ``append_event``). Without it the session page sits blank
        # for that whole window; with it the FE renders the session shell + a
        # "starting…" state immediately. It rides the same
        # ``append_event`` → ``SessionEventBus`` → SSE choke-point as every other
        # event, so no new transport is needed — additive only. Skipped when a run
        # is already active (the registry would refuse the start) so we never emit
        # a marker for a run that does not begin. Shared by both backends below.
        #
        # THE ACCEPTANCE SEAM PERSISTS WHAT IT ACCEPTED. The ``user`` event — the
        # record of the text the operator typed — is written HERE, adjacent to the
        # marker, not by the EXECUTOR: the orchestration body runs only after that
        # same heavy setup, so a turn appended there is invisible for the whole
        # cold-start window and every client renders an empty session. Writing it
        # at acceptance is what makes a 202 mean the turn is durable.
        # ``user_turn_persisted`` below then tells the body to skip its own
        # append, so the turn is written exactly once.
        #
        # The flag is the RESULT of the guard, never a hardcoded ``True``: this
        # pre-check is not atomic with the registry's own locked accept, so an
        # active run that finishes in between lets the start succeed after we
        # skipped the write. Deriving the flag means the body writes the turn in
        # exactly that case instead of the turn being lost outright. The opposite
        # skew — we wrote, then the registry refused — costs a turn record for a
        # run that never began, which is the same benign shape the marker has
        # always had, and strictly better than dropping a real turn.
        user_turn_persisted = not self._run_registry.is_running(session_id)
        if user_turn_persisted:
            self.append_event(
                session_id,
                {
                    "type": "run_accepted",
                    "payload": {
                        "session_id": session_id,
                        "run_id": run_id,
                        "ts": _utc_now(),
                    },
                },
            )
            self._session_store.append_user_turn(session_id, user_query, attachments)

        def _relaunch(query: str) -> str:
            # Attachments are deliberately dropped: they were persisted onto the
            # turn this session already holds, and replaying them would attach
            # the same files to a second turn. Everything else is replayed as-is.
            return self.start_async(
                **{**run_arguments, "user_query": query, "attachments": None}
            )

        def _on_release() -> None:
            self._goal_retry_gate.maybe_retry(session_id, _relaunch)

        # Backend selection: ONLY a single-threaded async host (Pyodide's
        # WebLoop, ``sys.platform == "emscripten"``) with an already-running
        # event loop drives the orchestration as an asyncio task via
        # ``orchestrate_session_async`` — no daemon thread, no nested
        # ``asyncio.run``. Gating on the platform (not merely "is a loop
        # running") is deliberate: CPython always takes the daemon-thread path
        # below even when called from within a running loop (e.g. a future
        # async CPython caller), so blocking orchestration never runs on that
        # caller's own loop. Cancellation is surfaced through the same
        # ``RunRegistry`` cancel_event either way.
        running_loop: asyncio.AbstractEventLoop | None = None
        if sys.platform == "emscripten":
            try:
                running_loop = asyncio.get_running_loop()
            except RuntimeError:
                running_loop = None

        if running_loop is not None:
            cancel_event = threading.Event()
            if not self._run_registry.register_loop_run(
                session_id,
                cancel_event=cancel_event,
                message_queue=msg_queue,
                interrupt_step=interrupt_event,
            ):
                return ""
            coro = orchestrate_session_async(
                user_query=user_query,
                model_name=model_name,
                fallback_models=fallback_models,
                max_iters=max_iters,
                initial_plan=initial_plan,
                session_id=session_id,
                session_store=self._session_store,
                tool_registry=tool_registry,
                permission_policy=permission_policy,
                approval_callback=approval_callback,
                hook_manager=hook_manager,
                mode=mode,
                should_cancel=cancel_event.is_set,
                allowed_tools=allowed_tools,
                denied_tools=denied_tools,
                strict_tool_scope=strict_tool_scope,
                capability_mode=capability_mode,
                skill_instructions=skill_instructions,
                message_queue=msg_queue,
                interrupt_step=interrupt_event,
                cwd=cwd,
                session_step_budget=session_step_budget,
                user_id=user_id,
                source_platform=source_platform,
                invocation_id=invocation_id,
                extra_session_tools=extra_session_tools,
                enable_skills=enable_skills,
                project_autoselect=project_autoselect,
                attachments=attachments,
                user_turn_persisted=user_turn_persisted,
                session_mcp_servers=session_mcp_servers,
            )
            task = running_loop.create_task(coro)
            # Hold a strong reference on the registry-owned handle so asyncio
            # can't GC this fire-and-forget task mid-run (mirrors
            # AgentHandle.asyncio_task / SpawnAgentTool._lifecycle_tasks).
            self._run_registry.attach_loop_task(session_id, task)
            task.add_done_callback(
                lambda t: self._on_loop_run_done(session_id, t, _on_release)
            )
            return run_id

        def _run(cancel_event: threading.Event) -> None:
            self.run_sync(
                user_query=user_query,
                session_id=session_id,
                model_name=model_name,
                fallback_models=fallback_models,
                max_iters=max_iters,
                initial_plan=initial_plan,
                tool_registry=tool_registry,
                permission_policy=permission_policy,
                approval_callback=approval_callback,
                hook_manager=hook_manager,
                mode=mode,
                should_cancel=cancel_event.is_set,
                allowed_tools=allowed_tools,
                denied_tools=denied_tools,
                strict_tool_scope=strict_tool_scope,
                capability_mode=capability_mode,
                skill_instructions=skill_instructions,
                message_queue=msg_queue,
                interrupt_step=interrupt_event,
                cwd=cwd,
                session_step_budget=session_step_budget,
                user_id=user_id,
                source_platform=source_platform,
                invocation_id=invocation_id,
                extra_session_tools=extra_session_tools,
                enable_skills=enable_skills,
                project_autoselect=project_autoselect,
                attachments=attachments,
                user_turn_persisted=user_turn_persisted,
                session_mcp_servers=session_mcp_servers,
            )

        started = self._run_registry.start(
            session_id,
            target=_run,
            message_queue=msg_queue,
            interrupt_step=interrupt_event,
            on_release=_on_release,
        )
        return run_id if started else ""

    def _on_loop_run_done(
        self,
        session_id: str,
        task: asyncio.Task,
        on_release: Callable[[], None] | None = None,
    ) -> None:
        """Done-callback for a loop-backed run: finalize + surface exceptions.

        *on_release* is the loop-backed twin of the thread path's post-release
        callback and fires for the same reason and in the same order — after
        ``finalize_loop_run`` has dropped the handle, so the session's slot is
        free for a follow-up run.

        ``Orchestrator``/``orchestrate_session_async`` already catch and log
        everything that happens INSIDE the orchestration body, so an
        exception surfacing here means something failed BEFORE that
        try/except (task construction, an unawaited-setup bug). Without
        reading it here, asyncio only emits a bare "Task exception was never
        retrieved" warning with no session context — read + log it via the
        module logger instead.
        """
        self._run_registry.finalize_loop_run(session_id)
        if task.cancelled():
            return
        exc = task.exception()
        if exc is not None:
            logging.error(
                "Loop-backed orchestration run failed for session {}: {}: {}",
                session_id,
                type(exc).__name__,
                exc,
            )

    def _mint_run_id(self, session_id: str) -> str:
        """Mint ``"<session_id>:r<seq>"`` for a run about to start.

        *seq* is 1-based and counts this run: it is one more than the number
        of ``user`` events already in the transcript (each prior run appended
        exactly one, and runs are serialized — the registry refuses a
        concurrent start). A fresh session has zero prior user events → ``r1``.
        Storeless: derived from the transcript, no separate counter to persist.
        """
        try:
            events = self._session_store.load_transcript(session_id)
            prior_runs = sum(1 for e in events if e.get("type") == "user")
        except Exception:  # pragma: no cover - defensive; never block a start
            prior_runs = 0
        return f"{session_id}:r{prior_runs + 1}"

    def run_sync(
        self,
        *,
        user_query: str,
        session_id: str,
        model_name: str | None = None,
        fallback_models: tuple[str, ...] | None = None,
        max_iters: int = 3,
        initial_plan: Plan | None = None,
        tool_registry=None,
        permission_policy=None,
        approval_callback=None,
        hook_manager=None,
        mode: str | None = None,
        should_cancel: Callable[[], bool] | None = None,
        allowed_tools: list[str] | None = None,
        denied_tools: list[str] | None = None,
        strict_tool_scope: bool = False,
        capability_mode: str = "all",
        skill_instructions: str | None = None,
        message_queue: queue.Queue[str] | None = None,
        interrupt_step: threading.Event | None = None,
        cwd: str | None = None,
        session_step_budget: int = 0,
        user_id: str | None = None,
        source_platform: str | None = None,
        invocation_id: str | None = None,
        extra_session_tools: list[SessionTool] | None = None,
        enable_skills: bool = True,
        project_autoselect: bool = False,
        attachments: list[dict] | None = None,
        user_turn_persisted: bool = False,
        session_mcp_servers: dict[str, dict] | None = None,
    ) -> TaskQueue:
        """Run an orchestration request synchronously.

        ``enable_skills=False`` opts a headless product drive (search/wiki) out
        of auto-skill injection so the agent never burns its first step
        activating a host ``~/.claude`` skill (default ``True`` — CLI/channel
        behavior unchanged). ``capability_mode`` is the root delegation-privilege
        ceiling (default ``"all"`` — no filtering); see :meth:`start_async`.

        ``project_autoselect=True`` binds ``list_projects`` / ``switch_project``
        on the ROOT agent so the run chooses its own workspace; default
        ``False`` binds neither.

        ``user_turn_persisted=True`` says an upstream seam already wrote this
        turn's ``user`` event, so the orchestration body must not write a second
        one. Default ``False`` — a direct caller (the CLI turn engine, the
        structured runners) still owns the append.
        """
        return orchestrate_session(
            user_query=user_query,
            model_name=model_name,
            fallback_models=fallback_models,
            max_iters=max_iters,
            initial_plan=initial_plan,
            session_id=session_id,
            session_store=self._session_store,
            tool_registry=tool_registry,
            permission_policy=permission_policy,
            approval_callback=approval_callback,
            hook_manager=hook_manager,
            mode=mode,
            should_cancel=should_cancel,
            allowed_tools=allowed_tools,
            denied_tools=denied_tools,
            strict_tool_scope=strict_tool_scope,
            capability_mode=capability_mode,
            skill_instructions=skill_instructions,
            message_queue=message_queue,
            interrupt_step=interrupt_step,
            cwd=cwd,
            session_step_budget=session_step_budget,
            user_id=user_id,
            source_platform=source_platform,
            invocation_id=invocation_id,
            extra_session_tools=extra_session_tools,
            enable_skills=enable_skills,
            project_autoselect=project_autoselect,
            attachments=attachments,
            user_turn_persisted=user_turn_persisted,
            session_mcp_servers=session_mcp_servers,
        )

    async def arun(
        self,
        *,
        user_query: str,
        session_id: str,
        model_name: str | None = None,
        fallback_models: tuple[str, ...] | None = None,
        max_iters: int = 3,
        initial_plan: Plan | None = None,
        tool_registry=None,
        permission_policy=None,
        approval_callback=None,
        hook_manager=None,
        mode: str | None = None,
        should_cancel: Callable[[], bool] | None = None,
        allowed_tools: list[str] | None = None,
        denied_tools: list[str] | None = None,
        strict_tool_scope: bool = False,
        capability_mode: str = "all",
        skill_instructions: str | None = None,
        message_queue: queue.Queue[str] | None = None,
        interrupt_step: threading.Event | None = None,
        cwd: str | None = None,
        session_step_budget: int = 0,
        user_id: str | None = None,
        source_platform: str | None = None,
        invocation_id: str | None = None,
        extra_session_tools: list[SessionTool] | None = None,
        enable_skills: bool = True,
        project_autoselect: bool = False,
        attachments: list[dict] | None = None,
        user_turn_persisted: bool = False,
        session_mcp_servers: dict[str, dict] | None = None,
    ) -> TaskQueue:
        """Run an orchestration request inside an existing event loop.

        Async counterpart of :meth:`run_sync`. Use this when the caller is
        itself a coroutine (e.g. an in-browser Flask handler running on
        Pyodide's WebLoop) so the orchestration can ``await`` without
        ``asyncio.run``. Same signature and semantics as :meth:`run_sync`.
        """
        return await orchestrate_session_async(
            user_query=user_query,
            model_name=model_name,
            fallback_models=fallback_models,
            max_iters=max_iters,
            initial_plan=initial_plan,
            session_id=session_id,
            session_store=self._session_store,
            tool_registry=tool_registry,
            permission_policy=permission_policy,
            approval_callback=approval_callback,
            hook_manager=hook_manager,
            mode=mode,
            should_cancel=should_cancel,
            allowed_tools=allowed_tools,
            denied_tools=denied_tools,
            strict_tool_scope=strict_tool_scope,
            capability_mode=capability_mode,
            skill_instructions=skill_instructions,
            message_queue=message_queue,
            interrupt_step=interrupt_step,
            cwd=cwd,
            session_step_budget=session_step_budget,
            user_id=user_id,
            source_platform=source_platform,
            invocation_id=invocation_id,
            extra_session_tools=extra_session_tools,
            enable_skills=enable_skills,
            project_autoselect=project_autoselect,
            attachments=attachments,
            user_turn_persisted=user_turn_persisted,
            session_mcp_servers=session_mcp_servers,
        )

    def cancel(self, session_id: str) -> bool:
        """Cancel an active run if present."""
        return self._run_registry.cancel(session_id)

    def is_running(self, session_id: str) -> bool:
        """Return True if session has an active run."""
        return self._run_registry.is_running(session_id)

    def active_run_handle(self, session_id: str) -> RunHandle | None:
        """The live :class:`RunHandle` for *session_id*, if a run is active.

        Read-only steering-signal access for in-run collaborators — the
        ask-user question dispatcher polls the handle's ``cancel_event`` /
        ``interrupt_step`` / ``message_queue`` while blocked so a steer or
        interrupt supersedes a pending question instead of deadlocking behind
        it. Callers must treat the handle as read-only.
        """
        return self._run_registry.get_handle(session_id)

    def is_terminated(self, session_id: str) -> bool:
        """Return True iff the session was permanently terminated.

        The ONE choke point every entry-point guard reads — a thin read-through
        to the store so there is zero duplicated status derivation. An unknown
        session reads ``False`` (no stamp).
        """
        return self._session_store.is_terminated(session_id)

    def register_on_terminate(self, callback: Callable[[str], int | None]) -> None:
        """Register a callback fired once when a session is terminated.

        The Wave-2 cascade-cancel seam: the callback receives the terminated
        ``session_id`` and MAY return an int count of downstream artifacts it
        cancelled (e.g. scheduled triggers), which :meth:`terminate_session`
        sums into ``cancelled_triggers``. Best-effort — a raising callback is
        logged and skipped, never blocking termination.
        """
        self._on_terminate.append(callback)

    def terminate_session(self, session_id: str) -> dict[str, object]:
        """Permanently terminate a session — idempotent and irreversible.

        First call: stamps ``terminated_at``, cancels any live run (cooperative,
        via the run registry's cancel event), fires the ``on_terminate``
        callbacks (summing any returned cancelled-artifact counts), and appends
        a ``session_terminated`` transcript event — which rides the standard
        ``append_event`` → ``SessionEventBus`` → SSE choke-point, so a live
        stream observes the termination with no new transport. A repeat call is
        a no-op that returns the SAME shape with the ORIGINAL ``terminated_at``
        and ``cancelled_triggers: 0`` (the side effects never re-fire).

        Side effects fire exactly once, arbitrated by the store's set-once
        write: ``SessionStoreBase.terminate_session`` returns ``True`` only to
        the call that newly stamped the timestamp, so two concurrent FIRST
        calls (Flask is threaded even at ``--workers 1``) can't both pass an
        unlocked read and duplicate the event + callback fan-out — exactly one
        of them wins the race and runs the block below.

        Callers guard unknown-session (404) upstream; this assumes the session
        exists. Returns ``{session_id, status:"terminated", terminated_at,
        cancelled_triggers}``.
        """
        newly_terminated = self._session_store.terminate_session(session_id)
        terminated_at = self._session_store.get_terminated_at(session_id)
        if not newly_terminated:
            return {
                "session_id": session_id,
                "status": "terminated",
                "terminated_at": terminated_at,
                "cancelled_triggers": 0,
            }
        # Cancel any in-flight run so the terminated session stops working; the
        # cancel event unwinds the loop cooperatively (best-effort, no-op if idle).
        self._run_registry.cancel(session_id)
        cancelled_triggers = 0
        for callback in self._on_terminate:
            try:
                result = callback(session_id)
            except Exception:
                logging.warning(
                    "on_terminate callback failed for session {}", session_id, exc_info=True
                )
                continue
            if isinstance(result, int):
                cancelled_triggers += result
        # append_terminal_event, NOT append_event: the store already cached
        # this session as terminated (line above), so a guarded append would
        # drop the tombstone event it is itself trying to write.
        self._session_store.append_terminal_event(
            session_id,
            {
                "type": "session_terminated",
                "payload": {
                    "session_id": session_id,
                    "terminated_at": terminated_at,
                    "cancelled_triggers": cancelled_triggers,
                },
            },
        )
        return {
            "session_id": session_id,
            "status": "terminated",
            "terminated_at": terminated_at,
            "cancelled_triggers": cancelled_triggers,
        }

    def start_command(
        self,
        session_id: str,
        target: Callable[[threading.Event], None],
    ) -> bool:
        """Start a non-orchestration background run for a slash command.

        Reuses the same RunRegistry as ``start_async`` so ``is_running()``
        and the events-polling pipeline treat command runs identically to
        query runs. The FE drives all in-flight UI off the same
        authoritative server state — no browser-side patching required.
        """
        return self._run_registry.start(session_id, target=target)

    def resolve_recovery_query(
        self,
        session_id: str,
        action: RecoveryAction,
        *,
        from_ts: str | None = None,
        replacement_text: str | None = None,
        trigger: RecoveryTrigger | None = None,
    ) -> str:
        """Resolve the user query text for a retry/continue recovery action.

        Appends a ``recovery`` audit event to the transcript and returns the
        query text the caller should pass to :meth:`start_async`. The
        orchestrator automatically picks up prior events via
        :class:`ContextBuilder`, so the caller does not need to trim the
        transcript.

        *replacement_text* substitutes the query this method would otherwise
        build: for ``retry`` the original user message (enabling "edit and
        regenerate"), for ``continue`` the generic resume prompt. An automatic
        re-drive supplies its own deterministic prompt that way rather than
        appending a second turn of its own.

        *trigger* names a NON-HUMAN originator on the ``recovery`` marker (see
        :data:`RecoveryTrigger`); omitted for every operator-driven recovery.
        It is what an automatic re-drive reads back to know it has already
        fired.

        Raises :class:`ValueError` when ``action`` is unrecognised, there is
        no prior user message to recover from, or (for ``retry``) the last
        user message is empty.
        Raises :class:`RuntimeError` if a run is already active for the
        session — cancel it first.
        """
        if action not in ("retry", "continue"):
            raise ValueError(f"unknown recovery action: {action!r}; expected 'retry' or 'continue'")
        if self.is_running(session_id):
            raise RuntimeError(f"session {session_id} is running; cancel before recovering")
        events = self._session_store.load_transcript(session_id)

        # Find the target user event.  When *from_ts* is given (only
        # meaningful for "retry"), locate the user event at that exact
        # timestamp so the caller can retry from any turn — not just the
        # last one.  Otherwise fall back to the most recent user event.
        if from_ts and action == "retry":
            last_user = next(
                (e for e in events if e.get("type") == "user" and e.get("ts") == from_ts),
                None,
            )
            if last_user is None:
                raise ValueError(f"no user event at ts={from_ts!r}")
        else:
            last_user = next(
                (e for e in reversed(events) if e.get("type") == "user"),
                None,
            )
        if last_user is None:
            raise ValueError(
                "no prior user message to recover from — start with a fresh query instead"
            )
        user_payload = last_user.get("payload") or {}
        original_text = user_payload.get("text", "") if isinstance(user_payload, dict) else ""

        if action == "retry":
            # ----------------------------------------------------------
            # Retry = time-travel: delete the failed turn so the session
            # looks like it ended right before that user message was
            # sent. ``Orchestrator.run`` re-appends the user event +
            # runs a fresh attempt. Prior successful turns stay intact.
            # ----------------------------------------------------------
            if not replacement_text and not original_text:
                raise ValueError("last user message has empty text; cannot retry")
            last_user_ts = last_user.get("ts", "")
            if last_user_ts:
                # Delete the user event itself + everything after it
                # (tool_results, completion, recovery events, …). Use
                # ``ts >= last_user_ts`` semantics by truncating after
                # the timestamp just BEFORE the user event.
                #
                # Find the event immediately before the last user.
                prev_ts = ""
                for ev in events:
                    if ev is last_user:
                        break
                    prev_ts = ev.get("ts", "")
                if prev_ts:
                    self._session_store.truncate_after(session_id, prev_ts)
                else:
                    # The user event is the first event — nuke everything
                    # by truncating after an impossibly-early timestamp.
                    self._session_store.truncate_after(session_id, "0000-00-00T00:00:00+00:00")
            query_text = replacement_text or original_text
        else:
            # ----------------------------------------------------------
            # Continue = stitch: keep the failed run's traces, delete
            # only a STALE PRIOR continue attempt, then start a new
            # continuation turn.
            #
            # A prior continue leaves a ``recovery`` audit marker + the
            # synthetic turn it drove; re-continuing drops that stale
            # attempt so the transcript doesn't accrete duplicate recovery
            # prompts. The cut anchors on the LAST ``recovery`` marker —
            # NEVER the last ``completion``. A run killed mid-flight
            # (process restart) emits neither a completion nor a recovery
            # marker for its turn, so a completion anchor falls back to the
            # PREVIOUS completed turn and deletes the entire interrupted
            # turn: its user message and every tool call/result. No prior
            # recovery marker ⇒ nothing stale ⇒ preserve all traces (the
            # interrupted turn stays an open turn the continuation resumes
            # from).
            #
            # The marker alone is NOT licence to cut. It records that a
            # continue was once triggered — never that it was the last
            # thing to happen. An attempt that went on to produce work
            # leaves that work AFTER the marker, so cutting at the marker
            # deletes it outright. A marker is stale only when its attempt
            # produced NOTHING DURABLE: no ``assistant`` message, no
            # ``completion`` and no ``tool_result``. Only a bare synthetic
            # re-prompt that the model never answered is safe to drop.
            #
            # ``assistant`` counts as durable even with no ``completion``
            # behind it. ``Orchestrator.run`` appends the assistant event
            # and the completion as SEPARATE writes with a title-generation
            # call between them, so a run that dies in that window leaves a
            # real answer the user already read on screen — the same window
            # the startup sweep closes with a synthetic completion.
            # Anything the user could have read must never be deleted.
            #
            # Staleness is a PREDICATE, not an anchor: the cut still lands
            # on the recovery marker.
            #
            # Both scans key off APPEND ORDER rather than a ts comparison.
            # The transcript is an append-only log, so position is its
            # ground truth, whereas ISO strings sort chronologically only
            # while every producer spells the UTC offset identically —
            # ``Z`` sorts above ``+00:00``, so a single same-second mix is
            # enough to widen the cut past the marker.
            # ----------------------------------------------------------
            last_recovery_idx = -1
            for idx, ev in enumerate(events):
                if ev.get("type") == "recovery":
                    last_recovery_idx = idx
            # A marker at index 0 leaves no earlier event to anchor the
            # cut on, so it is likewise treated as nothing to drop.
            if last_recovery_idx > 0 and not any(
                ev.get("type") in {"assistant", "completion", "tool_result"}
                for ev in events[last_recovery_idx + 1 :]
            ):
                # Truncate just BEFORE the stale recovery marker so its
                # own synthetic continue-turn is removed but the real
                # work preceding it is kept.
                prev_ts = events[last_recovery_idx - 1].get("ts", "")
                if prev_ts:
                    self._session_store.truncate_after(session_id, prev_ts)
            query_text = replacement_text or get_prompt_registry().render(
                "loop.recovery_continue"
            )
            # Audit marker so the transcript records when a continue was
            # triggered. Not appended for retry (the failed turn is deleted
            # entirely — no trace left to annotate), which is also why an
            # automatic re-drive must use ``continue``: this marker is its
            # one-shot ledger, and retry would delete it.
            payload: dict[str, object] = {"action": action}
            if trigger is not None:
                payload["trigger"] = trigger
            self._session_store.append_event(
                session_id,
                {"type": "recovery", "payload": payload},
            )

        return query_text

    def reinject_recovery_context(self, session_id: str) -> None:
        """Re-emit capability-gating fields so a recovered run keeps them.

        The orchestrator reads the *most-recent* ``context`` event when it
        builds the system prompt and resolves capability-gated AgentDefs.
        After a recovery turn the latest context event may be a plain
        ``mode``/``model`` update that does NOT carry the original
        ``client_capabilities`` / ``structured_workspace`` — so a recovered
        wiki/QA/structured session would silently lose its capability and
        ``spawn_agent`` lookups for the gated AgentDefs would fail.

        This scans the transcript for the latest value of each gating key
        (preserving exactly what the session already had — no origin→capability
        map) and appends a single fresh ``context`` event carrying them, so the
        most-recent context event after recovery still gates correctly. A no-op
        when the session never declared any gating field.

        Shared by BOTH the API recover endpoint and the CLI recovery command —
        the single source of truth for recovery capability re-injection.
        """
        events = self._session_store.load_transcript(session_id)
        gating: dict[str, object] = {}
        for event in events:
            if event.get("type") != "context":
                continue
            payload = event.get("payload")
            if not isinstance(payload, dict):
                continue
            for key in _RECOVERY_GATING_KEYS:
                if key in payload:
                    gating[key] = payload[key]
        if gating:
            self.append_context_event(session_id, gating)

    def enqueue_message(self, session_id: str, text: str) -> bool:
        """Enqueue a steering message for the root agent of a running session.

        The message is also persisted as a ``"user"`` event so it appears in
        the session transcript (console timeline, CLI history, Langfuse).

        Returns False if no active run or no message queue.
        """
        handle = self._run_registry.get_handle(session_id)
        if handle and handle.message_queue is not None:
            handle.message_queue.put_nowait(text)
            self._session_store.append_event(
                session_id, {"type": "user_steer", "payload": {"text": text}}
            )
            return True
        return False

    def interrupt_step(self, session_id: str) -> bool:
        """Interrupt the current tool execution step.

        The loop continues after the interrupted step with error results.
        Returns False if no active run or no interrupt event.
        """
        handle = self._run_registry.get_handle(session_id)
        if handle and handle.interrupt_step is not None:
            handle.interrupt_step.set()
            self._session_store.append_event(
                session_id,
                {"type": "user_steer", "payload": {"text": "[Interrupted by user]"}},
            )
            return True
        return False

    def _has_pending_plan_proposal(self, session_id: str) -> tuple[bool, int, str]:
        """Check if the session has an unresolved plan_proposed event.

        Returns ``(has_pending, revision, plan_path)`` where *revision* is the
        latest unresolved ``plan_proposed`` revision number, or 0 if none.
        """
        events = self._session_store.load_transcript(session_id)
        proposed_revisions: set[int] = set()
        resolved_revisions: set[int] = set()
        plan_path = ""
        for event in events:
            etype = event.get("type")
            payload = event.get("payload") or {}
            if not isinstance(payload, dict):
                continue
            rev = payload.get("revision", 0)
            if etype == "plan_proposed":
                proposed_revisions.add(rev)
                plan_path = payload.get("plan_path", "")
            elif etype in ("plan_approved", "plan_rejected"):
                resolved_revisions.add(rev)
        pending = proposed_revisions - resolved_revisions
        if pending:
            return True, max(pending), plan_path
        return False, 0, ""

    def approve_plan(self, session_id: str) -> bool:
        """Approve a pending plan proposal episodically.

        Emits a ``plan_approved`` event to the transcript. Does NOT start
        a new run — the caller (API endpoint) is responsible for starting
        the act-mode run via ``start_async``.

        Returns False if no pending plan proposal exists or a run is
        already active.
        """
        if self.is_running(session_id):
            return False
        has_pending, revision, plan_path = self._has_pending_plan_proposal(session_id)
        if not has_pending:
            return False
        self._session_store.append_event(
            session_id,
            {
                "type": "plan_approved",
                "payload": {"plan_path": plan_path, "revision": revision},
            },
        )
        # Signal mode transition so all clients pick up the change.
        self._session_store.append_event(
            session_id,
            {"type": "context", "payload": {"mode": "act"}},
        )
        return True

    def reject_plan(self, session_id: str) -> bool:
        """Reject a pending plan proposal.

        Emits a ``plan_rejected`` event. The session stays dormant —
        the user can type refinement guidance as a new message, which
        starts a fresh plan-mode run.

        Returns False if no pending plan proposal exists or a run is
        already active.
        """
        if self.is_running(session_id):
            return False
        has_pending, revision, plan_path = self._has_pending_plan_proposal(session_id)
        if not has_pending:
            return False
        self._session_store.append_event(
            session_id,
            {
                "type": "plan_rejected",
                "payload": {"plan_path": plan_path, "revision": revision},
            },
        )
        return True

session_store: SessionStoreBase property

Expose the underlying session store.

__init__(*, session_store: SessionStoreBase | None = None, run_registry: RunRegistry | None = None, goal_retry_gate: GoalRetryGate | None = None) -> None

Initialize the runtime with session storage and optional run registry.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
def __init__(
    self,
    *,
    session_store: SessionStoreBase | None = None,
    run_registry: RunRegistry | None = None,
    goal_retry_gate: GoalRetryGate | None = None,
) -> None:
    """Initialize the runtime with session storage and optional run registry."""
    self._session_store = session_store or create_session_store()
    self._run_registry = run_registry or RunRegistry()
    # The one-shot re-invocation of a session that ended without meeting its
    # goal. Injected so a caller can substitute or disable it; the default
    # binds to this runtime, which owns every seam the gate needs.
    self._goal_retry_gate = goal_retry_gate or GoalRetryGate(runtime=self)
    # Callbacks fired once when a session is permanently terminated. The
    # seam a Wave-2 trigger service registers on to cascade-cancel a
    # session's scheduled triggers; empty by default so termination has no
    # extra side effects until something opts in.
    self._on_terminate: list[Callable[[str], int | None]] = []

active_run_handle(session_id: str) -> RunHandle | None

The live :class:RunHandle for session_id, if a run is active.

Read-only steering-signal access for in-run collaborators — the ask-user question dispatcher polls the handle's cancel_event / interrupt_step / message_queue while blocked so a steer or interrupt supersedes a pending question instead of deadlocking behind it. Callers must treat the handle as read-only.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
def active_run_handle(self, session_id: str) -> RunHandle | None:
    """The live :class:`RunHandle` for *session_id*, if a run is active.

    Read-only steering-signal access for in-run collaborators — the
    ask-user question dispatcher polls the handle's ``cancel_event`` /
    ``interrupt_step`` / ``message_queue`` while blocked so a steer or
    interrupt supersedes a pending question instead of deadlocking behind
    it. Callers must treat the handle as read-only.
    """
    return self._run_registry.get_handle(session_id)

append_context_event(session_id: str, context: dict[str, object]) -> None

Append a context event to the session transcript.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
615
616
617
618
619
def append_context_event(self, session_id: str, context: dict[str, object]) -> None:
    """Append a context event to the session transcript."""
    if not context:
        return
    self._session_store.append_event(session_id, {"type": "context", "payload": context})

append_event(session_id: str, event: dict[str, object]) -> None

Append a raw transcript event verbatim.

Unlike :meth:append_context_event (which wraps payloads as {"type": "context", ...}), this writes the event as-is, so a completion event reaches :meth:summarize_session — the single status authority — instead of being hidden inside a context payload.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
632
633
634
635
636
637
638
639
640
def append_event(self, session_id: str, event: dict[str, object]) -> None:
    """Append a raw transcript event verbatim.

    Unlike :meth:`append_context_event` (which wraps payloads as
    ``{"type": "context", ...}``), this writes the event as-is, so a
    ``completion`` event reaches :meth:`summarize_session` — the single
    status authority — instead of being hidden inside a context payload.
    """
    self._session_store.append_event(session_id, event)

approve_plan(session_id: str) -> bool

Approve a pending plan proposal episodically.

Emits a plan_approved event to the transcript. Does NOT start a new run — the caller (API endpoint) is responsible for starting the act-mode run via start_async.

Returns False if no pending plan proposal exists or a run is already active.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
def approve_plan(self, session_id: str) -> bool:
    """Approve a pending plan proposal episodically.

    Emits a ``plan_approved`` event to the transcript. Does NOT start
    a new run — the caller (API endpoint) is responsible for starting
    the act-mode run via ``start_async``.

    Returns False if no pending plan proposal exists or a run is
    already active.
    """
    if self.is_running(session_id):
        return False
    has_pending, revision, plan_path = self._has_pending_plan_proposal(session_id)
    if not has_pending:
        return False
    self._session_store.append_event(
        session_id,
        {
            "type": "plan_approved",
            "payload": {"plan_path": plan_path, "revision": revision},
        },
    )
    # Signal mode transition so all clients pick up the change.
    self._session_store.append_event(
        session_id,
        {"type": "context", "payload": {"mode": "act"}},
    )
    return True

arun(*, user_query: str, session_id: str, model_name: str | None = None, fallback_models: tuple[str, ...] | None = None, max_iters: int = 3, initial_plan: Plan | None = None, tool_registry=None, permission_policy=None, approval_callback=None, hook_manager=None, mode: str | None = None, should_cancel: Callable[[], bool] | None = None, allowed_tools: list[str] | None = None, denied_tools: list[str] | None = None, strict_tool_scope: bool = False, capability_mode: str = 'all', skill_instructions: str | None = None, message_queue: queue.Queue[str] | None = None, interrupt_step: threading.Event | None = None, cwd: str | None = None, session_step_budget: int = 0, user_id: str | None = None, source_platform: str | None = None, invocation_id: str | None = None, extra_session_tools: list[SessionTool] | None = None, enable_skills: bool = True, project_autoselect: bool = False, attachments: list[dict] | None = None, user_turn_persisted: bool = False, session_mcp_servers: dict[str, dict] | None = None) -> TaskQueue async

Run an orchestration request inside an existing event loop.

Async counterpart of :meth:run_sync. Use this when the caller is itself a coroutine (e.g. an in-browser Flask handler running on Pyodide's WebLoop) so the orchestration can await without asyncio.run. Same signature and semantics as :meth:run_sync.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
async def arun(
    self,
    *,
    user_query: str,
    session_id: str,
    model_name: str | None = None,
    fallback_models: tuple[str, ...] | None = None,
    max_iters: int = 3,
    initial_plan: Plan | None = None,
    tool_registry=None,
    permission_policy=None,
    approval_callback=None,
    hook_manager=None,
    mode: str | None = None,
    should_cancel: Callable[[], bool] | None = None,
    allowed_tools: list[str] | None = None,
    denied_tools: list[str] | None = None,
    strict_tool_scope: bool = False,
    capability_mode: str = "all",
    skill_instructions: str | None = None,
    message_queue: queue.Queue[str] | None = None,
    interrupt_step: threading.Event | None = None,
    cwd: str | None = None,
    session_step_budget: int = 0,
    user_id: str | None = None,
    source_platform: str | None = None,
    invocation_id: str | None = None,
    extra_session_tools: list[SessionTool] | None = None,
    enable_skills: bool = True,
    project_autoselect: bool = False,
    attachments: list[dict] | None = None,
    user_turn_persisted: bool = False,
    session_mcp_servers: dict[str, dict] | None = None,
) -> TaskQueue:
    """Run an orchestration request inside an existing event loop.

    Async counterpart of :meth:`run_sync`. Use this when the caller is
    itself a coroutine (e.g. an in-browser Flask handler running on
    Pyodide's WebLoop) so the orchestration can ``await`` without
    ``asyncio.run``. Same signature and semantics as :meth:`run_sync`.
    """
    return await orchestrate_session_async(
        user_query=user_query,
        model_name=model_name,
        fallback_models=fallback_models,
        max_iters=max_iters,
        initial_plan=initial_plan,
        session_id=session_id,
        session_store=self._session_store,
        tool_registry=tool_registry,
        permission_policy=permission_policy,
        approval_callback=approval_callback,
        hook_manager=hook_manager,
        mode=mode,
        should_cancel=should_cancel,
        allowed_tools=allowed_tools,
        denied_tools=denied_tools,
        strict_tool_scope=strict_tool_scope,
        capability_mode=capability_mode,
        skill_instructions=skill_instructions,
        message_queue=message_queue,
        interrupt_step=interrupt_step,
        cwd=cwd,
        session_step_budget=session_step_budget,
        user_id=user_id,
        source_platform=source_platform,
        invocation_id=invocation_id,
        extra_session_tools=extra_session_tools,
        enable_skills=enable_skills,
        project_autoselect=project_autoselect,
        attachments=attachments,
        user_turn_persisted=user_turn_persisted,
        session_mcp_servers=session_mcp_servers,
    )

cancel(session_id: str) -> bool

Cancel an active run if present.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1560
1561
1562
def cancel(self, session_id: str) -> bool:
    """Cancel an active run if present."""
    return self._run_registry.cancel(session_id)

enqueue_message(session_id: str, text: str) -> bool

Enqueue a steering message for the root agent of a running session.

The message is also persisted as a "user" event so it appears in the session transcript (console timeline, CLI history, Langfuse).

Returns False if no active run or no message queue.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
def enqueue_message(self, session_id: str, text: str) -> bool:
    """Enqueue a steering message for the root agent of a running session.

    The message is also persisted as a ``"user"`` event so it appears in
    the session transcript (console timeline, CLI history, Langfuse).

    Returns False if no active run or no message queue.
    """
    handle = self._run_registry.get_handle(session_id)
    if handle and handle.message_queue is not None:
        handle.message_queue.put_nowait(text)
        self._session_store.append_event(
            session_id, {"type": "user_steer", "payload": {"text": text}}
        )
        return True
    return False

ensure_session(session_id: str) -> None

Idempotently materialise a session record for a pre-minted id.

Thin delegate to the store's ensure_session. A caller that minted a session_id outside resolve_session (e.g. the realtime recorder, which pre-mints to open a Langfuse trace before any store write) calls this so the session is a real RECORD — visible to list_sessions and every read surface — not an orphan transcript.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
604
605
606
607
608
609
610
611
612
613
def ensure_session(self, session_id: str) -> None:
    """Idempotently materialise a session record for a pre-minted id.

    Thin delegate to the store's ``ensure_session``. A caller that minted a
    ``session_id`` outside ``resolve_session`` (e.g. the realtime recorder,
    which pre-mints to open a Langfuse trace before any store write) calls
    this so the session is a real RECORD — visible to ``list_sessions`` and
    every read surface — not an orphan transcript.
    """
    self._session_store.ensure_session(session_id)

interrupt_step(session_id: str) -> bool

Interrupt the current tool execution step.

The loop continues after the interrupted step with error results. Returns False if no active run or no interrupt event.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
def interrupt_step(self, session_id: str) -> bool:
    """Interrupt the current tool execution step.

    The loop continues after the interrupted step with error results.
    Returns False if no active run or no interrupt event.
    """
    handle = self._run_registry.get_handle(session_id)
    if handle and handle.interrupt_step is not None:
        handle.interrupt_step.set()
        self._session_store.append_event(
            session_id,
            {"type": "user_steer", "payload": {"text": "[Interrupted by user]"}},
        )
        return True
    return False

is_running(session_id: str) -> bool

Return True if session has an active run.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1564
1565
1566
def is_running(self, session_id: str) -> bool:
    """Return True if session has an active run."""
    return self._run_registry.is_running(session_id)

is_terminated(session_id: str) -> bool

Return True iff the session was permanently terminated.

The ONE choke point every entry-point guard reads — a thin read-through to the store so there is zero duplicated status derivation. An unknown session reads False (no stamp).

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1579
1580
1581
1582
1583
1584
1585
1586
def is_terminated(self, session_id: str) -> bool:
    """Return True iff the session was permanently terminated.

    The ONE choke point every entry-point guard reads — a thin read-through
    to the store so there is zero duplicated status derivation. An unknown
    session reads ``False`` (no stamp).
    """
    return self._session_store.is_terminated(session_id)

list_sessions(query: SessionQuery | None = None, *, limit: int | None = None, offset: int = 0) -> list[dict[str, object]]

List sessions with summary metadata, narrowed by query.

O(collection) in rows, and never O(all history) in reads. TWO batched store calls serve the whole page, and neither scales with how much any session has recorded:

  • :meth:~SessionStoreBase.list_session_digests returns each session's transcript reduced to the events a summary folds. A per-id load_transcript here would make a listing read every event ever stored, and degrade monotonically forever.
  • load_session_records returns the metadata that is stored ON the session rather than folded from events (title, archived, terminated, owner, tags, pinned_at, projects) in one call for the page instead of per-row reads.

query narrows the digest fetch itself, and that ORDER is what makes it compose with a page. The whole query goes to list_session_digests (still called EXACTLY ONCE per listing, which a test pins), so every predicate answerable from the session record — owner/archived/pinned/projects — is applied by the store before it decides which candidates to page, and a session the query rejects is never opened. Filtering afterwards instead would be worse than merely slower once limit exists: paging first and filtering the page would make pinned=True return only the pinned sessions inside the newest N candidates, which is usually none of them.

The fold itself is unchanged and still single-sourced. summarize_session sees a smaller event list, not a different derivation — re-deriving a row's fields with a store-side aggregation would be a second implementation, free to drift until a row disagrees with the session it names.

limit/offset bound how many CANDIDATES this call examines, not how many rows it is guaranteed to return. They page list_session_digests after query has narrowed the candidate set but before the two VISIBILITY rules below run, which is what lets a Mongo-backed store scope its expensive read to the page instead of the whole collection (see that method's docstring for the measured saving) — but a candidate dropped below (no visible event and not running, or no created_at) shrinks the page rather than being backfilled from the next one. Those two rules stay here because they are the only ones that need the transcript's EVENTS, so no store query can decide them; they are hygiene rather than user filters, which is why shrinkage is acceptable for them and would not have been for pinned. limit=None (the default) is an unpaginated call.

Ordering is pinned first, then newest first. Two stable sorts rather than one composite key: the second pass only has to move pinned rows to the front, and stability preserves the recency order the first pass established within each group. Pinning is applied HERE, as an ordering over what the query already admitted — never as a way around it. A caller that hides an origin keeps hiding it when the row is pinned, because filtering a sorted list preserves relative order. That is what makes "a mobile surface shows only its own pinned sessions" true by construction, with no surface-specific branch anywhere in this method.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
def list_sessions(
    self,
    query: SessionQuery | None = None,
    *,
    limit: int | None = None,
    offset: int = 0,
) -> list[dict[str, object]]:
    """List sessions with summary metadata, narrowed by *query*.

    ``O(collection)`` in rows, and never ``O(all history)`` in reads. TWO
    batched store calls serve the whole page, and neither scales with how
    much any session has recorded:

    * :meth:`~SessionStoreBase.list_session_digests` returns each session's
      transcript reduced to the events a summary folds. A per-id
      ``load_transcript`` here would make a listing read every event ever
      stored, and degrade monotonically forever.
    * ``load_session_records`` returns the metadata that is stored ON the
      session rather than folded from events (title, archived, terminated,
      owner, tags, ``pinned_at``, ``projects``) in one call for the page
      instead of per-row reads.

    **`query` narrows the digest fetch itself, and that ORDER is what makes
    it compose with a page.** The whole query goes to
    ``list_session_digests`` (still called EXACTLY ONCE per listing, which a
    test pins), so every predicate answerable from the session record —
    ``owner``/``archived``/``pinned``/``projects`` — is applied by the store
    before it decides which candidates to page, and a session the query
    rejects is never opened. Filtering afterwards instead would be worse
    than merely slower once ``limit`` exists: paging first and filtering the
    page would make ``pinned=True`` return only the pinned sessions inside
    the newest N candidates, which is usually none of them.

    **The fold itself is unchanged and still single-sourced.**
    ``summarize_session`` sees a smaller event list, not a different
    derivation — re-deriving a row's fields with a store-side aggregation
    would be a second implementation, free to drift until a row disagrees
    with the session it names.

    **``limit``/``offset`` bound how many CANDIDATES this call examines, not
    how many rows it is guaranteed to return.** They page
    ``list_session_digests`` after *query* has narrowed the candidate set but
    before the two VISIBILITY rules below run, which is what lets a
    Mongo-backed store scope its expensive read to the page instead of the
    whole collection (see that method's docstring for the measured saving) —
    but a candidate dropped below (no visible event and not running, or no
    ``created_at``) shrinks the page rather than being backfilled from the
    next one. Those two rules stay here because they are the only ones that
    need the transcript's EVENTS, so no store query can decide them; they
    are hygiene rather than user filters, which is why shrinkage is
    acceptable for them and would not have been for ``pinned``.
    ``limit=None`` (the default) is an unpaginated call.

    Ordering is **pinned first, then newest first.** Two stable sorts
    rather than one composite key: the second pass only has to move pinned
    rows to the front, and stability preserves the recency order the first
    pass established within each group. Pinning is applied HERE, as an
    ordering over what the query already admitted — never as a way around
    it. A caller that hides an origin keeps hiding it when the row is
    pinned, because filtering a sorted list preserves relative order. That
    is what makes "a mobile surface shows only its own pinned sessions"
    true by construction, with no surface-specific branch anywhere in this
    method.
    """
    # Resolved HERE rather than left to the store: an omitted query means
    # "the default listing" (archived hidden), while the store's own
    # ``None`` means the unnarrowed read a self-filtering caller wants.
    query = query or SessionQuery()
    summaries: list[dict[str, object]] = []
    digests = self._session_store.list_session_digests(
        query, limit=limit, offset=offset
    )
    records = self._session_store.load_session_records(
        [digest.session_id for digest in digests]
    )
    for digest in digests:
        session_id = digest.session_id
        events = digest.events
        record = records[session_id]
        summary = self.summarize_session(session_id, events=events, record=record)
        # Evaluated over the DIGEST, and that is exact rather than merely
        # close: a session with a visible event has a ``user`` event, which
        # the digest always carries, so this can only read False for a
        # session the next check drops anyway (no user turn and not running
        # ⇒ ``created_at`` is None). The two filters are kept separate
        # regardless — the digest's contents are a store concern, and a
        # listing rule that silently depended on them would be one
        # projection change away from dropping rows.
        has_visible_event = any(
            event.get("type") not in {"session", "context"} for event in events
        )
        if not has_visible_event and not summary.get("running"):
            continue
        if summary.get("created_at") is None and not summary.get("running"):
            continue
        summaries.append(summary)
    summaries.sort(key=lambda s: str(s.get("created_at") or ""), reverse=True)
    summaries.sort(key=lambda s: str(s.get("pinned_at") or ""), reverse=True)
    return summaries

load_events(session_id: str, after: str | None = None) -> list[EventRecord]

Load a session's events, narrowed to those newer than after.

O(matched events) on a driver that can range-read, O(one session's events) on one that cannot — the store decides, which is the point. Materialising the whole transcript and filtering it in Python would make after narrow the RESPONSE while the read stayed the record's entire history: a cursor in name only. Measure the TIME, not the payload — a shrinking response hides constant work.

An unparseable after returns everything, unchanged: an in-process caller has no channel to be refused on, and most of them pass no cursor at all. That widening is a DEFENSIVE default and nothing may rely on it — the refusal belongs to the surface that accepted the value from a client, and the /events route 400s before reaching here.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
def load_events(self, session_id: str, after: str | None = None) -> list[EventRecord]:
    """Load a session's events, narrowed to those newer than *after*.

    ``O(matched events)`` on a driver that can range-read, ``O(one session's
    events)`` on one that cannot — the store decides, which is the point.
    Materialising the whole transcript and filtering it in Python would
    make ``after`` narrow the RESPONSE while the read stayed the record's
    entire history: a cursor in name only. Measure the TIME, not the
    payload — a shrinking response hides constant work.

    An unparseable *after* returns everything, unchanged: an in-process
    caller has no channel to be refused on, and most of them pass no cursor
    at all. That widening is a DEFENSIVE default and nothing may rely on it
    — the refusal belongs to the surface that accepted the value from a
    client, and the ``/events`` route 400s before reaching here.
    """
    cursor = EventCursor.parse(after) if after else None
    return self._session_store.load_events_after(session_id, cursor)

register_on_terminate(callback: Callable[[str], int | None]) -> None

Register a callback fired once when a session is terminated.

The Wave-2 cascade-cancel seam: the callback receives the terminated session_id and MAY return an int count of downstream artifacts it cancelled (e.g. scheduled triggers), which :meth:terminate_session sums into cancelled_triggers. Best-effort — a raising callback is logged and skipped, never blocking termination.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
def register_on_terminate(self, callback: Callable[[str], int | None]) -> None:
    """Register a callback fired once when a session is terminated.

    The Wave-2 cascade-cancel seam: the callback receives the terminated
    ``session_id`` and MAY return an int count of downstream artifacts it
    cancelled (e.g. scheduled triggers), which :meth:`terminate_session`
    sums into ``cancelled_triggers``. Best-effort — a raising callback is
    logged and skipped, never blocking termination.
    """
    self._on_terminate.append(callback)

reinject_recovery_context(session_id: str) -> None

Re-emit capability-gating fields so a recovered run keeps them.

The orchestrator reads the most-recent context event when it builds the system prompt and resolves capability-gated AgentDefs. After a recovery turn the latest context event may be a plain mode/model update that does NOT carry the original client_capabilities / structured_workspace — so a recovered wiki/QA/structured session would silently lose its capability and spawn_agent lookups for the gated AgentDefs would fail.

This scans the transcript for the latest value of each gating key (preserving exactly what the session already had — no origin→capability map) and appends a single fresh context event carrying them, so the most-recent context event after recovery still gates correctly. A no-op when the session never declared any gating field.

Shared by BOTH the API recover endpoint and the CLI recovery command — the single source of truth for recovery capability re-injection.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
def reinject_recovery_context(self, session_id: str) -> None:
    """Re-emit capability-gating fields so a recovered run keeps them.

    The orchestrator reads the *most-recent* ``context`` event when it
    builds the system prompt and resolves capability-gated AgentDefs.
    After a recovery turn the latest context event may be a plain
    ``mode``/``model`` update that does NOT carry the original
    ``client_capabilities`` / ``structured_workspace`` — so a recovered
    wiki/QA/structured session would silently lose its capability and
    ``spawn_agent`` lookups for the gated AgentDefs would fail.

    This scans the transcript for the latest value of each gating key
    (preserving exactly what the session already had — no origin→capability
    map) and appends a single fresh ``context`` event carrying them, so the
    most-recent context event after recovery still gates correctly. A no-op
    when the session never declared any gating field.

    Shared by BOTH the API recover endpoint and the CLI recovery command —
    the single source of truth for recovery capability re-injection.
    """
    events = self._session_store.load_transcript(session_id)
    gating: dict[str, object] = {}
    for event in events:
        if event.get("type") != "context":
            continue
        payload = event.get("payload")
        if not isinstance(payload, dict):
            continue
        for key in _RECOVERY_GATING_KEYS:
            if key in payload:
                gating[key] = payload[key]
    if gating:
        self.append_context_event(session_id, gating)

reject_plan(session_id: str) -> bool

Reject a pending plan proposal.

Emits a plan_rejected event. The session stays dormant — the user can type refinement guidance as a new message, which starts a fresh plan-mode run.

Returns False if no pending plan proposal exists or a run is already active.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
def reject_plan(self, session_id: str) -> bool:
    """Reject a pending plan proposal.

    Emits a ``plan_rejected`` event. The session stays dormant —
    the user can type refinement guidance as a new message, which
    starts a fresh plan-mode run.

    Returns False if no pending plan proposal exists or a run is
    already active.
    """
    if self.is_running(session_id):
        return False
    has_pending, revision, plan_path = self._has_pending_plan_proposal(session_id)
    if not has_pending:
        return False
    self._session_store.append_event(
        session_id,
        {
            "type": "plan_rejected",
            "payload": {"plan_path": plan_path, "revision": revision},
        },
    )
    return True

resolve_recovery_query(session_id: str, action: RecoveryAction, *, from_ts: str | None = None, replacement_text: str | None = None, trigger: RecoveryTrigger | None = None) -> str

Resolve the user query text for a retry/continue recovery action.

Appends a recovery audit event to the transcript and returns the query text the caller should pass to :meth:start_async. The orchestrator automatically picks up prior events via :class:ContextBuilder, so the caller does not need to trim the transcript.

replacement_text substitutes the query this method would otherwise build: for retry the original user message (enabling "edit and regenerate"), for continue the generic resume prompt. An automatic re-drive supplies its own deterministic prompt that way rather than appending a second turn of its own.

trigger names a NON-HUMAN originator on the recovery marker (see :data:RecoveryTrigger); omitted for every operator-driven recovery. It is what an automatic re-drive reads back to know it has already fired.

Raises :class:ValueError when action is unrecognised, there is no prior user message to recover from, or (for retry) the last user message is empty. Raises :class:RuntimeError if a run is already active for the session — cancel it first.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
def resolve_recovery_query(
    self,
    session_id: str,
    action: RecoveryAction,
    *,
    from_ts: str | None = None,
    replacement_text: str | None = None,
    trigger: RecoveryTrigger | None = None,
) -> str:
    """Resolve the user query text for a retry/continue recovery action.

    Appends a ``recovery`` audit event to the transcript and returns the
    query text the caller should pass to :meth:`start_async`. The
    orchestrator automatically picks up prior events via
    :class:`ContextBuilder`, so the caller does not need to trim the
    transcript.

    *replacement_text* substitutes the query this method would otherwise
    build: for ``retry`` the original user message (enabling "edit and
    regenerate"), for ``continue`` the generic resume prompt. An automatic
    re-drive supplies its own deterministic prompt that way rather than
    appending a second turn of its own.

    *trigger* names a NON-HUMAN originator on the ``recovery`` marker (see
    :data:`RecoveryTrigger`); omitted for every operator-driven recovery.
    It is what an automatic re-drive reads back to know it has already
    fired.

    Raises :class:`ValueError` when ``action`` is unrecognised, there is
    no prior user message to recover from, or (for ``retry``) the last
    user message is empty.
    Raises :class:`RuntimeError` if a run is already active for the
    session — cancel it first.
    """
    if action not in ("retry", "continue"):
        raise ValueError(f"unknown recovery action: {action!r}; expected 'retry' or 'continue'")
    if self.is_running(session_id):
        raise RuntimeError(f"session {session_id} is running; cancel before recovering")
    events = self._session_store.load_transcript(session_id)

    # Find the target user event.  When *from_ts* is given (only
    # meaningful for "retry"), locate the user event at that exact
    # timestamp so the caller can retry from any turn — not just the
    # last one.  Otherwise fall back to the most recent user event.
    if from_ts and action == "retry":
        last_user = next(
            (e for e in events if e.get("type") == "user" and e.get("ts") == from_ts),
            None,
        )
        if last_user is None:
            raise ValueError(f"no user event at ts={from_ts!r}")
    else:
        last_user = next(
            (e for e in reversed(events) if e.get("type") == "user"),
            None,
        )
    if last_user is None:
        raise ValueError(
            "no prior user message to recover from — start with a fresh query instead"
        )
    user_payload = last_user.get("payload") or {}
    original_text = user_payload.get("text", "") if isinstance(user_payload, dict) else ""

    if action == "retry":
        # ----------------------------------------------------------
        # Retry = time-travel: delete the failed turn so the session
        # looks like it ended right before that user message was
        # sent. ``Orchestrator.run`` re-appends the user event +
        # runs a fresh attempt. Prior successful turns stay intact.
        # ----------------------------------------------------------
        if not replacement_text and not original_text:
            raise ValueError("last user message has empty text; cannot retry")
        last_user_ts = last_user.get("ts", "")
        if last_user_ts:
            # Delete the user event itself + everything after it
            # (tool_results, completion, recovery events, …). Use
            # ``ts >= last_user_ts`` semantics by truncating after
            # the timestamp just BEFORE the user event.
            #
            # Find the event immediately before the last user.
            prev_ts = ""
            for ev in events:
                if ev is last_user:
                    break
                prev_ts = ev.get("ts", "")
            if prev_ts:
                self._session_store.truncate_after(session_id, prev_ts)
            else:
                # The user event is the first event — nuke everything
                # by truncating after an impossibly-early timestamp.
                self._session_store.truncate_after(session_id, "0000-00-00T00:00:00+00:00")
        query_text = replacement_text or original_text
    else:
        # ----------------------------------------------------------
        # Continue = stitch: keep the failed run's traces, delete
        # only a STALE PRIOR continue attempt, then start a new
        # continuation turn.
        #
        # A prior continue leaves a ``recovery`` audit marker + the
        # synthetic turn it drove; re-continuing drops that stale
        # attempt so the transcript doesn't accrete duplicate recovery
        # prompts. The cut anchors on the LAST ``recovery`` marker —
        # NEVER the last ``completion``. A run killed mid-flight
        # (process restart) emits neither a completion nor a recovery
        # marker for its turn, so a completion anchor falls back to the
        # PREVIOUS completed turn and deletes the entire interrupted
        # turn: its user message and every tool call/result. No prior
        # recovery marker ⇒ nothing stale ⇒ preserve all traces (the
        # interrupted turn stays an open turn the continuation resumes
        # from).
        #
        # The marker alone is NOT licence to cut. It records that a
        # continue was once triggered — never that it was the last
        # thing to happen. An attempt that went on to produce work
        # leaves that work AFTER the marker, so cutting at the marker
        # deletes it outright. A marker is stale only when its attempt
        # produced NOTHING DURABLE: no ``assistant`` message, no
        # ``completion`` and no ``tool_result``. Only a bare synthetic
        # re-prompt that the model never answered is safe to drop.
        #
        # ``assistant`` counts as durable even with no ``completion``
        # behind it. ``Orchestrator.run`` appends the assistant event
        # and the completion as SEPARATE writes with a title-generation
        # call between them, so a run that dies in that window leaves a
        # real answer the user already read on screen — the same window
        # the startup sweep closes with a synthetic completion.
        # Anything the user could have read must never be deleted.
        #
        # Staleness is a PREDICATE, not an anchor: the cut still lands
        # on the recovery marker.
        #
        # Both scans key off APPEND ORDER rather than a ts comparison.
        # The transcript is an append-only log, so position is its
        # ground truth, whereas ISO strings sort chronologically only
        # while every producer spells the UTC offset identically —
        # ``Z`` sorts above ``+00:00``, so a single same-second mix is
        # enough to widen the cut past the marker.
        # ----------------------------------------------------------
        last_recovery_idx = -1
        for idx, ev in enumerate(events):
            if ev.get("type") == "recovery":
                last_recovery_idx = idx
        # A marker at index 0 leaves no earlier event to anchor the
        # cut on, so it is likewise treated as nothing to drop.
        if last_recovery_idx > 0 and not any(
            ev.get("type") in {"assistant", "completion", "tool_result"}
            for ev in events[last_recovery_idx + 1 :]
        ):
            # Truncate just BEFORE the stale recovery marker so its
            # own synthetic continue-turn is removed but the real
            # work preceding it is kept.
            prev_ts = events[last_recovery_idx - 1].get("ts", "")
            if prev_ts:
                self._session_store.truncate_after(session_id, prev_ts)
        query_text = replacement_text or get_prompt_registry().render(
            "loop.recovery_continue"
        )
        # Audit marker so the transcript records when a continue was
        # triggered. Not appended for retry (the failed turn is deleted
        # entirely — no trace left to annotate), which is also why an
        # automatic re-drive must use ``continue``: this marker is its
        # one-shot ledger, and retry would delete it.
        payload: dict[str, object] = {"action": action}
        if trigger is not None:
            payload["trigger"] = trigger
        self._session_store.append_event(
            session_id,
            {"type": "recovery", "payload": payload},
        )

    return query_text

resolve_session(*, session_id: str | None = None, session_tag: str | None = None, fork_from: str | None = None, fork_at_ts: str | None = None, owner: str | None = None) -> str

Resolve session identifiers, tags, and forks to a session id.

When fork_at_ts is provided alongside fork_from, only events up to (and including) that timestamp are copied into the new session.

owner stamps whichever NEW session this call mints — a fresh one or a fork. It is an opaque subject string; core never learns what a principal is (see SessionStoreBase.create_session). Resolving to an EXISTING session ignores it: ownership is established once, at creation, so re-engaging someone else's session can never quietly re-stamp it.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
def resolve_session(
    self,
    *,
    session_id: str | None = None,
    session_tag: str | None = None,
    fork_from: str | None = None,
    fork_at_ts: str | None = None,
    owner: str | None = None,
) -> str:
    """Resolve session identifiers, tags, and forks to a session id.

    When *fork_at_ts* is provided alongside *fork_from*, only events up to
    (and including) that timestamp are copied into the new session.

    *owner* stamps whichever NEW session this call mints — a fresh one or a
    fork. It is an opaque subject string; core never learns what a principal
    is (see ``SessionStoreBase.create_session``). Resolving to an EXISTING
    session ignores it: ownership is established once, at creation, so
    re-engaging someone else's session can never quietly re-stamp it.
    """
    if fork_from:
        source_session_id = self._session_store.resolve_tag(fork_from) or fork_from
        if self.is_terminated(source_session_id):
            raise SessionTerminatedError(
                f"session {source_session_id} is permanently terminated; "
                "a terminated session cannot be forked"
            )
        if fork_at_ts:
            session_id = self._session_store.fork_session_at(
                source_session_id, fork_at_ts, owner
            )
        else:
            session_id = self._session_store.fork_session(source_session_id, owner)
    if session_tag and not session_id:
        resolved = self._session_store.resolve_tag(session_tag)
        session_id = resolved if resolved else None
    if not session_id:
        session_id = self._session_store.create_session(owner)
    if session_tag:
        self._session_store.tag_session(session_id, session_tag)
    assert session_id is not None
    return session_id

run_sync(*, user_query: str, session_id: str, model_name: str | None = None, fallback_models: tuple[str, ...] | None = None, max_iters: int = 3, initial_plan: Plan | None = None, tool_registry=None, permission_policy=None, approval_callback=None, hook_manager=None, mode: str | None = None, should_cancel: Callable[[], bool] | None = None, allowed_tools: list[str] | None = None, denied_tools: list[str] | None = None, strict_tool_scope: bool = False, capability_mode: str = 'all', skill_instructions: str | None = None, message_queue: queue.Queue[str] | None = None, interrupt_step: threading.Event | None = None, cwd: str | None = None, session_step_budget: int = 0, user_id: str | None = None, source_platform: str | None = None, invocation_id: str | None = None, extra_session_tools: list[SessionTool] | None = None, enable_skills: bool = True, project_autoselect: bool = False, attachments: list[dict] | None = None, user_turn_persisted: bool = False, session_mcp_servers: dict[str, dict] | None = None) -> TaskQueue

Run an orchestration request synchronously.

enable_skills=False opts a headless product drive (search/wiki) out of auto-skill injection so the agent never burns its first step activating a host ~/.claude skill (default True — CLI/channel behavior unchanged). capability_mode is the root delegation-privilege ceiling (default "all" — no filtering); see :meth:start_async.

project_autoselect=True binds list_projects / switch_project on the ROOT agent so the run chooses its own workspace; default False binds neither.

user_turn_persisted=True says an upstream seam already wrote this turn's user event, so the orchestration body must not write a second one. Default False — a direct caller (the CLI turn engine, the structured runners) still owns the append.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
def run_sync(
    self,
    *,
    user_query: str,
    session_id: str,
    model_name: str | None = None,
    fallback_models: tuple[str, ...] | None = None,
    max_iters: int = 3,
    initial_plan: Plan | None = None,
    tool_registry=None,
    permission_policy=None,
    approval_callback=None,
    hook_manager=None,
    mode: str | None = None,
    should_cancel: Callable[[], bool] | None = None,
    allowed_tools: list[str] | None = None,
    denied_tools: list[str] | None = None,
    strict_tool_scope: bool = False,
    capability_mode: str = "all",
    skill_instructions: str | None = None,
    message_queue: queue.Queue[str] | None = None,
    interrupt_step: threading.Event | None = None,
    cwd: str | None = None,
    session_step_budget: int = 0,
    user_id: str | None = None,
    source_platform: str | None = None,
    invocation_id: str | None = None,
    extra_session_tools: list[SessionTool] | None = None,
    enable_skills: bool = True,
    project_autoselect: bool = False,
    attachments: list[dict] | None = None,
    user_turn_persisted: bool = False,
    session_mcp_servers: dict[str, dict] | None = None,
) -> TaskQueue:
    """Run an orchestration request synchronously.

    ``enable_skills=False`` opts a headless product drive (search/wiki) out
    of auto-skill injection so the agent never burns its first step
    activating a host ``~/.claude`` skill (default ``True`` — CLI/channel
    behavior unchanged). ``capability_mode`` is the root delegation-privilege
    ceiling (default ``"all"`` — no filtering); see :meth:`start_async`.

    ``project_autoselect=True`` binds ``list_projects`` / ``switch_project``
    on the ROOT agent so the run chooses its own workspace; default
    ``False`` binds neither.

    ``user_turn_persisted=True`` says an upstream seam already wrote this
    turn's ``user`` event, so the orchestration body must not write a second
    one. Default ``False`` — a direct caller (the CLI turn engine, the
    structured runners) still owns the append.
    """
    return orchestrate_session(
        user_query=user_query,
        model_name=model_name,
        fallback_models=fallback_models,
        max_iters=max_iters,
        initial_plan=initial_plan,
        session_id=session_id,
        session_store=self._session_store,
        tool_registry=tool_registry,
        permission_policy=permission_policy,
        approval_callback=approval_callback,
        hook_manager=hook_manager,
        mode=mode,
        should_cancel=should_cancel,
        allowed_tools=allowed_tools,
        denied_tools=denied_tools,
        strict_tool_scope=strict_tool_scope,
        capability_mode=capability_mode,
        skill_instructions=skill_instructions,
        message_queue=message_queue,
        interrupt_step=interrupt_step,
        cwd=cwd,
        session_step_budget=session_step_budget,
        user_id=user_id,
        source_platform=source_platform,
        invocation_id=invocation_id,
        extra_session_tools=extra_session_tools,
        enable_skills=enable_skills,
        project_autoselect=project_autoselect,
        attachments=attachments,
        user_turn_persisted=user_turn_persisted,
        session_mcp_servers=session_mcp_servers,
    )

set_session_pinned(session_id: str, pinned: bool) -> str | None

Pin or unpin a session, returning the resulting pinned_at stamp.

Lives on the runtime for the same reason archive/rename/fork do: it is the ONE place a session's record is mutated, so every surface reaches the store through it rather than around it.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1076
1077
1078
1079
1080
1081
1082
1083
1084
def set_session_pinned(self, session_id: str, pinned: bool) -> str | None:
    """Pin or unpin a session, returning the resulting ``pinned_at`` stamp.

    Lives on the runtime for the same reason archive/rename/fork do: it is
    the ONE place a session's record is mutated, so every surface reaches the
    store through it rather than around it.
    """
    self._session_store.set_pinned(session_id, pinned)
    return self._session_store.get_pinned_at(session_id)

start_async(*, session_id: str, user_query: str, model_name: str | None = None, fallback_models: tuple[str, ...] | None = None, max_iters: int = 3, initial_plan: Plan | None = None, tool_registry=None, permission_policy=None, approval_callback=None, hook_manager=None, mode: str | None = None, allowed_tools: list[str] | None = None, denied_tools: list[str] | None = None, strict_tool_scope: bool = False, capability_mode: str = 'all', skill_instructions: str | None = None, cwd: str | None = None, session_step_budget: int = 0, user_id: str | None = None, source_platform: str | None = None, invocation_id: str | None = None, extra_session_tools: list[SessionTool] | None = None, enable_skills: bool = True, project_autoselect: bool = False, attachments: list[dict] | None = None, session_mcp_servers: dict[str, dict] | None = None) -> str

Start an asynchronous orchestration run for the session.

capability_mode is the ROOT delegation-privilege ceiling (default "all" — no filtering). A caller that resolved a principal's role into a narrower tier (read_only for a viewer) passes it here so the session's own tools — not only spawned children — are capped; it travels unchanged into AgentContext.root alongside workspace_mode.

fallback_models (when provided) opts this run into cross-model fallback; None defers to the resolved config policy.

Returns a storeless per-run handle run_id of the form "<session_id>:r<seq>" where seq is the 1-based count of runs started on this session so far (so one session can host many runs). The run_id is resolvable back to the session id by splitting on the first : — no new store or index is required. When the run registry refuses the start (a run is already active for this session), an empty string is returned so existing if not started callers still detect the refusal.

Accepting a turn PERSISTS it: the user event is written here, before the run is handed off, so a caller that got a run_id can rely on the transcript already holding the turn. See the block below for why the executor's own append is suppressed rather than deduplicated.

When the run ends without meeting an explicit goal, :class:GoalRetryGate re-drives the session ONCE from the post-release seam below. That retry replays THIS call's arguments unchanged, so it inherits the same tool registry, hooks, capability ceiling and workspace the operator's run had.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
def start_async(
    self,
    *,
    session_id: str,
    user_query: str,
    model_name: str | None = None,
    fallback_models: tuple[str, ...] | None = None,
    max_iters: int = 3,
    initial_plan: Plan | None = None,
    tool_registry=None,
    permission_policy=None,
    approval_callback=None,
    hook_manager=None,
    mode: str | None = None,
    allowed_tools: list[str] | None = None,
    denied_tools: list[str] | None = None,
    strict_tool_scope: bool = False,
    capability_mode: str = "all",
    skill_instructions: str | None = None,
    cwd: str | None = None,
    session_step_budget: int = 0,
    user_id: str | None = None,
    source_platform: str | None = None,
    invocation_id: str | None = None,
    extra_session_tools: list[SessionTool] | None = None,
    enable_skills: bool = True,
    project_autoselect: bool = False,
    attachments: list[dict] | None = None,
    session_mcp_servers: dict[str, dict] | None = None,
) -> str:
    """Start an asynchronous orchestration run for the session.

    ``capability_mode`` is the ROOT delegation-privilege ceiling (default
    ``"all"`` — no filtering). A caller that resolved a
    principal's role into a narrower tier (``read_only`` for a viewer)
    passes it here so the session's own tools — not only spawned children —
    are capped; it travels unchanged into ``AgentContext.root`` alongside
    ``workspace_mode``.

    ``fallback_models`` (when provided) opts this run into cross-model
    fallback; ``None`` defers to the resolved config policy.

    Returns a storeless per-run handle ``run_id`` of the form
    ``"<session_id>:r<seq>"`` where *seq* is the 1-based count of runs
    started on this session so far (so one session can host many runs).
    The run_id is resolvable back to the session id by splitting on the
    first ``:`` — no new store or index is required. When the run
    registry refuses the start (a run is already active for this
    session), an empty string is returned so existing ``if not started``
    callers still detect the refusal.

    Accepting a turn PERSISTS it: the ``user`` event is written here, before
    the run is handed off, so a caller that got a run_id can rely on the
    transcript already holding the turn. See the block below for why the
    executor's own append is suppressed rather than deduplicated.

    When the run ends without meeting an explicit goal, :class:`GoalRetryGate`
    re-drives the session ONCE from the post-release seam below. That retry
    replays THIS call's arguments unchanged, so it inherits the same tool
    registry, hooks, capability ceiling and workspace the operator's run had.
    """
    # Snapshot the accepted arguments BEFORE any local is bound — at this
    # point ``locals()`` is exactly this call's parameters. A one-shot goal
    # retry has to replay all of them, and re-listing thirty names here
    # would be a second home for the same contract, drifting silently the
    # day a parameter is added.
    run_arguments: dict[str, Any] = dict(locals())
    run_arguments.pop("self", None)
    run_id = self._mint_run_id(session_id)
    # A caller-supplied invocation_id wins; otherwise seed the Langfuse
    # trace id from the run id we just minted, so each run gets its own
    # trace instead of every run of a session collapsing onto one. The
    # snapshot above already captured the unresolved (possibly None)
    # value, so a goal-retry replay mints its OWN distinct run/trace id
    # rather than inheriting this one.
    if invocation_id is None:
        invocation_id = run_id
    msg_queue: queue.Queue[str] = queue.Queue()
    interrupt_event = threading.Event()

    # Emit an instant ``run_accepted`` lifecycle marker BEFORE the background
    # run begins its heavy synchronous setup (``Orchestrator.__init__``
    # → tool-registry build + project-instruction discovery, which precede the
    # first run-phase ``append_event``). Without it the session page sits blank
    # for that whole window; with it the FE renders the session shell + a
    # "starting…" state immediately. It rides the same
    # ``append_event`` → ``SessionEventBus`` → SSE choke-point as every other
    # event, so no new transport is needed — additive only. Skipped when a run
    # is already active (the registry would refuse the start) so we never emit
    # a marker for a run that does not begin. Shared by both backends below.
    #
    # THE ACCEPTANCE SEAM PERSISTS WHAT IT ACCEPTED. The ``user`` event — the
    # record of the text the operator typed — is written HERE, adjacent to the
    # marker, not by the EXECUTOR: the orchestration body runs only after that
    # same heavy setup, so a turn appended there is invisible for the whole
    # cold-start window and every client renders an empty session. Writing it
    # at acceptance is what makes a 202 mean the turn is durable.
    # ``user_turn_persisted`` below then tells the body to skip its own
    # append, so the turn is written exactly once.
    #
    # The flag is the RESULT of the guard, never a hardcoded ``True``: this
    # pre-check is not atomic with the registry's own locked accept, so an
    # active run that finishes in between lets the start succeed after we
    # skipped the write. Deriving the flag means the body writes the turn in
    # exactly that case instead of the turn being lost outright. The opposite
    # skew — we wrote, then the registry refused — costs a turn record for a
    # run that never began, which is the same benign shape the marker has
    # always had, and strictly better than dropping a real turn.
    user_turn_persisted = not self._run_registry.is_running(session_id)
    if user_turn_persisted:
        self.append_event(
            session_id,
            {
                "type": "run_accepted",
                "payload": {
                    "session_id": session_id,
                    "run_id": run_id,
                    "ts": _utc_now(),
                },
            },
        )
        self._session_store.append_user_turn(session_id, user_query, attachments)

    def _relaunch(query: str) -> str:
        # Attachments are deliberately dropped: they were persisted onto the
        # turn this session already holds, and replaying them would attach
        # the same files to a second turn. Everything else is replayed as-is.
        return self.start_async(
            **{**run_arguments, "user_query": query, "attachments": None}
        )

    def _on_release() -> None:
        self._goal_retry_gate.maybe_retry(session_id, _relaunch)

    # Backend selection: ONLY a single-threaded async host (Pyodide's
    # WebLoop, ``sys.platform == "emscripten"``) with an already-running
    # event loop drives the orchestration as an asyncio task via
    # ``orchestrate_session_async`` — no daemon thread, no nested
    # ``asyncio.run``. Gating on the platform (not merely "is a loop
    # running") is deliberate: CPython always takes the daemon-thread path
    # below even when called from within a running loop (e.g. a future
    # async CPython caller), so blocking orchestration never runs on that
    # caller's own loop. Cancellation is surfaced through the same
    # ``RunRegistry`` cancel_event either way.
    running_loop: asyncio.AbstractEventLoop | None = None
    if sys.platform == "emscripten":
        try:
            running_loop = asyncio.get_running_loop()
        except RuntimeError:
            running_loop = None

    if running_loop is not None:
        cancel_event = threading.Event()
        if not self._run_registry.register_loop_run(
            session_id,
            cancel_event=cancel_event,
            message_queue=msg_queue,
            interrupt_step=interrupt_event,
        ):
            return ""
        coro = orchestrate_session_async(
            user_query=user_query,
            model_name=model_name,
            fallback_models=fallback_models,
            max_iters=max_iters,
            initial_plan=initial_plan,
            session_id=session_id,
            session_store=self._session_store,
            tool_registry=tool_registry,
            permission_policy=permission_policy,
            approval_callback=approval_callback,
            hook_manager=hook_manager,
            mode=mode,
            should_cancel=cancel_event.is_set,
            allowed_tools=allowed_tools,
            denied_tools=denied_tools,
            strict_tool_scope=strict_tool_scope,
            capability_mode=capability_mode,
            skill_instructions=skill_instructions,
            message_queue=msg_queue,
            interrupt_step=interrupt_event,
            cwd=cwd,
            session_step_budget=session_step_budget,
            user_id=user_id,
            source_platform=source_platform,
            invocation_id=invocation_id,
            extra_session_tools=extra_session_tools,
            enable_skills=enable_skills,
            project_autoselect=project_autoselect,
            attachments=attachments,
            user_turn_persisted=user_turn_persisted,
            session_mcp_servers=session_mcp_servers,
        )
        task = running_loop.create_task(coro)
        # Hold a strong reference on the registry-owned handle so asyncio
        # can't GC this fire-and-forget task mid-run (mirrors
        # AgentHandle.asyncio_task / SpawnAgentTool._lifecycle_tasks).
        self._run_registry.attach_loop_task(session_id, task)
        task.add_done_callback(
            lambda t: self._on_loop_run_done(session_id, t, _on_release)
        )
        return run_id

    def _run(cancel_event: threading.Event) -> None:
        self.run_sync(
            user_query=user_query,
            session_id=session_id,
            model_name=model_name,
            fallback_models=fallback_models,
            max_iters=max_iters,
            initial_plan=initial_plan,
            tool_registry=tool_registry,
            permission_policy=permission_policy,
            approval_callback=approval_callback,
            hook_manager=hook_manager,
            mode=mode,
            should_cancel=cancel_event.is_set,
            allowed_tools=allowed_tools,
            denied_tools=denied_tools,
            strict_tool_scope=strict_tool_scope,
            capability_mode=capability_mode,
            skill_instructions=skill_instructions,
            message_queue=msg_queue,
            interrupt_step=interrupt_event,
            cwd=cwd,
            session_step_budget=session_step_budget,
            user_id=user_id,
            source_platform=source_platform,
            invocation_id=invocation_id,
            extra_session_tools=extra_session_tools,
            enable_skills=enable_skills,
            project_autoselect=project_autoselect,
            attachments=attachments,
            user_turn_persisted=user_turn_persisted,
            session_mcp_servers=session_mcp_servers,
        )

    started = self._run_registry.start(
        session_id,
        target=_run,
        message_queue=msg_queue,
        interrupt_step=interrupt_event,
        on_release=_on_release,
    )
    return run_id if started else ""

start_command(session_id: str, target: Callable[[threading.Event], None]) -> bool

Start a non-orchestration background run for a slash command.

Reuses the same RunRegistry as start_async so is_running() and the events-polling pipeline treat command runs identically to query runs. The FE drives all in-flight UI off the same authoritative server state — no browser-side patching required.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
def start_command(
    self,
    session_id: str,
    target: Callable[[threading.Event], None],
) -> bool:
    """Start a non-orchestration background run for a slash command.

    Reuses the same RunRegistry as ``start_async`` so ``is_running()``
    and the events-polling pipeline treat command runs identically to
    query runs. The FE drives all in-flight UI off the same
    authoritative server state — no browser-side patching required.
    """
    return self._run_registry.start(session_id, target=target)

summarize_session(session_id: str, *, events: list[EventRecord] | None = None, record: SessionRecord | None = None) -> dict[str, object]

Return a summarized view of a session.

record supplies the session's stored metadata (title, archived, terminated, owner, tags) when the caller has already batch-loaded it for a whole page — see :meth:list_sessions. Omitting it reads the same five facts one at a time, which is what every single-session caller does and what this method has always done, so the derived summary is identical either way; only the number of store reads differs.

With no events, the fold's input is the store's DIGEST of the session rather than its whole transcript — the same projection a listing row gets, and identical in result. It matters most on the poll path, which re-derives status once a second per open client: on the largest live session that read was 0.211 s of a 0.470 s poll, for a status that had not changed.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
def summarize_session(
    self,
    session_id: str,
    *,
    events: list[EventRecord] | None = None,
    record: SessionRecord | None = None,
) -> dict[str, object]:
    """Return a summarized view of a session.

    *record* supplies the session's stored metadata (title, archived,
    terminated, owner, tags) when the caller has already batch-loaded it for
    a whole page — see :meth:`list_sessions`. Omitting it reads the same
    five facts one at a time, which is what every single-session caller
    does and what this method has always done, so the derived summary is
    identical either way; only the number of store reads differs.

    With no *events*, the fold's input is the store's DIGEST of the session
    rather than its whole transcript — the same projection a listing row
    gets, and identical in result. It
    matters most on the poll path, which re-derives status once a second per
    open client: on the largest live session that read was 0.211 s of a
    0.470 s poll, for a status that had not changed.
    """
    if events is None:
        events = self._session_store.session_digest(session_id).events
    if record is None:
        record = self._session_store.load_session_records([session_id])[session_id]
    created_at = events[0]["ts"] if events else None
    stored_title = record.title
    title = stored_title
    status: SessionStatus = "idle"
    done_reason = None
    blocked_code: str | None = None
    has_user_event = False
    # Evidence for the model-attributable-failure projection, gathered in
    # this SAME pass — ``list_sessions`` calls this per session, so a second
    # walk of every transcript is a cost with no new information.
    failure_reason: str | None = None
    models_tried: list[str] = []
    outcome_assertion: OutcomeAssertion | None = None
    # Two more facts gathered in that same single pass, for the same reason:
    # how many lines the session changed, and which gated capabilities it
    # actually exercised (as opposed to which ones its client advertised).
    diff_stat = DiffStat()
    evidence = CapabilityEvidence()
    for event in events:
        evidence.observe(event)
        if event.get("type") == "tool_result":
            result_payload = event.get("payload", {})
            if isinstance(result_payload, dict):
                diff_stat = diff_stat + DiffStat.from_tool_result(result_payload)
        if event.get("type") == "user":
            has_user_event = True
            if title is None:
                payload = event.get("payload", {})
                if isinstance(payload, dict):
                    raw = payload.get("text")
                    if isinstance(raw, str):
                        title = raw[:120]
        if event.get("type") in RESILIENCE_EVENT_TYPES:
            failure_reason = self._read_resilience_event(
                event, models_tried, failure_reason
            )
        if event.get("type") == "outcome_assertion":
            assertion = self._read_outcome_assertion(event)
            if assertion is not None:
                outcome_assertion = assertion
        if event.get("type") == "completion":
            payload = event.get("payload", {})
            if isinstance(payload, dict):
                # A NEW terminal supersedes any assertion made about the
                # PREVIOUS one. A session hosts many turns, and turn 5's
                # status must not inherit turn 1's unmet purpose — the
                # assertion is appended right after the terminal it
                # describes, so append order alone scopes it correctly.
                outcome_assertion = None
                # EXACTLY ONE terminal event governs the derived state, and
                # every field below is re-read from THIS payload. A session
                # can carry two contradictory terminals — a boot sweep
                # stamps a synthetic "interrupted" completion for a run it
                # judged dead, and the run, still alive, appends its real
                # one minutes later — so deriving field-by-field across
                # events would blend the sweep's reason into the live run's
                # status. Last terminal wins, whole.
                done_reason = payload.get("done_reason")
                status, blocked_code = self._completion_status(payload)
                self._read_models_tried(payload, models_tried)
    # An outcome assertion is the ONLY signal that can contradict a terminal
    # the loop itself considers clean, so it is applied against the derived
    # status rather than folded into the reason table. It promotes ONLY a
    # claim of success: `failed`/`canceled` are already honest, and `blocked`
    # is both more specific and more actionable, so none of them is
    # overwritten by the coarser "purpose not met".
    if outcome_assertion is not None and status == "completed":
        status = "unmet_goal"
    running = self.is_running(session_id)
    if running:
        status = "running"
    # Permanent termination is the terminal-most state — it wins over
    # EVERYTHING, including a still-unwinding live run (cancellation is
    # cooperative, so ``is_running`` can briefly lag a terminate). Checked
    # AFTER the running override so ``terminated`` is never masked.
    terminated_at = record.terminated_at
    terminated = terminated_at is not None
    if terminated:
        status = "terminated"
    if not has_user_event and not running:
        created_at = None
    if not title:
        title = f"Session {session_id[:8]}"
    merged_context = self._session_store.merge_context_events(events)
    origin = SessionOrigin.classify(
        record.tags, merged_context
    )
    # ``recoverable`` = the FE/CLI may offer a Continue/Restart affordance.
    # True when the session is not running, did not complete successfully,
    # and has a prior user turn (so ``resolve_recovery_query`` won't raise).
    # The crucial case is a session that died mid-call with NO ``completion``
    # event at all (process killed) — status stays ``idle`` but a user turn
    # exists, so it must be recoverable.
    # ``awaiting_approval`` is recoverable: a plan-mode proposal (or a wiki QA
    # answer) parks here with no active run and no auto-exit. Recovery
    # ``continue`` re-engages it through the one loop, mirroring send_followup —
    # the only other way out. Only ``completed`` is genuinely terminal.
    # This is why the status vocabulary is where honesty has to be fixed
    # rather than the recovery rule: ``unmet_goal`` and ``blocked`` earn the
    # affordance purely by not being ``completed``, so every run the old
    # table laundered into success had lost it silently.
    recoverable = not running and status != "completed" and has_user_event
    # A terminated session is a hard dead-end — never offer Continue/Restart.
    if terminated:
        recoverable = False
    # Surface the EXERCISED capabilities + the workspace so the landing page
    # can show what a session actually did without re-reading the transcript.
    # Copying ``context.client_capabilities`` verbatim would be wrong: that
    # field is an ADVERTISEMENT, and the console sends the same fixed header
    # on every request, so every chat row would be chipped with four
    # capabilities it never touched.
    # ``CapabilityEvidence`` (fed above, in the one pass) holds each earnable
    # capability to a durable artifact event or a successful invocation of a
    # tool it gates, and passes the server-written scopes (``wiki``/``scg``)
    # through untouched — see its docstring for why the split falls there. A
    # runtime-granted capability is still intentionally NOT probed: it's a
    # live predicate, not a durable signal, already surfaced where it matters
    # (the run's Langfuse trace facet), and a per-row store probe in a session
    # LIST would be the wrong tradeoff.
    capabilities = evidence.resolve(merged_context.get("client_capabilities"))
    workspace = merged_context.get("structured_workspace") or merged_context.get(
        "workspace"
    )
    summary: dict[str, object] = {
        "session_id": session_id,
        "title": title,
        "created_at": created_at,
        "status": status,
        "done_reason": done_reason,
        "running": running,
        "recoverable": recoverable,
        "context": merged_context,
        "origin": origin.value,
        "capabilities": capabilities,
        "workspace": str(workspace) if workspace else None,
        "archived": record.archived,
        "terminated": terminated,
        "terminated_at": terminated_at,
    }
    # ``owner`` is APPENDED, and only when there IS one. An unowned session
    # is genuinely different from one owned by nobody, and reporting the
    # difference as an absent key rather than an explicit null is what keeps
    # this summary key-free for a deployment running without identity: no
    # principal exists there, so nothing is ever stamped, so the key never
    # appears. Appending (never inserting mid-dict) preserves key order too.
    owner = record.owner
    if owner is not None:
        summary["owner"] = owner
    # ``pinned`` follows the same append-when-present rule, and for a third
    # reason on top of the two above: pinning is a minority state, so an
    # unpinned session's summary carries neither key. Both keys travel
    # together — the boolean is what
    # a surface renders, the stamp is what it orders by — so a client never
    # has to infer one from the other. Read off *record* rather than a
    # direct store call — the whole point of batch-loading it is that a
    # listing's per-row cost stops scaling with row count; calling back into
    # the store here would reopen exactly that per-row round trip for these
    # two fields alone while every other field on the row stayed batched.
    if record.pinned_at is not None:
        summary["pinned"] = True
        summary["pinned_at"] = record.pinned_at
    # ``projects`` is the accumulated SET a project filter needs — never just
    # the current ``context.project`` — because an auto-select session may
    # have switched mid-task, and a filter reading only the latest binding
    # would miss every project it moved out of. Same append-when-present
    # rule: a session bound to nothing (a bare temp dir, or still on the
    # ``auto`` sentinel) carries no key at all.
    if record.projects:
        summary["projects"] = record.projects
    # The three failure facets follow the same append-when-present rule, for
    # the same reason: a session that never blocked and never re-tried a
    # model carries none of them.
    #
    # ``blocked_code`` says WHICH wall the run hit — the status alone says a
    # user can act, not what to act on. ``failure_reason`` + ``models_tried``
    # are what let a recovery affordance default AWAY from the model that
    # just failed instead of re-running the same one, which is the whole
    # point of recording them: retrying on a model whose failure was
    # model-specific amplifies the failure.
    if blocked_code is not None:
        summary["blocked_code"] = blocked_code
    if outcome_assertion is not None:
        # Surfaced whenever an assertion was made, even where it did not
        # change the status: a run that failed AND missed its purpose is
        # two facts, and dropping the second would hide the one a product
        # owner can act on.
        summary["unmet_goal_reason"] = outcome_assertion.reason
        if outcome_assertion.detail:
            summary["unmet_goal_detail"] = outcome_assertion.detail
    if failure_reason is not None:
        summary["failure_reason"] = failure_reason
    if models_tried:
        summary["models_tried"] = models_tried
    # Same append-when-present rule: a session that changed no lines carries
    # no ``diff_stat`` at all, so a consumer can treat the key's presence as
    # "this session edited something" without comparing against zero.
    if not diff_stat.is_empty:
        summary["diff_stat"] = {
            "additions": diff_stat.additions,
            "deletions": diff_stat.deletions,
        }
    return summary

tag_session(session_id: str, tag: str) -> None

Associate a provenance/lookup tag with a session.

Thin delegate to the store so callers that already hold a resolved session_id (e.g. a structured/realtime run stamping its origin tag) don't reach into session_store directly. resolve_session remains the seam for tag-keyed resolution; this is the write-only sibling for tagging a session you've already created.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
621
622
623
624
625
626
627
628
629
630
def tag_session(self, session_id: str, tag: str) -> None:
    """Associate a provenance/lookup tag with a session.

    Thin delegate to the store so callers that already hold a resolved
    ``session_id`` (e.g. a structured/realtime run stamping its origin tag)
    don't reach into ``session_store`` directly. ``resolve_session`` remains
    the seam for tag-keyed *resolution*; this is the write-only sibling for
    tagging a session you've already created.
    """
    self._session_store.tag_session(session_id, tag)

terminate_session(session_id: str) -> dict[str, object]

Permanently terminate a session — idempotent and irreversible.

First call: stamps terminated_at, cancels any live run (cooperative, via the run registry's cancel event), fires the on_terminate callbacks (summing any returned cancelled-artifact counts), and appends a session_terminated transcript event — which rides the standard append_event → SessionEventBus → SSE choke-point, so a live stream observes the termination with no new transport. A repeat call is a no-op that returns the SAME shape with the ORIGINAL terminated_at and cancelled_triggers: 0 (the side effects never re-fire).

Side effects fire exactly once, arbitrated by the store's set-once write: SessionStoreBase.terminate_session returns True only to the call that newly stamped the timestamp, so two concurrent FIRST calls (Flask is threaded even at --workers 1) can't both pass an unlocked read and duplicate the event + callback fan-out — exactly one of them wins the race and runs the block below.

Callers guard unknown-session (404) upstream; this assumes the session exists. Returns {session_id, status:"terminated", terminated_at, cancelled_triggers}.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
def terminate_session(self, session_id: str) -> dict[str, object]:
    """Permanently terminate a session — idempotent and irreversible.

    First call: stamps ``terminated_at``, cancels any live run (cooperative,
    via the run registry's cancel event), fires the ``on_terminate``
    callbacks (summing any returned cancelled-artifact counts), and appends
    a ``session_terminated`` transcript event — which rides the standard
    ``append_event`` → ``SessionEventBus`` → SSE choke-point, so a live
    stream observes the termination with no new transport. A repeat call is
    a no-op that returns the SAME shape with the ORIGINAL ``terminated_at``
    and ``cancelled_triggers: 0`` (the side effects never re-fire).

    Side effects fire exactly once, arbitrated by the store's set-once
    write: ``SessionStoreBase.terminate_session`` returns ``True`` only to
    the call that newly stamped the timestamp, so two concurrent FIRST
    calls (Flask is threaded even at ``--workers 1``) can't both pass an
    unlocked read and duplicate the event + callback fan-out — exactly one
    of them wins the race and runs the block below.

    Callers guard unknown-session (404) upstream; this assumes the session
    exists. Returns ``{session_id, status:"terminated", terminated_at,
    cancelled_triggers}``.
    """
    newly_terminated = self._session_store.terminate_session(session_id)
    terminated_at = self._session_store.get_terminated_at(session_id)
    if not newly_terminated:
        return {
            "session_id": session_id,
            "status": "terminated",
            "terminated_at": terminated_at,
            "cancelled_triggers": 0,
        }
    # Cancel any in-flight run so the terminated session stops working; the
    # cancel event unwinds the loop cooperatively (best-effort, no-op if idle).
    self._run_registry.cancel(session_id)
    cancelled_triggers = 0
    for callback in self._on_terminate:
        try:
            result = callback(session_id)
        except Exception:
            logging.warning(
                "on_terminate callback failed for session {}", session_id, exc_info=True
            )
            continue
        if isinstance(result, int):
            cancelled_triggers += result
    # append_terminal_event, NOT append_event: the store already cached
    # this session as terminated (line above), so a guarded append would
    # drop the tombstone event it is itself trying to write.
    self._session_store.append_terminal_event(
        session_id,
        {
            "type": "session_terminated",
            "payload": {
                "session_id": session_id,
                "terminated_at": terminated_at,
                "cancelled_triggers": cancelled_triggers,
            },
        },
    )
    return {
        "session_id": session_id,
        "status": "terminated",
        "terminated_at": terminated_at,
        "cancelled_triggers": cancelled_triggers,
    }

SessionTerminatedError

Bases: ValueError

Raised when an operation targets a permanently terminated session.

Termination is a kill switch: a dead session must not be runnable, steerable, recoverable, or fork-resurrectable — copying a terminated transcript into a fresh session would let an agent launder its way around the kill. Raised at the core seam so every caller inherits enforcement; HTTP surfaces map it to 410 Gone.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
137
138
139
140
141
142
143
144
145
class SessionTerminatedError(ValueError):
    """Raised when an operation targets a permanently terminated session.

    Termination is a kill switch: a dead session must not be runnable,
    steerable, recoverable, *or fork-resurrectable* — copying a terminated
    transcript into a fresh session would let an agent launder its way around
    the kill. Raised at the core seam so every caller inherits enforcement;
    HTTP surfaces map it to 410 Gone.
    """

parse_core_command(text: str) -> str | None

Return the core command token if present.

Source code in packages/mewbo_core/src/mewbo_core/loop/session_runtime.py
160
161
162
163
164
165
def parse_core_command(text: str) -> str | None:
    """Return the core command token if present."""
    if not text:
        return None
    command = text.strip().lower().split()[0]
    return command if command in CORE_COMMANDS else None

mewbo_core.session.session_store

Session transcript storage and management.

Provides a SessionStoreBase ABC, a filesystem-backed SessionStore implementation, and a create_session_store() factory that returns the configured driver (json or mongodb).

SessionPaths dataclass

Resolved filesystem paths for a session.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
@dataclass(frozen=True)
class SessionPaths:
    """Resolved filesystem paths for a session."""

    root: str
    session_id: str

    @property
    def session_dir(self) -> str:
        """Directory for session artifacts."""
        return os.path.join(self.root, self.session_id)

    @property
    def transcript_path(self) -> str:
        """Path to the JSONL transcript file."""
        return os.path.join(self.session_dir, "transcript.jsonl")

    @property
    def summary_path(self) -> str:
        """Path to the summary JSON file."""
        return os.path.join(self.session_dir, "summary.json")

    @property
    def title_path(self) -> str:
        """Path to the title JSON file."""
        return os.path.join(self.session_dir, "title.json")

session_dir: str property

Directory for session artifacts.

summary_path: str property

Path to the summary JSON file.

title_path: str property

Path to the title JSON file.

transcript_path: str property

Path to the JSONL transcript file.

SessionRecord dataclass

The per-session metadata a summary row needs, minus the transcript.

Everything here is stored ON the session record rather than derived by folding events, which is exactly why it can be batch-loaded: a listing needs all of it for every row and none of it depends on reading a transcript.

A plain frozen dataclass, not a Pydantic model: it never crosses a trust boundary — the store builds it from data it just read out of its own backend — so validating each field would buy nothing on a path whose entire purpose is to be cheap.

tags is the reverse tag index for this session (the same list tags_for_session returns), carried here so a batch resolves the index once instead of per row.

pinned_at/projects are the two filterable facets (session/CLAUDE.md → "Filterable facets") and belong here for the same reason as every other field: both are stored ON the record rather than folded from events, so a caller reading them one session at a time is paying the exact per-row round-trip cost this class exists to collapse.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
@dataclass(frozen=True)
class SessionRecord:
    """The per-session metadata a summary row needs, minus the transcript.

    Everything here is stored ON the session record rather than derived by
    folding events, which is exactly why it can be batch-loaded: a listing needs
    all of it for every row and none of it depends on reading a transcript.

    A plain frozen dataclass, not a Pydantic model: it never crosses a trust
    boundary — the store builds it from data it just read out of its own
    backend — so validating each field would buy nothing on a path whose entire
    purpose is to be cheap.

    ``tags`` is the reverse tag index for this session (the same list
    ``tags_for_session`` returns), carried here so a batch resolves the index
    once instead of per row.

    ``pinned_at``/``projects`` are the two filterable facets
    (``session/CLAUDE.md`` → "Filterable facets") and belong here for the same
    reason as every other field: both are stored ON the record rather than
    folded from events, so a caller reading them one session at a time is
    paying the exact per-row round-trip cost this class exists to collapse.
    """

    title: str | None = None
    archived: bool = False
    terminated_at: str | None = None
    owner: str | None = None
    tags: list[str] = field(default_factory=list)
    pinned_at: str | None = None
    projects: list[str] = field(default_factory=list)

SessionStore

Bases: SessionStoreBase

Filesystem-backed storage for session transcripts and summaries.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
class SessionStore(SessionStoreBase):
    """Filesystem-backed storage for session transcripts and summaries."""

    def __init__(self, root_dir: str | None = None) -> None:
        """Initialize the store and ensure the root directory exists."""
        super().__init__()
        if root_dir is None:
            root_dir = get_config_value("runtime", "session_dir", default="./data/sessions")
        self.root_dir = os.path.abspath(root_dir)
        os.makedirs(self.root_dir, exist_ok=True)

    def _index_path(self) -> str:
        """Return the path for the session index file."""
        return os.path.join(self.root_dir, "index.json")

    def _load_index(self) -> dict[str, dict[str, Any]]:
        """Load the session index from disk or return defaults.

        Bucket values are ``Any`` rather than ``str`` because ``projects`` maps a
        session to a LIST of identities while every other bucket maps it to a
        single stamp. A bucket missing from an index written before it existed
        reads as absent and is defaulted at each use site, so no migration is
        needed to start reading one.
        """
        index_path = self._index_path()
        if not os.path.exists(index_path):
            return {
                "tags": {},
                "archived": {},
                "terminated": {},
                "owners": {},
                "pinned": {},
                "projects": {},
            }
        with open(index_path, encoding="utf-8") as handle:
            return json.load(handle)

    def _save_index(self, data: dict[str, dict[str, Any]]) -> None:
        """Persist the session index to disk."""
        with open(self._index_path(), "w", encoding="utf-8") as handle:
            json.dump(data, handle, indent=2)

    def create_session(self, owner: str | None = None) -> str:
        """Create a new session directory and return its identifier."""
        session_id = uuid.uuid4().hex
        self.ensure_session(session_id, owner)
        return session_id

    def ensure_session(self, session_id: str, owner: str | None = None) -> None:
        """Idempotently create the session directory for a known id.

        For the filesystem driver a session "record" IS its directory (that's
        what :meth:`list_sessions` enumerates), so materialisation is a
        ``makedirs`` — ``exist_ok=True`` makes it a safe no-op on replay.

        The owner lands in its own index bucket beside ``archived``/``terminated``
        rather than in a per-session file, so the list filter costs one index read
        instead of one stat per session.

        The stamp is written ONLY when this call is the one that materialised the
        record — the filesystem analogue of the Mongo driver's ``$setOnInsert``,
        and the reason the directory's prior existence is checked before creating
        it. Stamping on any call would make an ALREADY-UNOWNED session claimable
        by whoever touched it next, and since an unowned session is visible to
        every lister, that is not necessarily its creator. An unowned call never
        writes at all, so a deployment with no identity configured leaves the
        index exactly as it was.
        """
        paths = self._paths(session_id)
        already_materialised = os.path.isdir(paths.session_dir)
        os.makedirs(paths.session_dir, exist_ok=True)
        if owner is None or already_materialised:
            return
        index = self._load_index()
        index.setdefault("owners", {})[session_id] = owner
        self._save_index(index)

    def get_owner(self, session_id: str) -> str | None:
        """Return the stamped owning subject, or ``None`` if unowned."""
        return self._load_index().get("owners", {}).get(session_id)

    def _paths(self, session_id: str) -> SessionPaths:
        """Build filesystem paths for a session."""
        return SessionPaths(root=self.root_dir, session_id=session_id)

    def session_dir(self, session_id: str) -> str:
        """Return the directory path for a session."""
        return self._paths(session_id).session_dir

    def _write_event(self, session_id: str, event: Event) -> None:
        """Append a single event record to the session transcript, unconditionally."""
        paths = self._paths(session_id)
        os.makedirs(paths.session_dir, exist_ok=True)
        payload: EventRecord = self.stamp(event)
        with open(paths.transcript_path, "a", encoding="utf-8") as handle:
            handle.write(json.dumps(payload) + "\n")
        self._publish_appended(session_id, payload)

    def load_transcript(self, session_id: str) -> list[EventRecord]:
        """Load all transcript events for a session. ``O(one session's events)``."""
        return list(self.stream_transcript(session_id))

    def stream_transcript(self, session_id: str) -> Iterator[EventRecord]:
        """Yield transcript events one parsed line at a time.

        ``O(1)`` in memory, which is what makes the base
        :meth:`~SessionStoreBase.list_session_digests` template usable here: a
        listing folds each session's file as it reads and retains only the
        relevant events, instead of holding a whole transcript to produce one
        row. The PARSE still costs ``O(one session's events)`` — a JSONL file
        carries no index, so there is no honest way to find a session's
        ``completion`` event without reading past everything before it, and this
        driver's listing therefore stays ``O(all history)`` in CPU. That is the
        floor for the default driver, and the reason MongoDB is the recommended
        backend for a store that has accumulated real history.
        """
        paths = self._paths(session_id)
        if not os.path.exists(paths.transcript_path):
            return
        with open(paths.transcript_path, encoding="utf-8") as handle:
            for line in handle:
                line = line.strip()
                if not line:
                    continue
                try:
                    yield json.loads(line)
                except json.JSONDecodeError:
                    logging.warning("Skipping malformed transcript line.")

    def truncate_after(self, session_id: str, cutoff_ts: str) -> int:
        """Rewrite the transcript keeping only events with ``ts <= cutoff_ts``."""
        paths = self._paths(session_id)
        if not os.path.exists(paths.transcript_path):
            return 0
        events = self.load_transcript(session_id)
        kept = [e for e in events if e.get("ts", "") <= cutoff_ts]
        removed = len(events) - len(kept)
        if removed:
            with open(paths.transcript_path, "w", encoding="utf-8") as handle:
                for event in kept:
                    handle.write(json.dumps(event) + "\n")
        return removed

    def save_summary(self, session_id: str, summary: str) -> None:
        """Persist a summary for a session."""
        paths = self._paths(session_id)
        os.makedirs(paths.session_dir, exist_ok=True)
        with open(paths.summary_path, "w", encoding="utf-8") as handle:
            json.dump({"summary": summary, "updated_at": _utc_now()}, handle, indent=2)

    def load_summary(self, session_id: str) -> str | None:
        """Load a previously saved summary, if present."""
        paths = self._paths(session_id)
        if not os.path.exists(paths.summary_path):
            return None
        with open(paths.summary_path, encoding="utf-8") as handle:
            data = json.load(handle)
        return data.get("summary")

    def save_title(self, session_id: str, title: str) -> None:
        """Persist a display title for a session."""
        paths = self._paths(session_id)
        os.makedirs(paths.session_dir, exist_ok=True)
        with open(paths.title_path, "w", encoding="utf-8") as handle:
            json.dump({"title": title, "updated_at": _utc_now()}, handle, indent=2)

    def load_title(self, session_id: str) -> str | None:
        """Load a previously saved title, if present."""
        paths = self._paths(session_id)
        if not os.path.exists(paths.title_path):
            return None
        with open(paths.title_path, encoding="utf-8") as handle:
            data = json.load(handle)
        title = data.get("title")
        return title if isinstance(title, str) and title else None

    def query_sessions(self, query: SessionQuery) -> list[str]:
        """List session IDs in the root directory, narrowed by *query*.

        Reads the index ONCE for the whole listing rather than once per
        predicate per session — the filesystem driver's every index-backed fact
        is a full read-and-parse of ``index.json``, so a per-session lookup would
        turn one file read into four times the session count.
        """
        if not os.path.exists(self.root_dir):
            return []
        session_ids = sorted(
            name
            for name in os.listdir(self.root_dir)
            if os.path.isdir(os.path.join(self.root_dir, name))
        )
        if query.is_unfiltered and query.include_archived:
            return session_ids
        index = self._load_index()
        owners = index.get("owners", {})
        archived = index.get("archived", {})
        pinned = index.get("pinned", {})
        projects = index.get("projects", {})
        return [
            sid
            for sid in session_ids
            if query.matches_record(
                # An id absent from the owners bucket is UNOWNED and stays
                # visible — see the base's contract for why that is the
                # migration semantic rather than a hole.
                owner=owners.get(sid),
                archived=sid in archived,
                pinned=sid in pinned,
                projects=list(projects.get(sid, [])),
            )
        ]

    def set_pinned(self, session_id: str, pinned: bool) -> None:
        """Stamp or clear ``pinned`` for a session in the index."""
        index = self._load_index()
        bucket = index.setdefault("pinned", {})
        if pinned:
            bucket[session_id] = _utc_now()
        elif session_id not in bucket:
            return  # Nothing to clear — don't rewrite the file for a no-op.
        else:
            bucket.pop(session_id, None)
        self._save_index(index)

    def get_pinned_at(self, session_id: str) -> str | None:
        """Return the stored ``pinned`` timestamp, or ``None``."""
        index = self._load_index()
        stamp = index.get("pinned", {}).get(session_id)
        return stamp if isinstance(stamp, str) else None

    def record_project(self, session_id: str, project: str) -> None:
        """Append *project* to the session's recorded set, if new.

        Returns without writing when the project is already recorded, which is
        the common case: every turn of a bound session re-emits the same context.
        """
        index = self._load_index()
        bucket = index.setdefault("projects", {})
        current = list(bucket.get(session_id, []))
        if project in current:
            return
        current.append(project)
        bucket[session_id] = current
        self._save_index(index)

    def projects_for_session(self, session_id: str) -> list[str]:
        """Return every project identity recorded for a session."""
        index = self._load_index()
        return list(index.get("projects", {}).get(session_id, []))

    def tag_session(self, session_id: str, tag: str) -> None:
        """Associate a tag with a session ID for quick lookup."""
        index = self._load_index()
        index.setdefault("tags", {})[tag] = session_id
        self._save_index(index)

    def resolve_tag(self, tag: str) -> str | None:
        """Resolve a tag to a session ID, if present."""
        index = self._load_index()
        return index.get("tags", {}).get(tag)

    def list_tags(self) -> dict[str, str]:
        """Return a mapping of tags to session IDs."""
        index = self._load_index()
        return dict(index.get("tags", {}))

    def archive_session(self, session_id: str) -> None:
        """Mark a session as archived."""
        index = self._load_index()
        archived = index.setdefault("archived", {})
        archived[session_id] = _utc_now()
        self._save_index(index)

    def unarchive_session(self, session_id: str) -> None:
        """Remove archived status from a session."""
        index = self._load_index()
        archived = index.get("archived", {})
        if session_id in archived:
            archived.pop(session_id, None)
            index["archived"] = archived
            self._save_index(index)

    def is_archived(self, session_id: str) -> bool:
        """Return True if a session is archived."""
        index = self._load_index()
        archived = index.get("archived", {})
        return session_id in archived

    def terminate_session(self, session_id: str) -> bool:
        """Stamp ``terminated_at`` in the index once (never overwrites).

        Returns whether THIS call inserted the stamp — checked before the
        write so a repeat call reports ``False`` without touching the file.
        """
        index = self._load_index()
        terminated = index.setdefault("terminated", {})
        if session_id in terminated:
            self._mark_terminated_cached(session_id)
            return False
        terminated[session_id] = _utc_now()
        self._save_index(index)
        self._mark_terminated_cached(session_id)
        return True

    def get_terminated_at(self, session_id: str) -> str | None:
        """Return the stored ``terminated_at`` timestamp, or ``None``."""
        index = self._load_index()
        return index.get("terminated", {}).get(session_id)

__init__(root_dir: str | None = None) -> None

Initialize the store and ensure the root directory exists.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
818
819
820
821
822
823
824
def __init__(self, root_dir: str | None = None) -> None:
    """Initialize the store and ensure the root directory exists."""
    super().__init__()
    if root_dir is None:
        root_dir = get_config_value("runtime", "session_dir", default="./data/sessions")
    self.root_dir = os.path.abspath(root_dir)
    os.makedirs(self.root_dir, exist_ok=True)

archive_session(session_id: str) -> None

Mark a session as archived.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1081
1082
1083
1084
1085
1086
def archive_session(self, session_id: str) -> None:
    """Mark a session as archived."""
    index = self._load_index()
    archived = index.setdefault("archived", {})
    archived[session_id] = _utc_now()
    self._save_index(index)

create_session(owner: str | None = None) -> str

Create a new session directory and return its identifier.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
857
858
859
860
861
def create_session(self, owner: str | None = None) -> str:
    """Create a new session directory and return its identifier."""
    session_id = uuid.uuid4().hex
    self.ensure_session(session_id, owner)
    return session_id

ensure_session(session_id: str, owner: str | None = None) -> None

Idempotently create the session directory for a known id.

For the filesystem driver a session "record" IS its directory (that's what :meth:list_sessions enumerates), so materialisation is a makedirs — exist_ok=True makes it a safe no-op on replay.

The owner lands in its own index bucket beside archived/terminated rather than in a per-session file, so the list filter costs one index read instead of one stat per session.

The stamp is written ONLY when this call is the one that materialised the record — the filesystem analogue of the Mongo driver's $setOnInsert, and the reason the directory's prior existence is checked before creating it. Stamping on any call would make an ALREADY-UNOWNED session claimable by whoever touched it next, and since an unowned session is visible to every lister, that is not necessarily its creator. An unowned call never writes at all, so a deployment with no identity configured leaves the index exactly as it was.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
def ensure_session(self, session_id: str, owner: str | None = None) -> None:
    """Idempotently create the session directory for a known id.

    For the filesystem driver a session "record" IS its directory (that's
    what :meth:`list_sessions` enumerates), so materialisation is a
    ``makedirs`` — ``exist_ok=True`` makes it a safe no-op on replay.

    The owner lands in its own index bucket beside ``archived``/``terminated``
    rather than in a per-session file, so the list filter costs one index read
    instead of one stat per session.

    The stamp is written ONLY when this call is the one that materialised the
    record — the filesystem analogue of the Mongo driver's ``$setOnInsert``,
    and the reason the directory's prior existence is checked before creating
    it. Stamping on any call would make an ALREADY-UNOWNED session claimable
    by whoever touched it next, and since an unowned session is visible to
    every lister, that is not necessarily its creator. An unowned call never
    writes at all, so a deployment with no identity configured leaves the
    index exactly as it was.
    """
    paths = self._paths(session_id)
    already_materialised = os.path.isdir(paths.session_dir)
    os.makedirs(paths.session_dir, exist_ok=True)
    if owner is None or already_materialised:
        return
    index = self._load_index()
    index.setdefault("owners", {})[session_id] = owner
    self._save_index(index)

get_owner(session_id: str) -> str | None

Return the stamped owning subject, or None if unowned.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
892
893
894
def get_owner(self, session_id: str) -> str | None:
    """Return the stamped owning subject, or ``None`` if unowned."""
    return self._load_index().get("owners", {}).get(session_id)

get_pinned_at(session_id: str) -> str | None

Return the stored pinned timestamp, or None.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1039
1040
1041
1042
1043
def get_pinned_at(self, session_id: str) -> str | None:
    """Return the stored ``pinned`` timestamp, or ``None``."""
    index = self._load_index()
    stamp = index.get("pinned", {}).get(session_id)
    return stamp if isinstance(stamp, str) else None

get_terminated_at(session_id: str) -> str | None

Return the stored terminated_at timestamp, or None.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1119
1120
1121
1122
def get_terminated_at(self, session_id: str) -> str | None:
    """Return the stored ``terminated_at`` timestamp, or ``None``."""
    index = self._load_index()
    return index.get("terminated", {}).get(session_id)

is_archived(session_id: str) -> bool

Return True if a session is archived.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1097
1098
1099
1100
1101
def is_archived(self, session_id: str) -> bool:
    """Return True if a session is archived."""
    index = self._load_index()
    archived = index.get("archived", {})
    return session_id in archived

list_tags() -> dict[str, str]

Return a mapping of tags to session IDs.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1076
1077
1078
1079
def list_tags(self) -> dict[str, str]:
    """Return a mapping of tags to session IDs."""
    index = self._load_index()
    return dict(index.get("tags", {}))

load_summary(session_id: str) -> str | None

Load a previously saved summary, if present.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
965
966
967
968
969
970
971
972
def load_summary(self, session_id: str) -> str | None:
    """Load a previously saved summary, if present."""
    paths = self._paths(session_id)
    if not os.path.exists(paths.summary_path):
        return None
    with open(paths.summary_path, encoding="utf-8") as handle:
        data = json.load(handle)
    return data.get("summary")

load_title(session_id: str) -> str | None

Load a previously saved title, if present.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
981
982
983
984
985
986
987
988
989
def load_title(self, session_id: str) -> str | None:
    """Load a previously saved title, if present."""
    paths = self._paths(session_id)
    if not os.path.exists(paths.title_path):
        return None
    with open(paths.title_path, encoding="utf-8") as handle:
        data = json.load(handle)
    title = data.get("title")
    return title if isinstance(title, str) and title else None

load_transcript(session_id: str) -> list[EventRecord]

Load all transcript events for a session. O(one session's events).

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
913
914
915
def load_transcript(self, session_id: str) -> list[EventRecord]:
    """Load all transcript events for a session. ``O(one session's events)``."""
    return list(self.stream_transcript(session_id))

projects_for_session(session_id: str) -> list[str]

Return every project identity recorded for a session.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1060
1061
1062
1063
def projects_for_session(self, session_id: str) -> list[str]:
    """Return every project identity recorded for a session."""
    index = self._load_index()
    return list(index.get("projects", {}).get(session_id, []))

query_sessions(query: SessionQuery) -> list[str]

List session IDs in the root directory, narrowed by query.

Reads the index ONCE for the whole listing rather than once per predicate per session — the filesystem driver's every index-backed fact is a full read-and-parse of index.json, so a per-session lookup would turn one file read into four times the session count.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
def query_sessions(self, query: SessionQuery) -> list[str]:
    """List session IDs in the root directory, narrowed by *query*.

    Reads the index ONCE for the whole listing rather than once per
    predicate per session — the filesystem driver's every index-backed fact
    is a full read-and-parse of ``index.json``, so a per-session lookup would
    turn one file read into four times the session count.
    """
    if not os.path.exists(self.root_dir):
        return []
    session_ids = sorted(
        name
        for name in os.listdir(self.root_dir)
        if os.path.isdir(os.path.join(self.root_dir, name))
    )
    if query.is_unfiltered and query.include_archived:
        return session_ids
    index = self._load_index()
    owners = index.get("owners", {})
    archived = index.get("archived", {})
    pinned = index.get("pinned", {})
    projects = index.get("projects", {})
    return [
        sid
        for sid in session_ids
        if query.matches_record(
            # An id absent from the owners bucket is UNOWNED and stays
            # visible — see the base's contract for why that is the
            # migration semantic rather than a hole.
            owner=owners.get(sid),
            archived=sid in archived,
            pinned=sid in pinned,
            projects=list(projects.get(sid, [])),
        )
    ]

record_project(session_id: str, project: str) -> None

Append project to the session's recorded set, if new.

Returns without writing when the project is already recorded, which is the common case: every turn of a bound session re-emits the same context.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
def record_project(self, session_id: str, project: str) -> None:
    """Append *project* to the session's recorded set, if new.

    Returns without writing when the project is already recorded, which is
    the common case: every turn of a bound session re-emits the same context.
    """
    index = self._load_index()
    bucket = index.setdefault("projects", {})
    current = list(bucket.get(session_id, []))
    if project in current:
        return
    current.append(project)
    bucket[session_id] = current
    self._save_index(index)

resolve_tag(tag: str) -> str | None

Resolve a tag to a session ID, if present.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1071
1072
1073
1074
def resolve_tag(self, tag: str) -> str | None:
    """Resolve a tag to a session ID, if present."""
    index = self._load_index()
    return index.get("tags", {}).get(tag)

save_summary(session_id: str, summary: str) -> None

Persist a summary for a session.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
958
959
960
961
962
963
def save_summary(self, session_id: str, summary: str) -> None:
    """Persist a summary for a session."""
    paths = self._paths(session_id)
    os.makedirs(paths.session_dir, exist_ok=True)
    with open(paths.summary_path, "w", encoding="utf-8") as handle:
        json.dump({"summary": summary, "updated_at": _utc_now()}, handle, indent=2)

save_title(session_id: str, title: str) -> None

Persist a display title for a session.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
974
975
976
977
978
979
def save_title(self, session_id: str, title: str) -> None:
    """Persist a display title for a session."""
    paths = self._paths(session_id)
    os.makedirs(paths.session_dir, exist_ok=True)
    with open(paths.title_path, "w", encoding="utf-8") as handle:
        json.dump({"title": title, "updated_at": _utc_now()}, handle, indent=2)

session_dir(session_id: str) -> str

Return the directory path for a session.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
900
901
902
def session_dir(self, session_id: str) -> str:
    """Return the directory path for a session."""
    return self._paths(session_id).session_dir

set_pinned(session_id: str, pinned: bool) -> None

Stamp or clear pinned for a session in the index.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
def set_pinned(self, session_id: str, pinned: bool) -> None:
    """Stamp or clear ``pinned`` for a session in the index."""
    index = self._load_index()
    bucket = index.setdefault("pinned", {})
    if pinned:
        bucket[session_id] = _utc_now()
    elif session_id not in bucket:
        return  # Nothing to clear — don't rewrite the file for a no-op.
    else:
        bucket.pop(session_id, None)
    self._save_index(index)

stream_transcript(session_id: str) -> Iterator[EventRecord]

Yield transcript events one parsed line at a time.

O(1) in memory, which is what makes the base :meth:~SessionStoreBase.list_session_digests template usable here: a listing folds each session's file as it reads and retains only the relevant events, instead of holding a whole transcript to produce one row. The PARSE still costs O(one session's events) — a JSONL file carries no index, so there is no honest way to find a session's completion event without reading past everything before it, and this driver's listing therefore stays O(all history) in CPU. That is the floor for the default driver, and the reason MongoDB is the recommended backend for a store that has accumulated real history.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
def stream_transcript(self, session_id: str) -> Iterator[EventRecord]:
    """Yield transcript events one parsed line at a time.

    ``O(1)`` in memory, which is what makes the base
    :meth:`~SessionStoreBase.list_session_digests` template usable here: a
    listing folds each session's file as it reads and retains only the
    relevant events, instead of holding a whole transcript to produce one
    row. The PARSE still costs ``O(one session's events)`` — a JSONL file
    carries no index, so there is no honest way to find a session's
    ``completion`` event without reading past everything before it, and this
    driver's listing therefore stays ``O(all history)`` in CPU. That is the
    floor for the default driver, and the reason MongoDB is the recommended
    backend for a store that has accumulated real history.
    """
    paths = self._paths(session_id)
    if not os.path.exists(paths.transcript_path):
        return
    with open(paths.transcript_path, encoding="utf-8") as handle:
        for line in handle:
            line = line.strip()
            if not line:
                continue
            try:
                yield json.loads(line)
            except json.JSONDecodeError:
                logging.warning("Skipping malformed transcript line.")

tag_session(session_id: str, tag: str) -> None

Associate a tag with a session ID for quick lookup.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1065
1066
1067
1068
1069
def tag_session(self, session_id: str, tag: str) -> None:
    """Associate a tag with a session ID for quick lookup."""
    index = self._load_index()
    index.setdefault("tags", {})[tag] = session_id
    self._save_index(index)

terminate_session(session_id: str) -> bool

Stamp terminated_at in the index once (never overwrites).

Returns whether THIS call inserted the stamp — checked before the write so a repeat call reports False without touching the file.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
def terminate_session(self, session_id: str) -> bool:
    """Stamp ``terminated_at`` in the index once (never overwrites).

    Returns whether THIS call inserted the stamp — checked before the
    write so a repeat call reports ``False`` without touching the file.
    """
    index = self._load_index()
    terminated = index.setdefault("terminated", {})
    if session_id in terminated:
        self._mark_terminated_cached(session_id)
        return False
    terminated[session_id] = _utc_now()
    self._save_index(index)
    self._mark_terminated_cached(session_id)
    return True

truncate_after(session_id: str, cutoff_ts: str) -> int

Rewrite the transcript keeping only events with ts <= cutoff_ts.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
944
945
946
947
948
949
950
951
952
953
954
955
956
def truncate_after(self, session_id: str, cutoff_ts: str) -> int:
    """Rewrite the transcript keeping only events with ``ts <= cutoff_ts``."""
    paths = self._paths(session_id)
    if not os.path.exists(paths.transcript_path):
        return 0
    events = self.load_transcript(session_id)
    kept = [e for e in events if e.get("ts", "") <= cutoff_ts]
    removed = len(events) - len(kept)
    if removed:
        with open(paths.transcript_path, "w", encoding="utf-8") as handle:
            for event in kept:
                handle.write(json.dumps(event) + "\n")
    return removed

unarchive_session(session_id: str) -> None

Remove archived status from a session.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1088
1089
1090
1091
1092
1093
1094
1095
def unarchive_session(self, session_id: str) -> None:
    """Remove archived status from a session."""
    index = self._load_index()
    archived = index.get("archived", {})
    if session_id in archived:
        archived.pop(session_id, None)
        index["archived"] = archived
        self._save_index(index)

SessionStoreBase

Bases: ABC

Abstract interface for session storage backends.

Each driver implements the 13 abstract storage primitives. Higher-level operations (fork_session, load_recent_events, compact_session) are concrete template methods built on top of those primitives.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
class SessionStoreBase(abc.ABC):
    """Abstract interface for session storage backends.

    Each driver implements the 13 abstract storage primitives.  Higher-level
    operations (``fork_session``, ``load_recent_events``, ``compact_session``)
    are concrete template methods built on top of those primitives.
    """

    root_dir: str  # local directory — always present (used for attachment files)

    def __init__(self) -> None:
        """Initialise the in-memory termination cache shared by every backend.

        Populated by :meth:`_mark_terminated_cached` (called from each
        backend's ``terminate_session``) and consulted by
        :meth:`_guard_append`, so a hot-path append never pays a Mongo
        round-trip / index-file read to learn whether its own session was
        just terminated. Best-effort and per-process only: a session
        terminated by a DIFFERENT process/worker is still caught by the
        durable ``terminated_at`` stamp at every other termination-aware seam
        (``resolve_session``, recovery) — this cache only shortcuts the
        common same-process race between a live run and a concurrent
        ``/terminate`` call.
        """
        self._terminated_cache: set[str] = set()
        self._terminated_logged: set[str] = set()

    def _mark_terminated_cached(self, session_id: str) -> None:
        """Record *session_id* as terminated in the in-memory cache."""
        self._terminated_cache.add(session_id)

    def _guard_append(self, session_id: str) -> bool:
        """Return True iff *session_id* may still accept an appended event.

        Consults ONLY the in-memory cache (never a store round-trip) — see
        :meth:`__init__`. Structured-logs once per session (not once per
        dropped event) the first time an append is refused.
        """
        if session_id not in self._terminated_cache:
            return True
        if session_id not in self._terminated_logged:
            self._terminated_logged.add(session_id)
            logging.warning("Dropping event(s) appended to terminated session {}.", session_id)
        return False

    # -- abstract primitives ------------------------------------------------

    @abc.abstractmethod
    def create_session(self, owner: str | None = None) -> str:
        """Create a new session and return its identifier.

        *owner* is an OPAQUE subject string, never an identity object: core sits
        below the identity kernel in the dependency DAG and must not learn what a
        principal is. The caller that HAS one resolves it to a string first.
        ``None`` means unowned, which is what every session created without an
        authenticated caller is — see :meth:`list_sessions` for what that implies
        on the read side.
        """

    @abc.abstractmethod
    def ensure_session(self, session_id: str, owner: str | None = None) -> None:
        """Idempotently materialise a session RECORD for a known id.

        ``create_session`` mints its own uuid; some callers (the realtime
        write-behind recorder) need to back a session whose id was pre-minted
        elsewhere — e.g. for an in-flight Langfuse trace opened before any store
        write. Without a record, ``list_sessions`` (and every read surface built
        on it) never sees the id, so the transcript is an orphan: events exist,
        the session is invisible.

        This is the seam that closes that gap. It is idempotent — calling it on
        an existing session is a no-op (never resets created_at / archived state).
        ``create_session`` is implemented on top of it (mint id → materialise).

        The owner stamp is SET-ONCE at materialisation, like ``created_at`` and
        unlike ``archived_at``: a replay must not be able to re-point an existing
        session at a different subject, which would be an ownership takeover
        written through an idempotent no-op path.
        """

    @abc.abstractmethod
    def _write_event(self, session_id: str, event: Event) -> None:
        """Durably write *event*, unconditionally (no termination guard).

        The shared primitive :meth:`append_event` (guarded) and
        :meth:`append_terminal_event` (the termination tombstone's one
        exemption) both write through — the actual per-backend I/O
        (file append / Mongo insert) plus the ``_publish_appended`` fan-out.
        """

    def append_event(self, session_id: str, event: Event) -> None:
        """Append a single event record to the session transcript.

        No-ops (after a dropped-event log, once per session) once
        :meth:`terminate_session` has stamped this session terminated — see
        :meth:`_guard_append`. The ONE deliberate exception is
        :meth:`append_terminal_event`, used for the termination tombstone
        itself.
        """
        if not self._guard_append(session_id):
            return
        self._write_event(session_id, event)
        self._record_context_project(session_id, event)

    def _record_context_project(self, session_id: str, event: Event) -> None:
        """Accumulate the project a ``context`` event declares onto the record.

        This is the ONE seam that keeps a session's project set current, and it
        works because both writes that can bind a project already funnel through
        here: session creation writes a ``context`` event carrying ``project``,
        and ``ToolUseLoop.rebind_workspace`` writes a mid-run switch as a MERGED
        context event. So an auto-select session that moves between projects
        accumulates each one with no second write site to keep in step.

        Best-effort by construction: a facet that fails to record must never
        break the append it rode in on, so the failure is logged and swallowed
        rather than raised into the caller's path. The cost of losing one is a
        session missing from a project filter, not a lost transcript event.
        """
        if event.get("type") != "context":
            return
        payload = event.get("payload")
        if not isinstance(payload, dict):
            return
        project = ProjectIdentity.from_context(payload)
        if project is None:
            return
        try:
            self.record_project(session_id, project)
        except Exception:
            logging.warning(
                "Failed to record project facet for session.", exc_info=True
            )

    def append_user_turn(
        self,
        session_id: str,
        text: str,
        attachments: list[dict] | None = None,
    ) -> None:
        """Append the ``user`` event that records ONE accepted turn.

        The single builder of that payload, because two seams write it: the
        acceptance seam (``SessionRuntime.start_async``, so the turn is durable
        before the executor's cold start) and the orchestration body itself (for
        every caller that reaches ``Orchestrator`` directly). A second
        hand-rolled ``{"type": "user", ...}`` literal in either place is free to
        drift from this one — silently, since both would still write a
        transcript event the readers accept.

        ``attachments`` is written only when non-empty, mirroring
        :class:`~mewbo_core.contracts.types.UserPayload`'s ``NotRequired`` key,
        so a turn without attachments carries no such key. It additively
        duplicates the sibling ``context`` event's
        descriptors so a client can render attachment cards above the user turn
        without joining across events.
        """
        payload: UserPayload = {"text": text}
        if attachments:
            payload["attachments"] = attachments
        self.append_event(session_id, {"type": "user", "payload": payload})

    def append_terminal_event(self, session_id: str, event: Event) -> None:
        """Append an event exempt from the termination guard.

        ``SessionRuntime.terminate_session`` writes the ``session_terminated``
        tombstone through this seam immediately after the store's own
        ``terminate_session`` has already cached the session as terminated —
        an ordinary :meth:`append_event` call would otherwise drop its own
        closing event. No other caller should use this.
        """
        self._write_event(session_id, event)

    @abc.abstractmethod
    def load_transcript(self, session_id: str) -> list[EventRecord]:
        """Load all transcript events for a session. ``O(one session's events)``.

        A PER-SESSION read. It must never be called in a loop over sessions —
        that is ``O(all history)`` wearing a listing's clothes, and it is exactly
        the shape :meth:`list_session_digests` exists to replace.
        """

    @abc.abstractmethod
    def save_summary(self, session_id: str, summary: str) -> None:
        """Persist a summary for a session."""

    @abc.abstractmethod
    def load_summary(self, session_id: str) -> str | None:
        """Load a previously saved summary, if present."""

    @abc.abstractmethod
    def save_title(self, session_id: str, title: str) -> None:
        """Persist a display title for a session."""

    @abc.abstractmethod
    def load_title(self, session_id: str) -> str | None:
        """Load a previously saved title, if present."""

    @abc.abstractmethod
    def get_owner(self, session_id: str) -> str | None:
        """Return the owning subject, or ``None`` if the session is unowned.

        The read-through sibling of the stamp ``ensure_session`` writes, mirroring
        how ``is_archived`` reads ``archive_session``'s write.
        """

    @abc.abstractmethod
    def session_dir(self, session_id: str) -> str:
        """Return the local directory path for a session (used for attachments)."""

    @abc.abstractmethod
    def tag_session(self, session_id: str, tag: str) -> None:
        """Associate a tag with a session ID for quick lookup."""

    @abc.abstractmethod
    def resolve_tag(self, tag: str) -> str | None:
        """Resolve a tag to a session ID, if present."""

    @abc.abstractmethod
    def list_tags(self) -> dict[str, str]:
        """Return a mapping of tags to session IDs."""

    def tags_for_session(self, session_id: str) -> list[str]:
        """Return every tag pointing at ``session_id`` (reverse of ``resolve_tag``).

        Concrete default over ``list_tags`` so all backends share one
        implementation; the tag set is small. Provenance classification reads
        this to recover a session's origin (see ``session_provenance``).
        """
        return [tag for tag, sid in self.list_tags().items() if sid == session_id]

    def load_session_records(self, session_ids: list[str]) -> dict[str, SessionRecord]:
        """Batch-load the per-session metadata one summary row needs.

        A LISTING derives every row's title, archived flag, termination stamp,
        owner, tags, pin stamp and project set. Fetching them one session at a
        time is seven store reads per row — on a networked driver, seven ROUND
        TRIPS per row, so the cost of listing scales with the session count
        times a per-call latency that has nothing to do with how much data is
        involved. This is the seam that collapses them: one call for the whole
        page.

        This default is the CORRECT-everywhere implementation, not the fast one.
        It still reads per session, but it already removes the worst repetition
        by resolving the reverse tag index ONCE for the whole batch instead of
        rescanning it per row. A driver that can answer the whole batch in a
        single query overrides this (see ``MongoSessionStore``); a driver that
        cannot inherits behaviour identical to what the caller did by hand, so
        adding a driver can never silently produce a WRONG row — only a slower
        one.

        Unknown ids are returned with an empty record rather than omitted, so a
        caller can index the result without a membership test.
        """
        tags_by_session: dict[str, list[str]] = {}
        for tag, sid in self.list_tags().items():
            tags_by_session.setdefault(sid, []).append(tag)
        return {
            session_id: SessionRecord(
                title=self.load_title(session_id),
                archived=self.is_archived(session_id),
                terminated_at=self.get_terminated_at(session_id),
                owner=self.get_owner(session_id),
                tags=tags_by_session.get(session_id, []),
                pinned_at=self.get_pinned_at(session_id),
                projects=self.projects_for_session(session_id),
            )
            for session_id in session_ids
        }

    @abc.abstractmethod
    def set_pinned(self, session_id: str, pinned: bool) -> None:
        """Pin or unpin a session, stamping ``pinned_at`` on the record.

        Stored as a TIMESTAMP rather than a bool, matching ``archived_at`` and
        ``terminated_at``, because it doubles as the ordering key: a surface
        shows most-recently-pinned first without a second field. Unlike those
        two the stamp is REVERSIBLE by design — a pin is a user's assertion
        about their own list, not a lifecycle terminal.
        """

    @abc.abstractmethod
    def get_pinned_at(self, session_id: str) -> str | None:
        """Return the stored ``pinned_at`` stamp, or ``None`` if unpinned.

        The read-through sibling of :meth:`set_pinned`, mirroring how
        ``is_archived`` reads ``archive_session``'s write.
        """

    @abc.abstractmethod
    def record_project(self, session_id: str, project: str) -> None:
        """Add *project* to the set this session has worked in (idempotent).

        A SET, not a field, because a session can move: an auto-select session
        starts with no project and the agent may switch several times, and a
        filter has to find it under every one of them. Recording is additive and
        order-free, so a replayed context event is a no-op.
        """

    @abc.abstractmethod
    def projects_for_session(self, session_id: str) -> list[str]:
        """Return every project identity recorded for a session."""

    @abc.abstractmethod
    def query_sessions(self, query: SessionQuery) -> list[str]:
        """Return the session IDs matching *query*, narrowed at the STORE.

        The point of narrowing here rather than at the caller is that the caller
        builds each row by loading that session's whole transcript — so a
        session rejected by the query is one that is never opened. Every field of
        :class:`SessionQuery` is answerable from the session record alone,
        precisely so this can be a single indexed read.

        ``query.owner=None`` lists EVERYTHING — what a caller holding a
        read-all authority (or no identity at all) gets.
        A non-``None`` owner narrows to that subject's own sessions PLUS every
        UNOWNED one. **Unowned is not a hole in the filter, it is the migration
        semantic:** sessions predating the owner stamp carry no subject, and no
        subject can be reconstructed for them after the fact. Hiding them would
        make a user's existing work vanish from their own list, which is a worse
        failure than showing a pre-existing session to someone who could already
        list it before the stamp existed. The unowned set is closed and shrinking
        — every session created by an authenticated caller from here on is
        stamped — so this is a fading allowance, not a permanent widening.
        """

    def list_sessions(self, owner: str | None = None) -> list[str]:
        """List session IDs, optionally narrowed to what *owner* may see.

        Concrete delegate over :meth:`query_sessions`, kept because it is the
        established call shape. ``include_archived=True`` is the contract: this
        method returns archived sessions too, and the archived filter is applied
        by the caller.
        """
        return self.query_sessions(SessionQuery(owner=owner, include_archived=True))

    @abc.abstractmethod
    def archive_session(self, session_id: str) -> None:
        """Mark a session as archived."""

    @abc.abstractmethod
    def unarchive_session(self, session_id: str) -> None:
        """Remove archived status from a session."""

    @abc.abstractmethod
    def is_archived(self, session_id: str) -> bool:
        """Return True if a session is archived."""

    @abc.abstractmethod
    def terminate_session(self, session_id: str) -> bool:
        """Permanently mark a session terminated (irreversible).

        Stamps ``terminated_at`` **once** — a repeat call never moves the
        original timestamp. There is deliberately NO un-terminate primitive:
        termination is a one-way door. Mirrors
        ``archive_session`` but without the reverse operation.

        Returns ``True`` iff THIS call was the one that newly stamped the
        timestamp, ``False`` if the session was already terminated. This is
        the arbitration signal ``SessionRuntime.terminate_session`` reads to
        run its side effects (cancel + callbacks + event) exactly once even
        under concurrent callers — the store's set-once write is the only
        thing racing safely, so the runtime must never decide on its own
        unlocked read.
        """

    @abc.abstractmethod
    def get_terminated_at(self, session_id: str) -> str | None:
        """Return the ISO ``terminated_at`` timestamp, or ``None`` if live.

        The read-through sibling of :meth:`terminate_session` (mirrors how
        ``is_archived`` reads ``archive_session``'s write).
        """

    def is_terminated(self, session_id: str) -> bool:
        """Return True iff a session was permanently terminated.

        Concrete over :meth:`get_terminated_at` so both backends share one
        implementation — the single derivation every termination guard reads.
        """
        return self.get_terminated_at(session_id) is not None

    @abc.abstractmethod
    def truncate_after(self, session_id: str, cutoff_ts: str) -> int:
        """Delete all events with ``ts > cutoff_ts``.

        Returns the number of deleted events. Used by the recovery
        pipeline to clean up a failed run before re-driving.
        """

    # -- concrete template methods ------------------------------------------

    @staticmethod
    def stamp(event: Event) -> EventRecord:
        """Build the durable record for *event*, stamping or canonicalising ``ts``.

        The ONE place a stored timestamp is spelled, called by every backend's
        ``_write_event``. An event that carries no ``ts`` — everything the engine
        appends — gets the clock's. One that DOES is normalised to the same UTC
        spelling rather than stored verbatim: the mirror-ingest endpoint accepts
        a client's own timestamps, and a differing UTC offset would order
        differently as TEXT than as an instant. Every ``ts`` comparison in this
        package is textual (``load_transcript``'s sort, ``truncate_after``'s
        range, the cursor floor a store pushes down), so a single foreign
        spelling would sort a real event into the wrong place — silently, since
        text comparison never raises. The INSTANT is preserved; only its
        spelling is fixed.

        A value that cannot be parsed is kept exactly as sent. It is still
        evidence of what a client claimed, and replacing it with the clock would
        forge a timestamp for an event that already has one.
        """
        raw = event.get("ts")
        if raw is None:
            return {"ts": _utc_now(), **event}
        return {**event, "ts": EventCursor.canonical(raw) or raw}

    @staticmethod
    def _publish_appended(session_id: str, record: EventRecord) -> None:
        """Fan a just-appended record out to the process-wide event bus.

        The universal append choke-point: every backend calls this right after
        the durable write so SSE waiters wake immediately and ``on_event`` hooks
        fire — driven by the SAME persisted ``record`` ``load_transcript``
        returns, so a live SSE event is byte-identical to the backlog one.
        Best-effort: a bus failure must never break a durable append.
        """
        from mewbo_core.session.session_event_bus import get_session_event_bus

        try:
            get_session_event_bus().publish(session_id, record)
        except Exception:
            logging.warning("Session event bus publish failed.", exc_info=True)

    @staticmethod
    def merge_context_events(events: list[EventRecord]) -> dict[str, object]:
        """Reduce a transcript to its current context (most-recent payload wins).

        One reducer shared by :meth:`latest_context` (the trace-provenance path)
        and ``SessionRuntime.summarize_session`` (the origin/recovery path) so the
        rule for "what context is this session running under" lives in exactly one
        place. ``context`` events are sparse, so a full scan is cheap.
        """
        merged: dict[str, object] = {}
        for event in events:
            if event.get("type") == "context":
                payload = event.get("payload")
                if isinstance(payload, dict):
                    merged.update(payload)
        return merged

    def latest_context(self, session_id: str) -> dict[str, object]:
        """Return the session's merged context (most-recent context event wins)."""
        return self.merge_context_events(self.load_transcript(session_id))

    def last_attestation_hash(self, session_id: str) -> str:
        """Return the attestation chain's head to re-seed on recovery.

        Concrete default over :meth:`load_transcript`: scans for the most
        recent ``type == "attestation"`` event and returns its persisted
        ``record_hash``, else the chain's genesis hash — mirrors the
        ``tags_for_session`` concrete-default idiom so every backend shares
        one scan; ``MongoSessionStore`` overrides with a targeted query.
        """
        from mewbo_core.agents.attestation import GENESIS_HASH

        for event in reversed(self.load_transcript(session_id)):
            if event.get("type") != "attestation":
                continue
            payload = event.get("payload")
            if isinstance(payload, dict):
                record_hash = payload.get("record_hash")
                if isinstance(record_hash, str) and record_hash:
                    return record_hash
        return GENESIS_HASH

    def fork_session(self, source_session_id: str, owner: str | None = None) -> str:
        """Create a new session by copying events, summary, and title from another.

        The fork is stamped for whoever asked for it, NOT for the source's owner:
        a fork is a new session that happens to start with a copy of a transcript.
        Leaving it unstamped would be worse than either — an unowned fork of an
        owned session is visible to every lister, so forking would launder a
        session out of its owner's scope.
        """
        events = self.load_transcript(source_session_id)
        summary = self.load_summary(source_session_id)
        title = self.load_title(source_session_id)
        new_session_id = self.create_session(owner)
        for event in events:
            self.append_event(new_session_id, event)
        if summary:
            self.save_summary(new_session_id, summary)
        if title:
            self.save_title(new_session_id, title)
        return new_session_id

    def fork_session_at(
        self, source_session_id: str, cutoff_ts: str, owner: str | None = None
    ) -> str:
        """Fork a session, keeping only events with ``ts <= cutoff_ts``.

        Composes :meth:`fork_session` + :meth:`truncate_after` and clears the
        copied summary (which may reference events beyond the cutoff).
        """
        new_id = self.fork_session(source_session_id, owner)
        self.truncate_after(new_id, cutoff_ts)
        self.save_summary(new_id, "")
        return new_id

    def list_session_digests(
        self,
        query: SessionQuery | None = None,
        *,
        limit: int | None = None,
        offset: int = 0,
    ) -> list[SessionDigest]:
        """Return one listing digest per session matching *query*.

        The seam a LISTING reads instead of calling :meth:`load_transcript` per
        id. One digest per session, in the order :meth:`query_sessions` returns
        them when *limit* is ``None``; a session with no events yields an empty
        digest rather than being dropped, because whether it belongs on a
        listing is the caller's rule (a running session has no events until its
        first append). ``None`` means an unnarrowed call — every session,
        archived included — leaving a caller that filters for itself to do so.

        **Cost is per DRIVER, and the base template is the expensive one.**
        This default folds every transcript, so it is ``O(all history)`` —
        honest, correct, and the reason ``MongoSessionStore`` overrides it with
        a projection that is ``O(collection)`` in rows plus ``O(summary-relevant
        events)`` in reads. A backend that gains an index owes an override; a
        backend that cannot must say so here rather than let a caller assume the
        listing is cheap.

        What the default still buys: the fold STREAMS, so peak memory is the
        RELEVANT events rather than the whole transcript — a 10,266-event
        session never has to be resident to contribute one row.

        *limit*/*offset* page the RESULT of the fold this driver already pays
        for — the file backend has no cheap way to learn a session's first
        timestamp without opening it, so paging here narrows what is RETURNED,
        not what is READ. ``MongoSessionStore`` overrides this to narrow the
        read too; see its docstring for why that split is real. Sort key is
        each digest's first event's ``ts`` (== the eventual ``created_at`` for
        every row that survives the caller's own visibility filter — see
        ``SessionRuntime.list_sessions``), descending, so a page here orders the
        same way the final listing does.

        **The page is cut from what *query* already admitted.** Even on a driver
        where paging cannot narrow the read, the two must compose in that order:
        filtering a page instead of paging the filtered set would make
        ``pinned=True`` with a ``limit`` return only the pinned sessions inside
        the newest N candidates, which is usually none of them.
        """
        digests = [
            SessionDigest.from_events(session_id, self.stream_transcript(session_id))
            for session_id in self.query_sessions(
                query or SessionQuery(include_archived=True)
            )
        ]
        if limit is None:
            return digests
        digests.sort(
            key=lambda digest: str(digest.events[0].get("ts")) if digest.events else "",
            reverse=True,
        )
        offset = max(offset, 0)
        return digests[offset : offset + max(limit, 0)]

    def session_digest(self, session_id: str) -> SessionDigest:
        """Return ONE session's listing digest — the single-session sibling above.

        ``O(one session's summary-relevant events)`` for a driver that projects,
        ``O(one session's events)`` for one that folds. Exists because a
        summary is not only a listing concern: the poll path re-derives status
        for a single session on every tick, and doing that from the full
        transcript is the same defect at a different scale — 0.211 s per poll on
        the largest live session, once a second, per open client.

        Deliberately NOT expressed as ``list_session_digests`` narrowed to one
        id: that route pays the listing's own setup (enumerating sessions) to
        answer a question about one.
        """
        return SessionDigest.from_events(session_id, self.stream_transcript(session_id))

    def load_events_after(
        self, session_id: str, cursor: EventCursor | None
    ) -> list[EventRecord]:
        """Return the events strictly newer than *cursor* (all of them if ``None``).

        ``O(one session's events)`` here, and that is the point of the seam
        rather than an acceptance of it: the base has no index to narrow with,
        so it streams and tests, but a driver that CAN turn the cursor into a
        range read overrides this and becomes ``O(matched events)``. Filtering a
        fully-materialised transcript in Python is a cursor in name only — the
        response shrinks and the work does not.

        A ``None`` cursor means "everything". It never means "the cursor was
        bad": an unparseable value is refused at the boundary that received it,
        because a filter that cannot be applied must not widen to the whole
        collection.
        """
        events = self.stream_transcript(session_id)
        if cursor is None:
            return list(events)
        return [event for event in events if cursor.matches(event)]

    def stream_transcript(self, session_id: str) -> Iterator[EventRecord]:
        """Yield a session's events in ts order without materialising them.

        ``O(one session's events)`` in time, ``O(1)`` in memory for a driver that
        overrides it. The default just iterates :meth:`load_transcript`, so it
        buys nothing on its own — it exists so :meth:`list_session_digests` has
        one seam a line-oriented backend can make genuinely streaming.
        """
        yield from self.load_transcript(session_id)

    def load_recent_events(
        self,
        session_id: str,
        limit: int = 8,
        include_types: set[str] | None = None,
    ) -> list[EventRecord]:
        """Load the most recent events, optionally filtered by type."""
        events = self.load_transcript(session_id)
        if include_types:
            events = [event for event in events if event.get("type") in include_types]
        if limit <= 0:
            return []
        return events[-limit:]

    @staticmethod
    def payload_key_is_set(event: EventRecord, key: str) -> bool:
        """Is *key* present on this event's payload with a value worth reading?

        The ONE spelling of the narrowing :meth:`latest_event_of_type` applies
        for *payload_key*, published here so ``MongoSessionStore`` can push an
        equivalent query down rather than restate the rule. Present, not
        ``None``, not the empty string — deliberately no further judgement: what
        counts as a USABLE value belongs to the caller that knows what the key
        means, and a store that guessed would silently skip an event a caller
        would have accepted.
        """
        payload = event.get("payload")
        if not isinstance(payload, dict) or key not in payload:
            return False
        value = payload[key]
        return value is not None and value != ""

    def latest_event_of_type(
        self, session_id: str, event_type: str, *, payload_key: str | None = None
    ) -> EventRecord | None:
        """Return the NEWEST event of *event_type*, or ``None`` when there is none.

        ``O(one session's events)`` here and ``O(1)`` on a driver that can walk
        an index backwards (``MongoSessionStore`` overrides). Same honest split
        as :meth:`load_events_after` and :meth:`load_recent_events`: the base
        has no cheap reverse read, so it streams and keeps the last match — a
        caller must not read this primitive as cheap on every backend. Memory
        stays ``O(1)``: it holds one event, never the transcript.

        **Bounded by the TYPE, never by a count.** Tailing the last N events is
        the tempting cheap answer and it is a wrong one — the newest event of a
        sparse type sits arbitrarily far back in a session with a long run since
        the last one, so a window that misses it reports "no such event". That
        is a wrong ANSWER, not a slow one, and a caller turns it into a refusal.

        *payload_key* narrows further, to events whose payload satisfies
        :meth:`payload_key_is_set`. It exists because "the newest event of this
        type" and "the newest event of this type that CARRIES the field I came
        for" are different questions, and a caller answering the second with the
        first is wrong whenever a later event of the same type omits the field.
        Context events are exactly that shape: they are merged key-by-key
        (:meth:`merge_context_events`), so a session can write a ``project``
        and then several context events that say nothing about one. Measured on
        the deployed store, 15 of the 204 sessions carrying a ``project``
        anywhere in context had a NEWER context event without it.

        Returns the EVENT, not a field off it, so a second caller wanting a
        different key out of the same event is served by the same read.
        """
        latest: EventRecord | None = None
        for event in self.stream_transcript(session_id):
            if event.get("type") != event_type:
                continue
            if payload_key is not None and not self.payload_key_is_set(event, payload_key):
                continue
            latest = event
        return latest

    async def compact_session(
        self,
        session_id: str,
        mode: CompactionMode | None = None,
        **kwargs: Any,
    ) -> CompactionResult:
        """Compact a session's transcript using structured summarization."""
        from mewbo_core.session.compact import (
            CompactionMode as CM,
            CompactionResult as CR,
            compact_conversation,
        )

        resolved_mode: CM = CM(mode) if mode is not None else CM.PARTIAL
        events = self.load_transcript(session_id)
        result: CR = await compact_conversation(events, resolved_mode, **kwargs)
        self.save_summary(session_id, result.summary)
        return result

__init__() -> None

Initialise the in-memory termination cache shared by every backend.

Populated by :meth:_mark_terminated_cached (called from each backend's terminate_session) and consulted by :meth:_guard_append, so a hot-path append never pays a Mongo round-trip / index-file read to learn whether its own session was just terminated. Best-effort and per-process only: a session terminated by a DIFFERENT process/worker is still caught by the durable terminated_at stamp at every other termination-aware seam (resolve_session, recovery) — this cache only shortcuts the common same-process race between a live run and a concurrent /terminate call.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
def __init__(self) -> None:
    """Initialise the in-memory termination cache shared by every backend.

    Populated by :meth:`_mark_terminated_cached` (called from each
    backend's ``terminate_session``) and consulted by
    :meth:`_guard_append`, so a hot-path append never pays a Mongo
    round-trip / index-file read to learn whether its own session was
    just terminated. Best-effort and per-process only: a session
    terminated by a DIFFERENT process/worker is still caught by the
    durable ``terminated_at`` stamp at every other termination-aware seam
    (``resolve_session``, recovery) — this cache only shortcuts the
    common same-process race between a live run and a concurrent
    ``/terminate`` call.
    """
    self._terminated_cache: set[str] = set()
    self._terminated_logged: set[str] = set()

append_event(session_id: str, event: Event) -> None

Append a single event record to the session transcript.

No-ops (after a dropped-event log, once per session) once :meth:terminate_session has stamped this session terminated — see :meth:_guard_append. The ONE deliberate exception is :meth:append_terminal_event, used for the termination tombstone itself.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
161
162
163
164
165
166
167
168
169
170
171
172
173
def append_event(self, session_id: str, event: Event) -> None:
    """Append a single event record to the session transcript.

    No-ops (after a dropped-event log, once per session) once
    :meth:`terminate_session` has stamped this session terminated — see
    :meth:`_guard_append`. The ONE deliberate exception is
    :meth:`append_terminal_event`, used for the termination tombstone
    itself.
    """
    if not self._guard_append(session_id):
        return
    self._write_event(session_id, event)
    self._record_context_project(session_id, event)

append_terminal_event(session_id: str, event: Event) -> None

Append an event exempt from the termination guard.

SessionRuntime.terminate_session writes the session_terminated tombstone through this seam immediately after the store's own terminate_session has already cached the session as terminated — an ordinary :meth:append_event call would otherwise drop its own closing event. No other caller should use this.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
233
234
235
236
237
238
239
240
241
242
def append_terminal_event(self, session_id: str, event: Event) -> None:
    """Append an event exempt from the termination guard.

    ``SessionRuntime.terminate_session`` writes the ``session_terminated``
    tombstone through this seam immediately after the store's own
    ``terminate_session`` has already cached the session as terminated —
    an ordinary :meth:`append_event` call would otherwise drop its own
    closing event. No other caller should use this.
    """
    self._write_event(session_id, event)

append_user_turn(session_id: str, text: str, attachments: list[dict] | None = None) -> None

Append the user event that records ONE accepted turn.

The single builder of that payload, because two seams write it: the acceptance seam (SessionRuntime.start_async, so the turn is durable before the executor's cold start) and the orchestration body itself (for every caller that reaches Orchestrator directly). A second hand-rolled {"type": "user", ...} literal in either place is free to drift from this one — silently, since both would still write a transcript event the readers accept.

attachments is written only when non-empty, mirroring :class:~mewbo_core.contracts.types.UserPayload's NotRequired key, so a turn without attachments carries no such key. It additively duplicates the sibling context event's descriptors so a client can render attachment cards above the user turn without joining across events.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
def append_user_turn(
    self,
    session_id: str,
    text: str,
    attachments: list[dict] | None = None,
) -> None:
    """Append the ``user`` event that records ONE accepted turn.

    The single builder of that payload, because two seams write it: the
    acceptance seam (``SessionRuntime.start_async``, so the turn is durable
    before the executor's cold start) and the orchestration body itself (for
    every caller that reaches ``Orchestrator`` directly). A second
    hand-rolled ``{"type": "user", ...}`` literal in either place is free to
    drift from this one — silently, since both would still write a
    transcript event the readers accept.

    ``attachments`` is written only when non-empty, mirroring
    :class:`~mewbo_core.contracts.types.UserPayload`'s ``NotRequired`` key,
    so a turn without attachments carries no such key. It additively
    duplicates the sibling ``context`` event's
    descriptors so a client can render attachment cards above the user turn
    without joining across events.
    """
    payload: UserPayload = {"text": text}
    if attachments:
        payload["attachments"] = attachments
    self.append_event(session_id, {"type": "user", "payload": payload})

archive_session(session_id: str) -> None abstractmethod

Mark a session as archived.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
407
408
409
@abc.abstractmethod
def archive_session(self, session_id: str) -> None:
    """Mark a session as archived."""

compact_session(session_id: str, mode: CompactionMode | None = None, **kwargs: Any) -> CompactionResult async

Compact a session's transcript using structured summarization.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
async def compact_session(
    self,
    session_id: str,
    mode: CompactionMode | None = None,
    **kwargs: Any,
) -> CompactionResult:
    """Compact a session's transcript using structured summarization."""
    from mewbo_core.session.compact import (
        CompactionMode as CM,
        CompactionResult as CR,
        compact_conversation,
    )

    resolved_mode: CM = CM(mode) if mode is not None else CM.PARTIAL
    events = self.load_transcript(session_id)
    result: CR = await compact_conversation(events, resolved_mode, **kwargs)
    self.save_summary(session_id, result.summary)
    return result

create_session(owner: str | None = None) -> str abstractmethod

Create a new session and return its identifier.

owner is an OPAQUE subject string, never an identity object: core sits below the identity kernel in the dependency DAG and must not learn what a principal is. The caller that HAS one resolves it to a string first. None means unowned, which is what every session created without an authenticated caller is — see :meth:list_sessions for what that implies on the read side.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
118
119
120
121
122
123
124
125
126
127
128
@abc.abstractmethod
def create_session(self, owner: str | None = None) -> str:
    """Create a new session and return its identifier.

    *owner* is an OPAQUE subject string, never an identity object: core sits
    below the identity kernel in the dependency DAG and must not learn what a
    principal is. The caller that HAS one resolves it to a string first.
    ``None`` means unowned, which is what every session created without an
    authenticated caller is — see :meth:`list_sessions` for what that implies
    on the read side.
    """

ensure_session(session_id: str, owner: str | None = None) -> None abstractmethod

Idempotently materialise a session RECORD for a known id.

create_session mints its own uuid; some callers (the realtime write-behind recorder) need to back a session whose id was pre-minted elsewhere — e.g. for an in-flight Langfuse trace opened before any store write. Without a record, list_sessions (and every read surface built on it) never sees the id, so the transcript is an orphan: events exist, the session is invisible.

This is the seam that closes that gap. It is idempotent — calling it on an existing session is a no-op (never resets created_at / archived state). create_session is implemented on top of it (mint id → materialise).

The owner stamp is SET-ONCE at materialisation, like created_at and unlike archived_at: a replay must not be able to re-point an existing session at a different subject, which would be an ownership takeover written through an idempotent no-op path.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
@abc.abstractmethod
def ensure_session(self, session_id: str, owner: str | None = None) -> None:
    """Idempotently materialise a session RECORD for a known id.

    ``create_session`` mints its own uuid; some callers (the realtime
    write-behind recorder) need to back a session whose id was pre-minted
    elsewhere — e.g. for an in-flight Langfuse trace opened before any store
    write. Without a record, ``list_sessions`` (and every read surface built
    on it) never sees the id, so the transcript is an orphan: events exist,
    the session is invisible.

    This is the seam that closes that gap. It is idempotent — calling it on
    an existing session is a no-op (never resets created_at / archived state).
    ``create_session`` is implemented on top of it (mint id → materialise).

    The owner stamp is SET-ONCE at materialisation, like ``created_at`` and
    unlike ``archived_at``: a replay must not be able to re-point an existing
    session at a different subject, which would be an ownership takeover
    written through an idempotent no-op path.
    """

fork_session(source_session_id: str, owner: str | None = None) -> str

Create a new session by copying events, summary, and title from another.

The fork is stamped for whoever asked for it, NOT for the source's owner: a fork is a new session that happens to start with a copy of a transcript. Leaving it unstamped would be worse than either — an unowned fork of an owned session is visible to every lister, so forking would launder a session out of its owner's scope.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
def fork_session(self, source_session_id: str, owner: str | None = None) -> str:
    """Create a new session by copying events, summary, and title from another.

    The fork is stamped for whoever asked for it, NOT for the source's owner:
    a fork is a new session that happens to start with a copy of a transcript.
    Leaving it unstamped would be worse than either — an unowned fork of an
    owned session is visible to every lister, so forking would launder a
    session out of its owner's scope.
    """
    events = self.load_transcript(source_session_id)
    summary = self.load_summary(source_session_id)
    title = self.load_title(source_session_id)
    new_session_id = self.create_session(owner)
    for event in events:
        self.append_event(new_session_id, event)
    if summary:
        self.save_summary(new_session_id, summary)
    if title:
        self.save_title(new_session_id, title)
    return new_session_id

fork_session_at(source_session_id: str, cutoff_ts: str, owner: str | None = None) -> str

Fork a session, keeping only events with ts <= cutoff_ts.

Composes :meth:fork_session + :meth:truncate_after and clears the copied summary (which may reference events beyond the cutoff).

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
568
569
570
571
572
573
574
575
576
577
578
579
def fork_session_at(
    self, source_session_id: str, cutoff_ts: str, owner: str | None = None
) -> str:
    """Fork a session, keeping only events with ``ts <= cutoff_ts``.

    Composes :meth:`fork_session` + :meth:`truncate_after` and clears the
    copied summary (which may reference events beyond the cutoff).
    """
    new_id = self.fork_session(source_session_id, owner)
    self.truncate_after(new_id, cutoff_ts)
    self.save_summary(new_id, "")
    return new_id

get_owner(session_id: str) -> str | None abstractmethod

Return the owning subject, or None if the session is unowned.

The read-through sibling of the stamp ensure_session writes, mirroring how is_archived reads archive_session's write.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
269
270
271
272
273
274
275
@abc.abstractmethod
def get_owner(self, session_id: str) -> str | None:
    """Return the owning subject, or ``None`` if the session is unowned.

    The read-through sibling of the stamp ``ensure_session`` writes, mirroring
    how ``is_archived`` reads ``archive_session``'s write.
    """

get_pinned_at(session_id: str) -> str | None abstractmethod

Return the stored pinned_at stamp, or None if unpinned.

The read-through sibling of :meth:set_pinned, mirroring how is_archived reads archive_session's write.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
352
353
354
355
356
357
358
@abc.abstractmethod
def get_pinned_at(self, session_id: str) -> str | None:
    """Return the stored ``pinned_at`` stamp, or ``None`` if unpinned.

    The read-through sibling of :meth:`set_pinned`, mirroring how
    ``is_archived`` reads ``archive_session``'s write.
    """

get_terminated_at(session_id: str) -> str | None abstractmethod

Return the ISO terminated_at timestamp, or None if live.

The read-through sibling of :meth:terminate_session (mirrors how is_archived reads archive_session's write).

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
437
438
439
440
441
442
443
@abc.abstractmethod
def get_terminated_at(self, session_id: str) -> str | None:
    """Return the ISO ``terminated_at`` timestamp, or ``None`` if live.

    The read-through sibling of :meth:`terminate_session` (mirrors how
    ``is_archived`` reads ``archive_session``'s write).
    """

is_archived(session_id: str) -> bool abstractmethod

Return True if a session is archived.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
415
416
417
@abc.abstractmethod
def is_archived(self, session_id: str) -> bool:
    """Return True if a session is archived."""

is_terminated(session_id: str) -> bool

Return True iff a session was permanently terminated.

Concrete over :meth:get_terminated_at so both backends share one implementation — the single derivation every termination guard reads.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
445
446
447
448
449
450
451
def is_terminated(self, session_id: str) -> bool:
    """Return True iff a session was permanently terminated.

    Concrete over :meth:`get_terminated_at` so both backends share one
    implementation — the single derivation every termination guard reads.
    """
    return self.get_terminated_at(session_id) is not None

last_attestation_hash(session_id: str) -> str

Return the attestation chain's head to re-seed on recovery.

Concrete default over :meth:load_transcript: scans for the most recent type == "attestation" event and returns its persisted record_hash, else the chain's genesis hash — mirrors the tags_for_session concrete-default idiom so every backend shares one scan; MongoSessionStore overrides with a targeted query.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
def last_attestation_hash(self, session_id: str) -> str:
    """Return the attestation chain's head to re-seed on recovery.

    Concrete default over :meth:`load_transcript`: scans for the most
    recent ``type == "attestation"`` event and returns its persisted
    ``record_hash``, else the chain's genesis hash — mirrors the
    ``tags_for_session`` concrete-default idiom so every backend shares
    one scan; ``MongoSessionStore`` overrides with a targeted query.
    """
    from mewbo_core.agents.attestation import GENESIS_HASH

    for event in reversed(self.load_transcript(session_id)):
        if event.get("type") != "attestation":
            continue
        payload = event.get("payload")
        if isinstance(payload, dict):
            record_hash = payload.get("record_hash")
            if isinstance(record_hash, str) and record_hash:
                return record_hash
    return GENESIS_HASH

latest_context(session_id: str) -> dict[str, object]

Return the session's merged context (most-recent context event wins).

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
522
523
524
def latest_context(self, session_id: str) -> dict[str, object]:
    """Return the session's merged context (most-recent context event wins)."""
    return self.merge_context_events(self.load_transcript(session_id))

latest_event_of_type(session_id: str, event_type: str, *, payload_key: str | None = None) -> EventRecord | None

Return the NEWEST event of event_type, or None when there is none.

O(one session's events) here and O(1) on a driver that can walk an index backwards (MongoSessionStore overrides). Same honest split as :meth:load_events_after and :meth:load_recent_events: the base has no cheap reverse read, so it streams and keeps the last match — a caller must not read this primitive as cheap on every backend. Memory stays O(1): it holds one event, never the transcript.

Bounded by the TYPE, never by a count. Tailing the last N events is the tempting cheap answer and it is a wrong one — the newest event of a sparse type sits arbitrarily far back in a session with a long run since the last one, so a window that misses it reports "no such event". That is a wrong ANSWER, not a slow one, and a caller turns it into a refusal.

payload_key narrows further, to events whose payload satisfies :meth:payload_key_is_set. It exists because "the newest event of this type" and "the newest event of this type that CARRIES the field I came for" are different questions, and a caller answering the second with the first is wrong whenever a later event of the same type omits the field. Context events are exactly that shape: they are merged key-by-key (:meth:merge_context_events), so a session can write a project and then several context events that say nothing about one. Measured on the deployed store, 15 of the 204 sessions carrying a project anywhere in context had a NEWER context event without it.

Returns the EVENT, not a field off it, so a second caller wanting a different key out of the same event is served by the same read.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
def latest_event_of_type(
    self, session_id: str, event_type: str, *, payload_key: str | None = None
) -> EventRecord | None:
    """Return the NEWEST event of *event_type*, or ``None`` when there is none.

    ``O(one session's events)`` here and ``O(1)`` on a driver that can walk
    an index backwards (``MongoSessionStore`` overrides). Same honest split
    as :meth:`load_events_after` and :meth:`load_recent_events`: the base
    has no cheap reverse read, so it streams and keeps the last match — a
    caller must not read this primitive as cheap on every backend. Memory
    stays ``O(1)``: it holds one event, never the transcript.

    **Bounded by the TYPE, never by a count.** Tailing the last N events is
    the tempting cheap answer and it is a wrong one — the newest event of a
    sparse type sits arbitrarily far back in a session with a long run since
    the last one, so a window that misses it reports "no such event". That
    is a wrong ANSWER, not a slow one, and a caller turns it into a refusal.

    *payload_key* narrows further, to events whose payload satisfies
    :meth:`payload_key_is_set`. It exists because "the newest event of this
    type" and "the newest event of this type that CARRIES the field I came
    for" are different questions, and a caller answering the second with the
    first is wrong whenever a later event of the same type omits the field.
    Context events are exactly that shape: they are merged key-by-key
    (:meth:`merge_context_events`), so a session can write a ``project``
    and then several context events that say nothing about one. Measured on
    the deployed store, 15 of the 204 sessions carrying a ``project``
    anywhere in context had a NEWER context event without it.

    Returns the EVENT, not a field off it, so a second caller wanting a
    different key out of the same event is served by the same read.
    """
    latest: EventRecord | None = None
    for event in self.stream_transcript(session_id):
        if event.get("type") != event_type:
            continue
        if payload_key is not None and not self.payload_key_is_set(event, payload_key):
            continue
        latest = event
    return latest

list_session_digests(query: SessionQuery | None = None, *, limit: int | None = None, offset: int = 0) -> list[SessionDigest]

Return one listing digest per session matching query.

The seam a LISTING reads instead of calling :meth:load_transcript per id. One digest per session, in the order :meth:query_sessions returns them when limit is None; a session with no events yields an empty digest rather than being dropped, because whether it belongs on a listing is the caller's rule (a running session has no events until its first append). None means an unnarrowed call — every session, archived included — leaving a caller that filters for itself to do so.

Cost is per DRIVER, and the base template is the expensive one. This default folds every transcript, so it is O(all history) — honest, correct, and the reason MongoSessionStore overrides it with a projection that is O(collection) in rows plus O(summary-relevant events) in reads. A backend that gains an index owes an override; a backend that cannot must say so here rather than let a caller assume the listing is cheap.

What the default still buys: the fold STREAMS, so peak memory is the RELEVANT events rather than the whole transcript — a 10,266-event session never has to be resident to contribute one row.

limit/offset page the RESULT of the fold this driver already pays for — the file backend has no cheap way to learn a session's first timestamp without opening it, so paging here narrows what is RETURNED, not what is READ. MongoSessionStore overrides this to narrow the read too; see its docstring for why that split is real. Sort key is each digest's first event's ts (== the eventual created_at for every row that survives the caller's own visibility filter — see SessionRuntime.list_sessions), descending, so a page here orders the same way the final listing does.

The page is cut from what query already admitted. Even on a driver where paging cannot narrow the read, the two must compose in that order: filtering a page instead of paging the filtered set would make pinned=True with a limit return only the pinned sessions inside the newest N candidates, which is usually none of them.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
def list_session_digests(
    self,
    query: SessionQuery | None = None,
    *,
    limit: int | None = None,
    offset: int = 0,
) -> list[SessionDigest]:
    """Return one listing digest per session matching *query*.

    The seam a LISTING reads instead of calling :meth:`load_transcript` per
    id. One digest per session, in the order :meth:`query_sessions` returns
    them when *limit* is ``None``; a session with no events yields an empty
    digest rather than being dropped, because whether it belongs on a
    listing is the caller's rule (a running session has no events until its
    first append). ``None`` means an unnarrowed call — every session,
    archived included — leaving a caller that filters for itself to do so.

    **Cost is per DRIVER, and the base template is the expensive one.**
    This default folds every transcript, so it is ``O(all history)`` —
    honest, correct, and the reason ``MongoSessionStore`` overrides it with
    a projection that is ``O(collection)`` in rows plus ``O(summary-relevant
    events)`` in reads. A backend that gains an index owes an override; a
    backend that cannot must say so here rather than let a caller assume the
    listing is cheap.

    What the default still buys: the fold STREAMS, so peak memory is the
    RELEVANT events rather than the whole transcript — a 10,266-event
    session never has to be resident to contribute one row.

    *limit*/*offset* page the RESULT of the fold this driver already pays
    for — the file backend has no cheap way to learn a session's first
    timestamp without opening it, so paging here narrows what is RETURNED,
    not what is READ. ``MongoSessionStore`` overrides this to narrow the
    read too; see its docstring for why that split is real. Sort key is
    each digest's first event's ``ts`` (== the eventual ``created_at`` for
    every row that survives the caller's own visibility filter — see
    ``SessionRuntime.list_sessions``), descending, so a page here orders the
    same way the final listing does.

    **The page is cut from what *query* already admitted.** Even on a driver
    where paging cannot narrow the read, the two must compose in that order:
    filtering a page instead of paging the filtered set would make
    ``pinned=True`` with a ``limit`` return only the pinned sessions inside
    the newest N candidates, which is usually none of them.
    """
    digests = [
        SessionDigest.from_events(session_id, self.stream_transcript(session_id))
        for session_id in self.query_sessions(
            query or SessionQuery(include_archived=True)
        )
    ]
    if limit is None:
        return digests
    digests.sort(
        key=lambda digest: str(digest.events[0].get("ts")) if digest.events else "",
        reverse=True,
    )
    offset = max(offset, 0)
    return digests[offset : offset + max(limit, 0)]

list_sessions(owner: str | None = None) -> list[str]

List session IDs, optionally narrowed to what owner may see.

Concrete delegate over :meth:query_sessions, kept because it is the established call shape. include_archived=True is the contract: this method returns archived sessions too, and the archived filter is applied by the caller.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
397
398
399
400
401
402
403
404
405
def list_sessions(self, owner: str | None = None) -> list[str]:
    """List session IDs, optionally narrowed to what *owner* may see.

    Concrete delegate over :meth:`query_sessions`, kept because it is the
    established call shape. ``include_archived=True`` is the contract: this
    method returns archived sessions too, and the archived filter is applied
    by the caller.
    """
    return self.query_sessions(SessionQuery(owner=owner, include_archived=True))

list_tags() -> dict[str, str] abstractmethod

Return a mapping of tags to session IDs.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
289
290
291
@abc.abstractmethod
def list_tags(self) -> dict[str, str]:
    """Return a mapping of tags to session IDs."""

load_events_after(session_id: str, cursor: EventCursor | None) -> list[EventRecord]

Return the events strictly newer than cursor (all of them if None).

O(one session's events) here, and that is the point of the seam rather than an acceptance of it: the base has no index to narrow with, so it streams and tests, but a driver that CAN turn the cursor into a range read overrides this and becomes O(matched events). Filtering a fully-materialised transcript in Python is a cursor in name only — the response shrinks and the work does not.

A None cursor means "everything". It never means "the cursor was bad": an unparseable value is refused at the boundary that received it, because a filter that cannot be applied must not widen to the whole collection.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
def load_events_after(
    self, session_id: str, cursor: EventCursor | None
) -> list[EventRecord]:
    """Return the events strictly newer than *cursor* (all of them if ``None``).

    ``O(one session's events)`` here, and that is the point of the seam
    rather than an acceptance of it: the base has no index to narrow with,
    so it streams and tests, but a driver that CAN turn the cursor into a
    range read overrides this and becomes ``O(matched events)``. Filtering a
    fully-materialised transcript in Python is a cursor in name only — the
    response shrinks and the work does not.

    A ``None`` cursor means "everything". It never means "the cursor was
    bad": an unparseable value is refused at the boundary that received it,
    because a filter that cannot be applied must not widen to the whole
    collection.
    """
    events = self.stream_transcript(session_id)
    if cursor is None:
        return list(events)
    return [event for event in events if cursor.matches(event)]

load_recent_events(session_id: str, limit: int = 8, include_types: set[str] | None = None) -> list[EventRecord]

Load the most recent events, optionally filtered by type.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
689
690
691
692
693
694
695
696
697
698
699
700
701
def load_recent_events(
    self,
    session_id: str,
    limit: int = 8,
    include_types: set[str] | None = None,
) -> list[EventRecord]:
    """Load the most recent events, optionally filtered by type."""
    events = self.load_transcript(session_id)
    if include_types:
        events = [event for event in events if event.get("type") in include_types]
    if limit <= 0:
        return []
    return events[-limit:]

load_session_records(session_ids: list[str]) -> dict[str, SessionRecord]

Batch-load the per-session metadata one summary row needs.

A LISTING derives every row's title, archived flag, termination stamp, owner, tags, pin stamp and project set. Fetching them one session at a time is seven store reads per row — on a networked driver, seven ROUND TRIPS per row, so the cost of listing scales with the session count times a per-call latency that has nothing to do with how much data is involved. This is the seam that collapses them: one call for the whole page.

This default is the CORRECT-everywhere implementation, not the fast one. It still reads per session, but it already removes the worst repetition by resolving the reverse tag index ONCE for the whole batch instead of rescanning it per row. A driver that can answer the whole batch in a single query overrides this (see MongoSessionStore); a driver that cannot inherits behaviour identical to what the caller did by hand, so adding a driver can never silently produce a WRONG row — only a slower one.

Unknown ids are returned with an empty record rather than omitted, so a caller can index the result without a membership test.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
def load_session_records(self, session_ids: list[str]) -> dict[str, SessionRecord]:
    """Batch-load the per-session metadata one summary row needs.

    A LISTING derives every row's title, archived flag, termination stamp,
    owner, tags, pin stamp and project set. Fetching them one session at a
    time is seven store reads per row — on a networked driver, seven ROUND
    TRIPS per row, so the cost of listing scales with the session count
    times a per-call latency that has nothing to do with how much data is
    involved. This is the seam that collapses them: one call for the whole
    page.

    This default is the CORRECT-everywhere implementation, not the fast one.
    It still reads per session, but it already removes the worst repetition
    by resolving the reverse tag index ONCE for the whole batch instead of
    rescanning it per row. A driver that can answer the whole batch in a
    single query overrides this (see ``MongoSessionStore``); a driver that
    cannot inherits behaviour identical to what the caller did by hand, so
    adding a driver can never silently produce a WRONG row — only a slower
    one.

    Unknown ids are returned with an empty record rather than omitted, so a
    caller can index the result without a membership test.
    """
    tags_by_session: dict[str, list[str]] = {}
    for tag, sid in self.list_tags().items():
        tags_by_session.setdefault(sid, []).append(tag)
    return {
        session_id: SessionRecord(
            title=self.load_title(session_id),
            archived=self.is_archived(session_id),
            terminated_at=self.get_terminated_at(session_id),
            owner=self.get_owner(session_id),
            tags=tags_by_session.get(session_id, []),
            pinned_at=self.get_pinned_at(session_id),
            projects=self.projects_for_session(session_id),
        )
        for session_id in session_ids
    }

load_summary(session_id: str) -> str | None abstractmethod

Load a previously saved summary, if present.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
257
258
259
@abc.abstractmethod
def load_summary(self, session_id: str) -> str | None:
    """Load a previously saved summary, if present."""

load_title(session_id: str) -> str | None abstractmethod

Load a previously saved title, if present.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
265
266
267
@abc.abstractmethod
def load_title(self, session_id: str) -> str | None:
    """Load a previously saved title, if present."""

load_transcript(session_id: str) -> list[EventRecord] abstractmethod

Load all transcript events for a session. O(one session's events).

A PER-SESSION read. It must never be called in a loop over sessions — that is O(all history) wearing a listing's clothes, and it is exactly the shape :meth:list_session_digests exists to replace.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
244
245
246
247
248
249
250
251
@abc.abstractmethod
def load_transcript(self, session_id: str) -> list[EventRecord]:
    """Load all transcript events for a session. ``O(one session's events)``.

    A PER-SESSION read. It must never be called in a loop over sessions —
    that is ``O(all history)`` wearing a listing's clothes, and it is exactly
    the shape :meth:`list_session_digests` exists to replace.
    """

merge_context_events(events: list[EventRecord]) -> dict[str, object] staticmethod

Reduce a transcript to its current context (most-recent payload wins).

One reducer shared by :meth:latest_context (the trace-provenance path) and SessionRuntime.summarize_session (the origin/recovery path) so the rule for "what context is this session running under" lives in exactly one place. context events are sparse, so a full scan is cheap.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
@staticmethod
def merge_context_events(events: list[EventRecord]) -> dict[str, object]:
    """Reduce a transcript to its current context (most-recent payload wins).

    One reducer shared by :meth:`latest_context` (the trace-provenance path)
    and ``SessionRuntime.summarize_session`` (the origin/recovery path) so the
    rule for "what context is this session running under" lives in exactly one
    place. ``context`` events are sparse, so a full scan is cheap.
    """
    merged: dict[str, object] = {}
    for event in events:
        if event.get("type") == "context":
            payload = event.get("payload")
            if isinstance(payload, dict):
                merged.update(payload)
    return merged

payload_key_is_set(event: EventRecord, key: str) -> bool staticmethod

Is key present on this event's payload with a value worth reading?

The ONE spelling of the narrowing :meth:latest_event_of_type applies for payload_key, published here so MongoSessionStore can push an equivalent query down rather than restate the rule. Present, not None, not the empty string — deliberately no further judgement: what counts as a USABLE value belongs to the caller that knows what the key means, and a store that guessed would silently skip an event a caller would have accepted.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
@staticmethod
def payload_key_is_set(event: EventRecord, key: str) -> bool:
    """Is *key* present on this event's payload with a value worth reading?

    The ONE spelling of the narrowing :meth:`latest_event_of_type` applies
    for *payload_key*, published here so ``MongoSessionStore`` can push an
    equivalent query down rather than restate the rule. Present, not
    ``None``, not the empty string — deliberately no further judgement: what
    counts as a USABLE value belongs to the caller that knows what the key
    means, and a store that guessed would silently skip an event a caller
    would have accepted.
    """
    payload = event.get("payload")
    if not isinstance(payload, dict) or key not in payload:
        return False
    value = payload[key]
    return value is not None and value != ""

projects_for_session(session_id: str) -> list[str] abstractmethod

Return every project identity recorded for a session.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
370
371
372
@abc.abstractmethod
def projects_for_session(self, session_id: str) -> list[str]:
    """Return every project identity recorded for a session."""

query_sessions(query: SessionQuery) -> list[str] abstractmethod

Return the session IDs matching query, narrowed at the STORE.

The point of narrowing here rather than at the caller is that the caller builds each row by loading that session's whole transcript — so a session rejected by the query is one that is never opened. Every field of :class:SessionQuery is answerable from the session record alone, precisely so this can be a single indexed read.

query.owner=None lists EVERYTHING — what a caller holding a read-all authority (or no identity at all) gets. A non-None owner narrows to that subject's own sessions PLUS every UNOWNED one. Unowned is not a hole in the filter, it is the migration semantic: sessions predating the owner stamp carry no subject, and no subject can be reconstructed for them after the fact. Hiding them would make a user's existing work vanish from their own list, which is a worse failure than showing a pre-existing session to someone who could already list it before the stamp existed. The unowned set is closed and shrinking — every session created by an authenticated caller from here on is stamped — so this is a fading allowance, not a permanent widening.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
@abc.abstractmethod
def query_sessions(self, query: SessionQuery) -> list[str]:
    """Return the session IDs matching *query*, narrowed at the STORE.

    The point of narrowing here rather than at the caller is that the caller
    builds each row by loading that session's whole transcript — so a
    session rejected by the query is one that is never opened. Every field of
    :class:`SessionQuery` is answerable from the session record alone,
    precisely so this can be a single indexed read.

    ``query.owner=None`` lists EVERYTHING — what a caller holding a
    read-all authority (or no identity at all) gets.
    A non-``None`` owner narrows to that subject's own sessions PLUS every
    UNOWNED one. **Unowned is not a hole in the filter, it is the migration
    semantic:** sessions predating the owner stamp carry no subject, and no
    subject can be reconstructed for them after the fact. Hiding them would
    make a user's existing work vanish from their own list, which is a worse
    failure than showing a pre-existing session to someone who could already
    list it before the stamp existed. The unowned set is closed and shrinking
    — every session created by an authenticated caller from here on is
    stamped — so this is a fading allowance, not a permanent widening.
    """

record_project(session_id: str, project: str) -> None abstractmethod

Add project to the set this session has worked in (idempotent).

A SET, not a field, because a session can move: an auto-select session starts with no project and the agent may switch several times, and a filter has to find it under every one of them. Recording is additive and order-free, so a replayed context event is a no-op.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
360
361
362
363
364
365
366
367
368
@abc.abstractmethod
def record_project(self, session_id: str, project: str) -> None:
    """Add *project* to the set this session has worked in (idempotent).

    A SET, not a field, because a session can move: an auto-select session
    starts with no project and the agent may switch several times, and a
    filter has to find it under every one of them. Recording is additive and
    order-free, so a replayed context event is a no-op.
    """

resolve_tag(tag: str) -> str | None abstractmethod

Resolve a tag to a session ID, if present.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
285
286
287
@abc.abstractmethod
def resolve_tag(self, tag: str) -> str | None:
    """Resolve a tag to a session ID, if present."""

save_summary(session_id: str, summary: str) -> None abstractmethod

Persist a summary for a session.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
253
254
255
@abc.abstractmethod
def save_summary(self, session_id: str, summary: str) -> None:
    """Persist a summary for a session."""

save_title(session_id: str, title: str) -> None abstractmethod

Persist a display title for a session.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
261
262
263
@abc.abstractmethod
def save_title(self, session_id: str, title: str) -> None:
    """Persist a display title for a session."""

session_digest(session_id: str) -> SessionDigest

Return ONE session's listing digest — the single-session sibling above.

O(one session's summary-relevant events) for a driver that projects, O(one session's events) for one that folds. Exists because a summary is not only a listing concern: the poll path re-derives status for a single session on every tick, and doing that from the full transcript is the same defect at a different scale — 0.211 s per poll on the largest live session, once a second, per open client.

Deliberately NOT expressed as list_session_digests narrowed to one id: that route pays the listing's own setup (enumerating sessions) to answer a question about one.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
def session_digest(self, session_id: str) -> SessionDigest:
    """Return ONE session's listing digest — the single-session sibling above.

    ``O(one session's summary-relevant events)`` for a driver that projects,
    ``O(one session's events)`` for one that folds. Exists because a
    summary is not only a listing concern: the poll path re-derives status
    for a single session on every tick, and doing that from the full
    transcript is the same defect at a different scale — 0.211 s per poll on
    the largest live session, once a second, per open client.

    Deliberately NOT expressed as ``list_session_digests`` narrowed to one
    id: that route pays the listing's own setup (enumerating sessions) to
    answer a question about one.
    """
    return SessionDigest.from_events(session_id, self.stream_transcript(session_id))

session_dir(session_id: str) -> str abstractmethod

Return the local directory path for a session (used for attachments).

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
277
278
279
@abc.abstractmethod
def session_dir(self, session_id: str) -> str:
    """Return the local directory path for a session (used for attachments)."""

set_pinned(session_id: str, pinned: bool) -> None abstractmethod

Pin or unpin a session, stamping pinned_at on the record.

Stored as a TIMESTAMP rather than a bool, matching archived_at and terminated_at, because it doubles as the ordering key: a surface shows most-recently-pinned first without a second field. Unlike those two the stamp is REVERSIBLE by design — a pin is a user's assertion about their own list, not a lifecycle terminal.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
341
342
343
344
345
346
347
348
349
350
@abc.abstractmethod
def set_pinned(self, session_id: str, pinned: bool) -> None:
    """Pin or unpin a session, stamping ``pinned_at`` on the record.

    Stored as a TIMESTAMP rather than a bool, matching ``archived_at`` and
    ``terminated_at``, because it doubles as the ordering key: a surface
    shows most-recently-pinned first without a second field. Unlike those
    two the stamp is REVERSIBLE by design — a pin is a user's assertion
    about their own list, not a lifecycle terminal.
    """

stamp(event: Event) -> EventRecord staticmethod

Build the durable record for event, stamping or canonicalising ts.

The ONE place a stored timestamp is spelled, called by every backend's _write_event. An event that carries no ts — everything the engine appends — gets the clock's. One that DOES is normalised to the same UTC spelling rather than stored verbatim: the mirror-ingest endpoint accepts a client's own timestamps, and a differing UTC offset would order differently as TEXT than as an instant. Every ts comparison in this package is textual (load_transcript's sort, truncate_after's range, the cursor floor a store pushes down), so a single foreign spelling would sort a real event into the wrong place — silently, since text comparison never raises. The INSTANT is preserved; only its spelling is fixed.

A value that cannot be parsed is kept exactly as sent. It is still evidence of what a client claimed, and replacing it with the clock would forge a timestamp for an event that already has one.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
@staticmethod
def stamp(event: Event) -> EventRecord:
    """Build the durable record for *event*, stamping or canonicalising ``ts``.

    The ONE place a stored timestamp is spelled, called by every backend's
    ``_write_event``. An event that carries no ``ts`` — everything the engine
    appends — gets the clock's. One that DOES is normalised to the same UTC
    spelling rather than stored verbatim: the mirror-ingest endpoint accepts
    a client's own timestamps, and a differing UTC offset would order
    differently as TEXT than as an instant. Every ``ts`` comparison in this
    package is textual (``load_transcript``'s sort, ``truncate_after``'s
    range, the cursor floor a store pushes down), so a single foreign
    spelling would sort a real event into the wrong place — silently, since
    text comparison never raises. The INSTANT is preserved; only its
    spelling is fixed.

    A value that cannot be parsed is kept exactly as sent. It is still
    evidence of what a client claimed, and replacing it with the clock would
    forge a timestamp for an event that already has one.
    """
    raw = event.get("ts")
    if raw is None:
        return {"ts": _utc_now(), **event}
    return {**event, "ts": EventCursor.canonical(raw) or raw}

stream_transcript(session_id: str) -> Iterator[EventRecord]

Yield a session's events in ts order without materialising them.

O(one session's events) in time, O(1) in memory for a driver that overrides it. The default just iterates :meth:load_transcript, so it buys nothing on its own — it exists so :meth:list_session_digests has one seam a line-oriented backend can make genuinely streaming.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
679
680
681
682
683
684
685
686
687
def stream_transcript(self, session_id: str) -> Iterator[EventRecord]:
    """Yield a session's events in ts order without materialising them.

    ``O(one session's events)`` in time, ``O(1)`` in memory for a driver that
    overrides it. The default just iterates :meth:`load_transcript`, so it
    buys nothing on its own — it exists so :meth:`list_session_digests` has
    one seam a line-oriented backend can make genuinely streaming.
    """
    yield from self.load_transcript(session_id)

tag_session(session_id: str, tag: str) -> None abstractmethod

Associate a tag with a session ID for quick lookup.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
281
282
283
@abc.abstractmethod
def tag_session(self, session_id: str, tag: str) -> None:
    """Associate a tag with a session ID for quick lookup."""

tags_for_session(session_id: str) -> list[str]

Return every tag pointing at session_id (reverse of resolve_tag).

Concrete default over list_tags so all backends share one implementation; the tag set is small. Provenance classification reads this to recover a session's origin (see session_provenance).

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
293
294
295
296
297
298
299
300
def tags_for_session(self, session_id: str) -> list[str]:
    """Return every tag pointing at ``session_id`` (reverse of ``resolve_tag``).

    Concrete default over ``list_tags`` so all backends share one
    implementation; the tag set is small. Provenance classification reads
    this to recover a session's origin (see ``session_provenance``).
    """
    return [tag for tag, sid in self.list_tags().items() if sid == session_id]

terminate_session(session_id: str) -> bool abstractmethod

Permanently mark a session terminated (irreversible).

Stamps terminated_at once — a repeat call never moves the original timestamp. There is deliberately NO un-terminate primitive: termination is a one-way door. Mirrors archive_session but without the reverse operation.

Returns True iff THIS call was the one that newly stamped the timestamp, False if the session was already terminated. This is the arbitration signal SessionRuntime.terminate_session reads to run its side effects (cancel + callbacks + event) exactly once even under concurrent callers — the store's set-once write is the only thing racing safely, so the runtime must never decide on its own unlocked read.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
@abc.abstractmethod
def terminate_session(self, session_id: str) -> bool:
    """Permanently mark a session terminated (irreversible).

    Stamps ``terminated_at`` **once** — a repeat call never moves the
    original timestamp. There is deliberately NO un-terminate primitive:
    termination is a one-way door. Mirrors
    ``archive_session`` but without the reverse operation.

    Returns ``True`` iff THIS call was the one that newly stamped the
    timestamp, ``False`` if the session was already terminated. This is
    the arbitration signal ``SessionRuntime.terminate_session`` reads to
    run its side effects (cancel + callbacks + event) exactly once even
    under concurrent callers — the store's set-once write is the only
    thing racing safely, so the runtime must never decide on its own
    unlocked read.
    """

truncate_after(session_id: str, cutoff_ts: str) -> int abstractmethod

Delete all events with ts > cutoff_ts.

Returns the number of deleted events. Used by the recovery pipeline to clean up a failed run before re-driving.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
453
454
455
456
457
458
459
@abc.abstractmethod
def truncate_after(self, session_id: str, cutoff_ts: str) -> int:
    """Delete all events with ``ts > cutoff_ts``.

    Returns the number of deleted events. Used by the recovery
    pipeline to clean up a failed run before re-driving.
    """

unarchive_session(session_id: str) -> None abstractmethod

Remove archived status from a session.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
411
412
413
@abc.abstractmethod
def unarchive_session(self, session_id: str) -> None:
    """Remove archived status from a session."""

create_session_store(root_dir: str | None = None) -> SessionStoreBase

Return the configured session store driver.

Reads storage.driver from the app config. Defaults to "json" (filesystem). Set to "mongodb" to use MongoDB.

Source code in packages/mewbo_core/src/mewbo_core/session/session_store.py
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
def create_session_store(root_dir: str | None = None) -> SessionStoreBase:
    """Return the configured session store driver.

    Reads ``storage.driver`` from the app config.  Defaults to ``"json"``
    (filesystem).  Set to ``"mongodb"`` to use MongoDB.
    """
    driver = get_config_value("storage", "driver", default="json")
    if driver == "mongodb":
        from mewbo_core.session.session_store_mongo import MongoSessionStore

        try:
            return MongoSessionStore(root_dir=root_dir)
        except Exception as exc:
            raise RuntimeError(
                f"Storage driver is 'mongodb' but MongoDB is not available. "
                f"Check MEWBO_MONGODB_URI and ensure MongoDB is running. "
                f"Error: {exc}"
            ) from exc
    return SessionStore(root_dir=root_dir)

mewbo_core.session.context

Context selection and rendering helpers.

ContextBuilder

Build short-term and selected context for a session.

Source code in packages/mewbo_core/src/mewbo_core/session/context.py
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
class ContextBuilder:
    """Build short-term and selected context for a session."""

    def __init__(self, session_store: SessionStoreBase) -> None:
        """Initialize the context builder."""
        self._session_store = session_store

    def build(
        self,
        session_id: str,
        user_query: str,
        model_name: str | None,
    ) -> ContextSnapshot:
        """Build a context snapshot for planning and synthesis."""
        events = self._session_store.load_transcript(session_id)
        summary = self._session_store.load_summary(session_id)
        # Compaction boundary: events at or before the most recent
        # ``context_compacted`` event are already represented in ``summary``.
        # Replaying their raw payloads under "Recent conversation:" would
        # double-count them and leak pre-compaction noise that the user
        # explicitly asked to summarize away. Slice forward past the marker
        # by POSITION: the transcript is an append-only log, so index order is
        # what "after the boundary" means. Comparing ISO strings instead holds
        # only while every producer spells the UTC offset identically — ``Z``
        # sorts above ``+00:00``, so one same-second mix silently drops a
        # post-boundary event out of context.
        last_compact_idx = -1
        for idx, event in enumerate(events):
            if event.get("type") == "context_compacted":
                last_compact_idx = idx
        if last_compact_idx >= 0:
            events = events[last_compact_idx + 1 :]
        context_events = [
            event
            for event in events
            if event.get("type")
            in {
                "user",
                "assistant",
                "tool_result",
                "step_reflection",
                # The promise-as-completion nudge rides this event. Without it in
                # the filter the marker reaches the console, CLI and forensics but
                # never the model — the one place the whole overhaul aims it.
                "run_note",
            }
        ]
        recent_limit = int(get_config_value("context", "recent_event_limit", default=8))
        recent_events = context_events[-recent_limit:] if recent_limit > 0 else []
        candidate_events = context_events[:-recent_limit] if recent_limit > 0 else context_events
        # Anchor the first user event so long sessions cannot FIFO-evict the
        # original task from recent_events. Cached prefix-friendly: the anchor
        # renders into the SystemMessage's "Recent conversation:" block, which
        # is inside the cacheable prefix rather than the per-turn HumanMessage.
        # After a compaction boundary the original task already lives inside
        # ``summary``; only anchor the post-boundary first user event so we
        # don't re-introduce pre-boundary content the user just summarized.
        first_user = next((e for e in context_events if e.get("type") == "user"), None)
        if first_user is not None and first_user not in recent_events:
            recent_events = [first_user, *recent_events]
        from mewbo_core.session.token_budget import read_last_input_tokens

        last_input_tokens = read_last_input_tokens(list(events))
        budget = get_token_budget(
            events,
            summary,
            model_name,
            last_input_tokens=last_input_tokens,
        )
        selected_events: list[EventRecord] | None = None
        selection_threshold = float(get_config_value("context", "selection_threshold", default=0.8))
        if (
            bool(get_config_value("context", "selection_enabled", default=True))
            and candidate_events
            and budget.utilization >= selection_threshold
        ):
            selected_events = self._select_context_events(
                candidate_events,
                user_query=user_query,
                model_name=model_name,
            )
        session_dir = self._session_store.session_dir(session_id)
        attachment_texts = _load_attachment_texts(session_dir, events, model_name)
        attachment_images = _load_attachment_images(session_dir, events, model_name)
        return ContextSnapshot(
            summary=summary,
            recent_events=recent_events,
            selected_events=selected_events,
            events=events,
            budget=budget,
            attachment_texts=attachment_texts,
            attachment_images=attachment_images,
        )

    def _select_context_events(
        self,
        events: list[EventRecord],
        user_query: str,
        model_name: str | None,
    ) -> list[EventRecord]:
        if not events:
            return []
        selector_model = (
            get_config_value("context", "context_selector_model")
            or model_name
            or get_config_value("llm", "action_plan_model")
            or get_config_value("llm", "default_model")
        )
        if not selector_model:
            return events
        parser = PydanticOutputParser(pydantic_object=ContextSelection)
        prompt = ChatPromptTemplate(
            messages=[
                SystemMessage(
                    content=(
                        "You select which prior events are still relevant to the user's "
                        "current request. Keep only events that directly help answer the "
                        "current query. If unsure, keep the event."
                    )
                ),
                HumanMessagePromptTemplate.from_template(
                    "User query:\n{user_query}\n\n"
                    "Candidate events:\n{candidates}\n\n"
                    "Return keep_ids and drop_ids.\n{format_instructions}"
                ),
            ],
            partial_variables={"format_instructions": parser.get_format_instructions()},
            input_variables=["user_query", "candidates"],
        )
        lines: list[str] = []
        for idx, event in enumerate(events, start=1):
            text = event_payload_text(event)
            if not text:
                continue
            lines.append(f"{idx}. {event.get('type', 'event')}: {text}")
        candidates_text = "\n".join(lines).strip()
        if not candidates_text:
            return events
        model = build_chat_model(model_name=selector_model)
        handler = build_langfuse_handler(
            user_id="mewbo-context",
            session_id=f"context-{os.getpid()}-{os.urandom(4).hex()}",
            trace_name="context-select",
            version=get_version(),
            release=get_config_value("runtime", "envmode", default="Not Specified"),
        )
        config: dict[str, object] = {}
        if handler is not None:
            config["callbacks"] = [handler]
            metadata = getattr(handler, "langfuse_metadata", None)
            if isinstance(metadata, dict) and metadata:
                config["metadata"] = metadata
        try:
            with langfuse_trace_span(
                "context-select",
                metadata={
                    "model": selector_model,
                    "candidates": str(len(lines)),
                },
                input_data={
                    "user_query": user_query.strip()[:200],
                    "candidate_count": len(lines),
                },
            ) as span:
                selection = (prompt | model | parser).invoke(
                    {"user_query": user_query.strip(), "candidates": candidates_text},
                    config=config or None,
                )
                if span is not None:
                    try:
                        span.update_trace(
                            output={
                                "keep_ids": selection.keep_ids,
                                "drop_ids": selection.drop_ids,
                            }
                        )
                    except Exception:
                        pass
        except Exception as exc:  # pragma: no cover - defensive
            logging.warning("Context selection failed: {}", exc)
            return events[-3:]

        keep_ids = set(selection.keep_ids or [])
        if not keep_ids:
            return events[-3:]
        kept: list[EventRecord] = []
        for idx, event in enumerate(events, start=1):
            if idx in keep_ids:
                kept.append(event)
        return kept or events[-3:]

__init__(session_store: SessionStoreBase) -> None

Initialize the context builder.

Source code in packages/mewbo_core/src/mewbo_core/session/context.py
197
198
199
def __init__(self, session_store: SessionStoreBase) -> None:
    """Initialize the context builder."""
    self._session_store = session_store

build(session_id: str, user_query: str, model_name: str | None) -> ContextSnapshot

Build a context snapshot for planning and synthesis.

Source code in packages/mewbo_core/src/mewbo_core/session/context.py
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
def build(
    self,
    session_id: str,
    user_query: str,
    model_name: str | None,
) -> ContextSnapshot:
    """Build a context snapshot for planning and synthesis."""
    events = self._session_store.load_transcript(session_id)
    summary = self._session_store.load_summary(session_id)
    # Compaction boundary: events at or before the most recent
    # ``context_compacted`` event are already represented in ``summary``.
    # Replaying their raw payloads under "Recent conversation:" would
    # double-count them and leak pre-compaction noise that the user
    # explicitly asked to summarize away. Slice forward past the marker
    # by POSITION: the transcript is an append-only log, so index order is
    # what "after the boundary" means. Comparing ISO strings instead holds
    # only while every producer spells the UTC offset identically — ``Z``
    # sorts above ``+00:00``, so one same-second mix silently drops a
    # post-boundary event out of context.
    last_compact_idx = -1
    for idx, event in enumerate(events):
        if event.get("type") == "context_compacted":
            last_compact_idx = idx
    if last_compact_idx >= 0:
        events = events[last_compact_idx + 1 :]
    context_events = [
        event
        for event in events
        if event.get("type")
        in {
            "user",
            "assistant",
            "tool_result",
            "step_reflection",
            # The promise-as-completion nudge rides this event. Without it in
            # the filter the marker reaches the console, CLI and forensics but
            # never the model — the one place the whole overhaul aims it.
            "run_note",
        }
    ]
    recent_limit = int(get_config_value("context", "recent_event_limit", default=8))
    recent_events = context_events[-recent_limit:] if recent_limit > 0 else []
    candidate_events = context_events[:-recent_limit] if recent_limit > 0 else context_events
    # Anchor the first user event so long sessions cannot FIFO-evict the
    # original task from recent_events. Cached prefix-friendly: the anchor
    # renders into the SystemMessage's "Recent conversation:" block, which
    # is inside the cacheable prefix rather than the per-turn HumanMessage.
    # After a compaction boundary the original task already lives inside
    # ``summary``; only anchor the post-boundary first user event so we
    # don't re-introduce pre-boundary content the user just summarized.
    first_user = next((e for e in context_events if e.get("type") == "user"), None)
    if first_user is not None and first_user not in recent_events:
        recent_events = [first_user, *recent_events]
    from mewbo_core.session.token_budget import read_last_input_tokens

    last_input_tokens = read_last_input_tokens(list(events))
    budget = get_token_budget(
        events,
        summary,
        model_name,
        last_input_tokens=last_input_tokens,
    )
    selected_events: list[EventRecord] | None = None
    selection_threshold = float(get_config_value("context", "selection_threshold", default=0.8))
    if (
        bool(get_config_value("context", "selection_enabled", default=True))
        and candidate_events
        and budget.utilization >= selection_threshold
    ):
        selected_events = self._select_context_events(
            candidate_events,
            user_query=user_query,
            model_name=model_name,
        )
    session_dir = self._session_store.session_dir(session_id)
    attachment_texts = _load_attachment_texts(session_dir, events, model_name)
    attachment_images = _load_attachment_images(session_dir, events, model_name)
    return ContextSnapshot(
        summary=summary,
        recent_events=recent_events,
        selected_events=selected_events,
        events=events,
        budget=budget,
        attachment_texts=attachment_texts,
        attachment_images=attachment_images,
    )

ContextSelection

Bases: BaseModel

Model output for selecting context events.

Source code in packages/mewbo_core/src/mewbo_core/session/context.py
31
32
33
34
35
class ContextSelection(BaseModel):
    """Model output for selecting context events."""

    keep_ids: list[int] = Field(default_factory=list)
    drop_ids: list[int] = Field(default_factory=list)

ContextSnapshot dataclass

Context snapshot for planning and synthesis.

Source code in packages/mewbo_core/src/mewbo_core/session/context.py
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
@dataclass(frozen=True)
class ContextSnapshot:
    """Context snapshot for planning and synthesis."""

    summary: str | None
    recent_events: list[EventRecord]
    selected_events: list[EventRecord] | None
    events: list[EventRecord]
    budget: TokenBudget
    # Markdown-rendered text from documents (PDF/DOCX/CSV/...) and raw
    # text files. Injected into the system prompt's "Attached files:" block.
    attachment_texts: list[str] = field(default_factory=list)
    # LiteLLM-style image content parts: {"type": "image_url", "image_url": {"url": ...}}
    # Spliced into the per-turn HumanMessage on vision-capable models.
    attachment_images: list[dict] = field(default_factory=list)

event_payload_text(event: EventRecord) -> str

Return a readable payload string for an event.

Source code in packages/mewbo_core/src/mewbo_core/session/context.py
55
56
57
58
59
60
61
62
63
64
65
def event_payload_text(event: EventRecord) -> str:
    """Return a readable payload string for an event."""
    payload = event.get("payload", "")
    if isinstance(payload, dict):
        if "tool_input" in payload:
            payload = dict(payload)
            payload["tool_input"] = format_tool_input(payload.get("tool_input"))
        return str(
            payload.get("text") or payload.get("message") or payload.get("result") or payload
        )
    return str(payload)

render_event_lines(events: list[EventRecord]) -> str

Render events into bullet lines for prompts.

Source code in packages/mewbo_core/src/mewbo_core/session/context.py
68
69
70
71
72
73
74
75
76
def render_event_lines(events: list[EventRecord]) -> str:
    """Render events into bullet lines for prompts."""
    lines: list[str] = []
    for event in events:
        text = event_payload_text(event)
        if not text:
            continue
        lines.append(f"- {event.get('type', 'event')}: {text}")
    return "\n".join(lines).strip()

mewbo_core.session.compaction

Lossless pre-compaction utilities.

This module intentionally contains only one thing: a pre_compact hook that strips ANSI escapes and truncates huge tool outputs before the LLM summarizer sees them. Zero tokens, zero risk.

Prior versions also shipped should_compact (an event-count heuristic) and summarize_events (a fallback "summary" that concatenated raw event text). Both were deleted: compaction decisions are now driven purely by the API-reported usage_metadata.input_tokens, and failed structured compaction must not be masked with raw-text noise.

micro_compact_events(events: list[EventRecord]) -> list[EventRecord]

Lossless pre-compaction: strip ANSI escapes and truncate large tool outputs.

Intended for use as a pre_compact hook — no LLM call, zero cost.

The two jobs are INDEPENDENT and applied as such. They used to be one branch, so an escape-laden result under the cap kept every escape — and raising the cap silently widened that band. Stripping is unconditional; only the truncation consults the cap.

Source code in packages/mewbo_core/src/mewbo_core/session/compaction.py
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
def micro_compact_events(events: list[EventRecord]) -> list[EventRecord]:
    """Lossless pre-compaction: strip ANSI escapes and truncate large tool outputs.

    Intended for use as a ``pre_compact`` hook — no LLM call, zero cost.

    The two jobs are INDEPENDENT and applied as such. They used to be one
    branch, so an escape-laden result under the cap kept every escape — and
    raising the cap silently widened that band. Stripping is unconditional;
    only the truncation consults the cap.
    """
    compacted: list[EventRecord] = []
    for event in events:
        if event.get("type") != "tool_result":
            compacted.append(event)
            continue
        payload = event.get("payload")
        if not isinstance(payload, dict):
            compacted.append(event)
            continue
        result = payload.get("result")
        if not isinstance(result, str):
            compacted.append(event)
            continue
        stripped = _ANSI_RE.sub("", result)
        if len(result) > _MAX_RESULT_CHARS:
            # State the omission's SIZE. A mute marker leaves a bounded memory
            # indistinguishable from a complete one, which is how a model ends
            # up reasoning from the first fraction of a result as though it
            # were the whole — and arguing with a tool that told it no such
            # limit existed. Sized against the ORIGINAL, which is what the
            # model saw before compaction.
            omitted = len(result) - _MAX_RESULT_CHARS
            stripped = (
                stripped[:_MAX_RESULT_CHARS]
                + f"\n[... {omitted} of {len(result)} characters omitted by compaction]"
                + "\n[truncated]"
            )
        elif stripped == result:
            compacted.append(event)
            continue
        payload = dict(payload)
        payload["result"] = stripped
        compacted.append({**event, "payload": payload})
    return compacted

mewbo_core.session.token_budget

Token budgeting anchored on LiteLLM-authoritative model metadata.

Philosophy: trust the API, not estimates. LiteLLM's get_model_info is the source of truth for each model's max_input_tokens; LangChain's response.usage_metadata.input_tokens is the source of truth for what the current prompt actually consumed. No char-count heuristics, no fake overhead additions, no regex guessing from the model name.

TokenBudget dataclass

Token accounting snapshot used to decide compaction.

Source code in packages/mewbo_core/src/mewbo_core/session/token_budget.py
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
@dataclass(frozen=True)
class TokenBudget:
    """Token accounting snapshot used to decide compaction."""

    total_tokens: int
    summary_tokens: int
    event_tokens: int
    context_window: int
    remaining_tokens: int
    utilization: float
    threshold: float

    @property
    def needs_compact(self) -> bool:
        """Return True when utilization meets or exceeds the configured threshold."""
        return self.utilization >= self.threshold

needs_compact: bool property

Return True when utilization meets or exceeds the configured threshold.

build_usage_numbers(events: list[EventRecord], root_model: str | None) -> dict[str, Any]

Walk a transcript once and return raw usage numbers.

Split by depth so clients can render root (hypervisor) vs sub-agents without conflation. Numbers only — no formatting, no color states, no labels. Clients format what they need.

Returned keys (input has two semantics, output only one):

Context-pressure (peak) — what matters for compaction / window math. input_tokens on an llm_call_end event is the prompt size for that call. Within a turn the prompt GROWS as tool results accumulate (step 1: ~13K baseline, step 11: ~27K), so summing gives a nonsense number that double-counts the baseline once per call. The peak (max across root calls) is the real context pressure: - root_peak_input_tokens: max input_tokens seen on any depth==0 call. - sub_peak_input_tokens: sum of per-sub-agent peaks (each sub-agent runs in an isolated context, so summing their peaks — not their sum inputs — represents "combined peak pressure of parallel sub-contexts").

Cumulative (billable) — what the provider charges for. Sum across all calls. Useful for cost dashboards. Named with _billed_in suffix to make the semantic explicit: - root_input_tokens_billed, sub_input_tokens_billed.

Output is additive everywhere. Each output token is produced once; summing is correct: - root_output_tokens, sub_output_tokens.

Other keys: root_model, root_max_input_tokens, root_last_input_tokens, root_utilization, tokens_until_compact, compact_threshold, root_llm_calls, sub_llm_calls, sub_agent_count, total_input_tokens_billed, total_output_tokens, compaction_count, compaction_tokens_saved, models_used (distinct model IDs in first-seen order, collected from context.model, sub_agent.model, and llm_fallback.to_model).

Source code in packages/mewbo_core/src/mewbo_core/session/token_budget.py
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
def build_usage_numbers(
    events: list[EventRecord],
    root_model: str | None,
) -> dict[str, Any]:
    """Walk a transcript once and return raw usage numbers.

    Split by depth so clients can render root (hypervisor) vs sub-agents
    without conflation. Numbers only — no formatting, no color states, no
    labels. Clients format what they need.

    Returned keys (input has two semantics, output only one):

    **Context-pressure (peak) — what matters for compaction / window math.**
    ``input_tokens`` on an ``llm_call_end`` event is the prompt size for
    that call. Within a turn the prompt GROWS as tool results accumulate
    (step 1: ~13K baseline, step 11: ~27K), so summing gives a nonsense
    number that double-counts the baseline once per call. The peak (max
    across root calls) is the real context pressure:
      - ``root_peak_input_tokens``: max input_tokens seen on any depth==0 call.
      - ``sub_peak_input_tokens``: sum of per-sub-agent peaks (each sub-agent
        runs in an isolated context, so summing their peaks — not their sum
        inputs — represents "combined peak pressure of parallel sub-contexts").

    **Cumulative (billable) — what the provider charges for.**
    Sum across all calls. Useful for cost dashboards. Named with
    ``_billed_in`` suffix to make the semantic explicit:
      - ``root_input_tokens_billed``, ``sub_input_tokens_billed``.

    **Output is additive everywhere.** Each output token is produced once;
    summing is correct:
      - ``root_output_tokens``, ``sub_output_tokens``.

    Other keys: ``root_model``, ``root_max_input_tokens``,
    ``root_last_input_tokens``, ``root_utilization``,
    ``tokens_until_compact``, ``compact_threshold``,
    ``root_llm_calls``, ``sub_llm_calls``, ``sub_agent_count``,
    ``total_input_tokens_billed``, ``total_output_tokens``,
    ``compaction_count``, ``compaction_tokens_saved``,
    ``models_used`` (distinct model IDs in first-seen order, collected from
    ``context.model``, ``sub_agent.model``, and ``llm_fallback.to_model``).
    """
    root_input_billed = root_output = root_calls = 0
    root_peak_input = 0
    root_cache_creation = root_cache_read = root_reasoning = 0
    sub_input_billed = sub_output = sub_calls = 0
    sub_peak_per_agent: dict[str, int] = {}
    sub_cache_creation = sub_cache_read = sub_reasoning = 0
    sub_agent_ids: set[str] = set()
    compaction_count = 0
    compaction_tokens_saved = 0
    # Order-preserving dedupe: dict keys preserve insertion order in Python 3.7+.
    models_seen: dict[str, None] = {}

    for event in events:
        etype = event.get("type")
        payload = event.get("payload")
        if not isinstance(payload, dict):
            continue
        if etype == "context":
            model = payload.get("model")
            if isinstance(model, str) and model:
                models_seen[model] = None
        elif etype == "sub_agent":
            model = payload.get("model")
            if isinstance(model, str) and model:
                models_seen[model] = None
        elif etype == "llm_fallback":
            model = payload.get("to_model")
            if isinstance(model, str) and model:
                models_seen[model] = None
        if etype == "llm_call_end":
            depth = payload.get("depth", 0)
            in_tok = int(payload.get("input_tokens", 0) or 0)
            out_tok = int(payload.get("output_tokens", 0) or 0)
            # Cache + reasoning subtotals default to 0 on an event that did not
            # capture them, so accumulating is a no-op there.
            cache_create = int(payload.get("cache_creation_input_tokens", 0) or 0)
            cache_read = int(payload.get("cache_read_input_tokens", 0) or 0)
            reasoning = int(payload.get("reasoning_output_tokens", 0) or 0)
            if depth == 0:
                root_input_billed += in_tok
                root_output += out_tok
                root_calls += 1
                root_cache_creation += cache_create
                root_cache_read += cache_read
                root_reasoning += reasoning
                if in_tok > root_peak_input:
                    root_peak_input = in_tok
            else:
                sub_input_billed += in_tok
                sub_output += out_tok
                sub_calls += 1
                sub_cache_creation += cache_create
                sub_cache_read += cache_read
                sub_reasoning += reasoning
                agent_id = payload.get("agent_id")
                if isinstance(agent_id, str):
                    sub_agent_ids.add(agent_id)
                    prev = sub_peak_per_agent.get(agent_id, 0)
                    if in_tok > prev:
                        sub_peak_per_agent[agent_id] = in_tok
        elif etype == "context_compacted":
            compaction_count += 1
            saved = payload.get("tokens_saved", 0)
            if isinstance(saved, int) and saved > 0:
                compaction_tokens_saved += saved

    sub_peak_input = sum(sub_peak_per_agent.values())
    max_input = get_model_max_input_tokens(root_model)
    last_input = read_last_input_tokens(events) or 0
    threshold = float(get_config_value("token_budget", "auto_compact_threshold", default=0.8))
    compact_at = int(max_input * threshold)
    utilization = (last_input / max_input) if max_input else 0.0
    tokens_until_compact = max(compact_at - last_input, 0)

    return {
        "root_model": root_model or "",
        "root_max_input_tokens": max_input,
        "root_last_input_tokens": last_input,
        "root_utilization": round(utilization, 4),
        "tokens_until_compact": tokens_until_compact,
        "compact_threshold": threshold,
        # Context-pressure (peak) — use these for window math.
        "root_peak_input_tokens": root_peak_input,
        "sub_peak_input_tokens": sub_peak_input,
        # Cumulative (billable) — use these for cost. ``input_tokens_billed``
        # is the raw provider count INCLUDING cached portions; pair it with
        # ``cache_read_tokens`` if you need to apply the discount client-side
        # (Anthropic cache reads bill at 0.1×, OpenAI at 0.5×).
        "root_input_tokens_billed": root_input_billed,
        "sub_input_tokens_billed": sub_input_billed,
        "total_input_tokens_billed": root_input_billed + sub_input_billed,
        # Output is additive everywhere.
        "root_output_tokens": root_output,
        "sub_output_tokens": sub_output,
        "total_output_tokens": root_output + sub_output,
        # Cache + reasoning subtotals (zero on an event that did not capture
        # them). Cache reads served from prompt cache; cache creation
        # tokens written to cache; reasoning tokens are the hidden output of
        # extended-thinking / o1-class models.
        "root_cache_creation_tokens": root_cache_creation,
        "root_cache_read_tokens": root_cache_read,
        "root_reasoning_tokens": root_reasoning,
        "sub_cache_creation_tokens": sub_cache_creation,
        "sub_cache_read_tokens": sub_cache_read,
        "sub_reasoning_tokens": sub_reasoning,
        "total_cache_creation_tokens": root_cache_creation + sub_cache_creation,
        "total_cache_read_tokens": root_cache_read + sub_cache_read,
        "total_reasoning_tokens": root_reasoning + sub_reasoning,
        "root_llm_calls": root_calls,
        "sub_llm_calls": sub_calls,
        "sub_agent_count": len(sub_agent_ids),
        "compaction_count": compaction_count,
        "compaction_tokens_saved": compaction_tokens_saved,
        "models_used": list(models_seen),
    }

estimate_event_tokens(events: Iterable[EventRecord]) -> int

Estimate total tokens for a sequence of events (fallback only).

Used only when no real usage_metadata is available (e.g. a fresh session before the first LLM call). After the first response lands, last_input_tokens from the API is the authoritative signal.

Source code in packages/mewbo_core/src/mewbo_core/session/token_budget.py
180
181
182
183
184
185
186
187
188
189
190
191
def estimate_event_tokens(events: Iterable[EventRecord]) -> int:
    """Estimate total tokens for a sequence of events (fallback only).

    Used only when no real ``usage_metadata`` is available (e.g. a fresh
    session before the first LLM call). After the first response lands,
    ``last_input_tokens`` from the API is the authoritative signal.
    """
    texts = [_event_to_text(event) for event in events]
    joined = "\n".join(text for text in texts if text)
    if not joined:
        return 0
    return num_tokens_from_string(joined)

estimate_summary_tokens(summary: str | None) -> int

Estimate token usage for the stored summary.

Source code in packages/mewbo_core/src/mewbo_core/session/token_budget.py
194
195
196
197
198
def estimate_summary_tokens(summary: str | None) -> int:
    """Estimate token usage for the stored summary."""
    if not summary:
        return 0
    return num_tokens_from_string(summary)

forget_cached_context_windows() -> None

Drop memoized catalogue answers after the catalogue changes.

_litellm_max_input_tokens memoizes a MISS as readily as a hit, so a lookup that ran before the proxy bridge hydrated the catalogue would pin the pre-hydration answer for the life of the process — and the API serves usage reads that resolve a window without ever constructing a client. Registration is the only event that invalidates those answers.

Source code in packages/mewbo_core/src/mewbo_core/session/token_budget.py
151
152
153
154
155
156
157
158
159
160
def forget_cached_context_windows() -> None:
    """Drop memoized catalogue answers after the catalogue changes.

    ``_litellm_max_input_tokens`` memoizes a MISS as readily as a hit, so a
    lookup that ran before the proxy bridge hydrated the catalogue would pin
    the pre-hydration answer for the life of the process — and the API serves
    usage reads that resolve a window without ever constructing a client.
    Registration is the only event that invalidates those answers.
    """
    _litellm_max_input_tokens.cache_clear()

get_model_max_input_tokens(model_name: str | None) -> int

Resolve the maximum input tokens for a model.

Priority: user override -> LiteLLM catalogue -> config default.

Source code in packages/mewbo_core/src/mewbo_core/session/token_budget.py
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
def get_model_max_input_tokens(model_name: str | None) -> int:
    """Resolve the maximum input tokens for a model.

    Priority: user override -> LiteLLM catalogue -> config default.
    """
    default_window = int(get_config_value("token_budget", "default_context_window", default=128000))
    if not model_name:
        return default_window
    overrides = _load_context_overrides()
    if model_name in overrides:
        return overrides[model_name]
    canonical = _strip_provider_prefix(model_name)
    if canonical in overrides:
        return overrides[canonical]
    # The PREFIXED name is tried first, and that order is load-bearing.
    # ``register_proxy_model_capabilities`` hydrates the catalogue under
    # ``openai/<name>``, so stripping before the lookup queried the one key the
    # proxy bridge never writes — a model absent from LiteLLM's bundled map
    # (a proxy-only route) silently fell back to the default window while its
    # real one sat in the catalogue under the prefixed spelling.
    for candidate in dict.fromkeys((model_name, canonical)):
        from_litellm = _litellm_max_input_tokens(candidate)
        if from_litellm is not None:
            return from_litellm
    _warn_fallback_context_window(model_name, default_window)
    return default_window

get_token_budget(events: Iterable[EventRecord], summary: str | None, model_name: str | None, threshold: float | None = None, *, last_input_tokens: int | None = None) -> TokenBudget

Calculate token utilization and remaining context budget.

When last_input_tokens is supplied (from a real response.usage_metadata read), it is used as the authoritative total. Otherwise we fall back to estimating from events + summary.

Source code in packages/mewbo_core/src/mewbo_core/session/token_budget.py
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
def get_token_budget(
    events: Iterable[EventRecord],
    summary: str | None,
    model_name: str | None,
    threshold: float | None = None,
    *,
    last_input_tokens: int | None = None,
) -> TokenBudget:
    """Calculate token utilization and remaining context budget.

    When ``last_input_tokens`` is supplied (from a real
    ``response.usage_metadata`` read), it is used as the authoritative
    total. Otherwise we fall back to estimating from events + summary.
    """
    if threshold is None:
        threshold = float(get_config_value("token_budget", "auto_compact_threshold", default=0.8))
    context_window = get_model_max_input_tokens(model_name)

    if last_input_tokens is not None and last_input_tokens > 0:
        # API-reported usage is authoritative. event_tokens/summary_tokens
        # are recomputed only for diagnostics.
        event_tokens = estimate_event_tokens(events)
        summary_tokens = estimate_summary_tokens(summary)
        total_tokens = last_input_tokens
    else:
        event_tokens = estimate_event_tokens(events)
        summary_tokens = estimate_summary_tokens(summary)
        total_tokens = event_tokens + summary_tokens

    remaining_tokens = max(context_window - total_tokens, 0)
    utilization = total_tokens / context_window if context_window else 0.0
    return TokenBudget(
        total_tokens=total_tokens,
        summary_tokens=summary_tokens,
        event_tokens=event_tokens,
        context_window=context_window,
        remaining_tokens=remaining_tokens,
        utilization=utilization,
        threshold=threshold,
    )

read_last_input_tokens(events: list[EventRecord]) -> int | None

Return the most recent llm_call_end event's input_tokens.

Session-store counterpart to the in-memory _last_input_tokens field on ToolUseLoop. Lets callers outside the loop (e.g. the orchestrator's compaction check) read the authoritative per-call token count that was already persisted.

Source code in packages/mewbo_core/src/mewbo_core/session/token_budget.py
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
def read_last_input_tokens(events: list[EventRecord]) -> int | None:
    """Return the most recent ``llm_call_end`` event's ``input_tokens``.

    Session-store counterpart to the in-memory ``_last_input_tokens`` field
    on ``ToolUseLoop``. Lets callers outside the loop (e.g. the orchestrator's
    compaction check) read the authoritative per-call token count that was
    already persisted.
    """
    for event in reversed(events):
        if event.get("type") != "llm_call_end":
            continue
        payload = event.get("payload")
        if not isinstance(payload, dict):
            continue
        value = payload.get("input_tokens")
        if isinstance(value, int) and value > 0:
            return value
    return None

mewbo_core.tooling.tool_registry

Tool registry and manifest loading for Mewbo.

ToolRegistry

Registry of configured tools and their instantiated runners.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
class ToolRegistry:
    """Registry of configured tools and their instantiated runners."""

    def __init__(self) -> None:
        """Initialize an empty registry."""
        self._tools: dict[str, ToolSpec] = {}
        self._instances: dict[str, ToolRunner] = {}

    def disable(self, tool_id: str, reason: str) -> None:
        """Disable a tool and store a reason for later reporting."""
        spec = self._tools.get(tool_id)
        if spec is None:
            return
        metadata = dict(spec.metadata)
        metadata["disabled_reason"] = reason
        self._tools[tool_id] = ToolSpec(
            tool_id=spec.tool_id,
            name=spec.name,
            description=spec.description,
            factory=spec.factory,
            enabled=False,
            kind=spec.kind,
            prompt_path=spec.prompt_path,
            metadata=metadata,
            concurrency_safe=spec.concurrency_safe,
            read_only=spec.read_only,
            interrupt_behavior=spec.interrupt_behavior,
            max_result_chars=spec.max_result_chars,
            timeout=spec.timeout,
            capability=spec.capability,
            poll=spec.poll,
            poll_when_args=spec.poll_when_args,
        )
        if tool_id in self._instances:
            self._instances.pop(tool_id, None)
        set_available_tools(
            [current_id for current_id, current_spec in self._tools.items() if current_spec.enabled]
        )

    def register(self, spec: ToolSpec) -> None:
        """Register a tool specification and update action validation."""
        self._tools[spec.tool_id] = spec
        set_available_tools(
            [tool_id for tool_id, tool_spec in self._tools.items() if tool_spec.enabled]
        )

    def get(self, tool_id: str) -> ToolRunner | None:
        """Return an enabled tool runner, instantiating it if needed."""
        spec = self._tools.get(tool_id)
        if spec is None or not spec.enabled:
            return None
        if tool_id not in self._instances:
            try:
                self._instances[tool_id] = spec.factory()
            except Exception as exc:  # pragma: no cover - defensive
                reason = f"Initialization failed: {exc}"
                logging.warning("Disabling tool {}: {}", tool_id, reason)
                self.disable(tool_id, reason)
                return None
        return self._instances[tool_id]

    def get_spec(self, tool_id: str) -> ToolSpec | None:
        """Return the tool specification, even if disabled."""
        return self._tools.get(tool_id)

    def list_specs(self, include_disabled: bool = False) -> list[ToolSpec]:
        """List tool specifications, optionally including disabled tools."""
        specs = list(self._tools.values())
        if include_disabled:
            return specs
        return [spec for spec in specs if spec.enabled]

    def tool_catalog(self) -> list[dict[str, str]]:
        """Return a serialized catalog of registered tool metadata."""
        return [
            {
                "tool_id": spec.tool_id,
                "name": spec.name,
                "description": spec.description,
            }
            for spec in self.list_specs()
        ]

__init__() -> None

Initialize an empty registry.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
208
209
210
211
def __init__(self) -> None:
    """Initialize an empty registry."""
    self._tools: dict[str, ToolSpec] = {}
    self._instances: dict[str, ToolRunner] = {}

disable(tool_id: str, reason: str) -> None

Disable a tool and store a reason for later reporting.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
def disable(self, tool_id: str, reason: str) -> None:
    """Disable a tool and store a reason for later reporting."""
    spec = self._tools.get(tool_id)
    if spec is None:
        return
    metadata = dict(spec.metadata)
    metadata["disabled_reason"] = reason
    self._tools[tool_id] = ToolSpec(
        tool_id=spec.tool_id,
        name=spec.name,
        description=spec.description,
        factory=spec.factory,
        enabled=False,
        kind=spec.kind,
        prompt_path=spec.prompt_path,
        metadata=metadata,
        concurrency_safe=spec.concurrency_safe,
        read_only=spec.read_only,
        interrupt_behavior=spec.interrupt_behavior,
        max_result_chars=spec.max_result_chars,
        timeout=spec.timeout,
        capability=spec.capability,
        poll=spec.poll,
        poll_when_args=spec.poll_when_args,
    )
    if tool_id in self._instances:
        self._instances.pop(tool_id, None)
    set_available_tools(
        [current_id for current_id, current_spec in self._tools.items() if current_spec.enabled]
    )

get(tool_id: str) -> ToolRunner | None

Return an enabled tool runner, instantiating it if needed.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
251
252
253
254
255
256
257
258
259
260
261
262
263
264
def get(self, tool_id: str) -> ToolRunner | None:
    """Return an enabled tool runner, instantiating it if needed."""
    spec = self._tools.get(tool_id)
    if spec is None or not spec.enabled:
        return None
    if tool_id not in self._instances:
        try:
            self._instances[tool_id] = spec.factory()
        except Exception as exc:  # pragma: no cover - defensive
            reason = f"Initialization failed: {exc}"
            logging.warning("Disabling tool {}: {}", tool_id, reason)
            self.disable(tool_id, reason)
            return None
    return self._instances[tool_id]

get_spec(tool_id: str) -> ToolSpec | None

Return the tool specification, even if disabled.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
266
267
268
def get_spec(self, tool_id: str) -> ToolSpec | None:
    """Return the tool specification, even if disabled."""
    return self._tools.get(tool_id)

list_specs(include_disabled: bool = False) -> list[ToolSpec]

List tool specifications, optionally including disabled tools.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
270
271
272
273
274
275
def list_specs(self, include_disabled: bool = False) -> list[ToolSpec]:
    """List tool specifications, optionally including disabled tools."""
    specs = list(self._tools.values())
    if include_disabled:
        return specs
    return [spec for spec in specs if spec.enabled]

register(spec: ToolSpec) -> None

Register a tool specification and update action validation.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
244
245
246
247
248
249
def register(self, spec: ToolSpec) -> None:
    """Register a tool specification and update action validation."""
    self._tools[spec.tool_id] = spec
    set_available_tools(
        [tool_id for tool_id, tool_spec in self._tools.items() if tool_spec.enabled]
    )

tool_catalog() -> list[dict[str, str]]

Return a serialized catalog of registered tool metadata.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
277
278
279
280
281
282
283
284
285
286
def tool_catalog(self) -> list[dict[str, str]]:
    """Return a serialized catalog of registered tool metadata."""
    return [
        {
            "tool_id": spec.tool_id,
            "name": spec.name,
            "description": spec.description,
        }
        for spec in self.list_specs()
    ]

ToolRegistryCache

Process-wide reuse of built ToolRegistry objects, keyed by build inputs.

Orchestrator.__init__ builds the tool registry on every run (per-query), paying load_registry's cost — reading the MCP manifest, constructing every ToolSpec, probing Home-Assistant/LSP status — before the run emits its first event. That work is identical across runs whose inputs are identical, so the result is cached and the SAME registry handed back on the next run within a session (kill the blank-shell wait).

The key is exactly the set of inputs that change what load_registry produces: the project cwd, any plugin-contributed extra_mcp_servers, and the resolved MCP-config fingerprint (so an edit to mcp.json still forces a rebuild — the same signal that gates the manifest rebuild in _ensure_auto_manifest). Allowed-tools scoping is applied per-run downstream (filter_specs over list_specs), never at build time, so it is deliberately NOT part of the key.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
class ToolRegistryCache:
    """Process-wide reuse of built ``ToolRegistry`` objects, keyed by build inputs.

    ``Orchestrator.__init__`` builds the tool registry on every run (per-query),
    paying ``load_registry``'s cost — reading the MCP manifest, constructing every
    ``ToolSpec``, probing Home-Assistant/LSP status — *before* the run emits its
    first event. That work is identical across runs whose inputs are identical, so
    the result is cached and the SAME registry handed back on the next run within a
    session (kill the blank-shell wait).

    The key is exactly the set of inputs that change what ``load_registry``
    produces: the project ``cwd``, any plugin-contributed ``extra_mcp_servers``,
    and the resolved MCP-config fingerprint (so an edit to ``mcp.json`` still forces
    a rebuild — the same signal that gates the manifest rebuild in
    ``_ensure_auto_manifest``). Allowed-tools scoping is applied per-run downstream
    (``filter_specs`` over ``list_specs``), never at build time, so it is
    deliberately NOT part of the key.
    """

    def __init__(self) -> None:
        """Initialize an empty, lock-guarded cache."""
        self._lock = threading.Lock()
        self._cache: dict[str, ToolRegistry] = {}

    @staticmethod
    def _key(
        cwd: str | None,
        extra_mcp_servers: dict[str, dict] | None,
        trust_cwd: bool = True,
    ) -> str:
        servers = (
            json.dumps(extra_mcp_servers, sort_keys=True, default=str)
            if extra_mcp_servers
            else ""
        )
        config = _resolve_mcp_config(
            get_mcp_config_path() or "",
            cwd=cwd,
            extra_mcp_servers=extra_mcp_servers,
            trust_cwd=trust_cwd,
        )
        fingerprint = _config_fingerprint(config) or ""
        # `trust_cwd` is part of the key, not just the build: two scopes sharing
        # a cwd but disagreeing about its trust must never share a registry —
        # the trusted one's cached MCP runners carry `trust_cwd=True`.
        trust = "trusted" if trust_cwd else "untrusted"
        return "\x00".join((cwd or "", servers, fingerprint, trust))

    def get_or_build(
        self,
        *,
        cwd: str | None = None,
        extra_mcp_servers: dict[str, dict] | None = None,
        trust_cwd: bool = True,
    ) -> ToolRegistry:
        """Return a cached registry for these inputs, building exactly once."""
        from mewbo_core.config import is_untrusted_cwd

        trust_cwd = trust_cwd and not is_untrusted_cwd(cwd)
        key = self._key(cwd, extra_mcp_servers, trust_cwd)
        with self._lock:
            cached = self._cache.get(key)
        if cached is not None:
            return cached
        # Build OUTSIDE the lock so a slow ``load_registry`` for one scope never
        # blocks another scope's lookup. A rare concurrent cold miss builds twice;
        # ``setdefault`` makes both callers converge on the first-stored instance.
        registry = load_registry(
            cwd=cwd, extra_mcp_servers=extra_mcp_servers, trust_cwd=trust_cwd
        )
        with self._lock:
            return self._cache.setdefault(key, registry)

    def clear(self) -> None:
        """Drop every cached registry."""
        with self._lock:
            self._cache.clear()

__init__() -> None

Initialize an empty, lock-guarded cache.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
1519
1520
1521
1522
def __init__(self) -> None:
    """Initialize an empty, lock-guarded cache."""
    self._lock = threading.Lock()
    self._cache: dict[str, ToolRegistry] = {}

clear() -> None

Drop every cached registry.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
1573
1574
1575
1576
def clear(self) -> None:
    """Drop every cached registry."""
    with self._lock:
        self._cache.clear()

get_or_build(*, cwd: str | None = None, extra_mcp_servers: dict[str, dict] | None = None, trust_cwd: bool = True) -> ToolRegistry

Return a cached registry for these inputs, building exactly once.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
def get_or_build(
    self,
    *,
    cwd: str | None = None,
    extra_mcp_servers: dict[str, dict] | None = None,
    trust_cwd: bool = True,
) -> ToolRegistry:
    """Return a cached registry for these inputs, building exactly once."""
    from mewbo_core.config import is_untrusted_cwd

    trust_cwd = trust_cwd and not is_untrusted_cwd(cwd)
    key = self._key(cwd, extra_mcp_servers, trust_cwd)
    with self._lock:
        cached = self._cache.get(key)
    if cached is not None:
        return cached
    # Build OUTSIDE the lock so a slow ``load_registry`` for one scope never
    # blocks another scope's lookup. A rare concurrent cold miss builds twice;
    # ``setdefault`` makes both callers converge on the first-stored instance.
    registry = load_registry(
        cwd=cwd, extra_mcp_servers=extra_mcp_servers, trust_cwd=trust_cwd
    )
    with self._lock:
        return self._cache.setdefault(key, registry)

ToolRunner

Bases: Protocol

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
79
80
81
82
83
84
85
86
87
88
class ToolRunner(Protocol):
    def run(self, action_step: ActionStep) -> MockSpeaker:  # pragma: no cover
        """Execute an action step and return a speaker response.

        Args:
            action_step: Action step payload to execute.

        Returns:
            MockSpeaker response from the tool.
        """

run(action_step: ActionStep) -> MockSpeaker

Execute an action step and return a speaker response.

Parameters:

Name Type Description Default
action_step ActionStep

Action step payload to execute.

required

Returns:

Type Description
MockSpeaker

MockSpeaker response from the tool.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
80
81
82
83
84
85
86
87
88
def run(self, action_step: ActionStep) -> MockSpeaker:  # pragma: no cover
    """Execute an action step and return a speaker response.

    Args:
        action_step: Action step payload to execute.

    Returns:
        MockSpeaker response from the tool.
    """

ToolSpec dataclass

Metadata describing a tool available to the assistant.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
@dataclass(frozen=True)
class ToolSpec:
    """Metadata describing a tool available to the assistant."""

    tool_id: str
    name: str
    description: str
    factory: Callable[[], ToolRunner]
    enabled: bool = True
    kind: str = "local"
    prompt_path: str | None = None
    metadata: dict[str, JsonValue] = field(default_factory=dict)
    concurrency_safe: bool = True  # Can run in parallel (True for backward compat)
    read_only: bool = False  # No side effects
    interrupt_behavior: str = "block"  # "cancel" or "block" on user interrupt
    max_result_chars: int = 2000  # Per-tool result size cap (0 = unlimited)
    timeout: float = 120.0  # Per-tool execution timeout in seconds
    # POLL-CLASS: this tool's documented contract is "call me again until the
    # thing I front settles" (a nested run's status probe, a wait primitive).
    # Repeated identical calls to it are honest waiting, so ``DoomLoopGuard``
    # drops them from the no-progress signature. Declared HERE rather than as
    # another hardcoded id in ``llm_resilience``: two sub-second "processing"
    # answers from a self-polling run are indistinguishable from a stuck loop
    # by repetition alone, and only the tool knows which it is.
    #
    # ``poll`` exempts EVERY call (a pure wait primitive). ``poll_when_args``
    # exempts only calls carrying one of the named arguments, which is what a
    # tool that both STARTS and POLLS the same work needs — one tool id, two
    # meanings, told apart by argument shape alone. Exempting such an id
    # outright would blind the guard to an agent re-issuing the same start
    # forever.
    poll: bool = False
    poll_when_args: tuple[str, ...] = ()
    # Declared privilege tier for delegation ``capability_mode`` filtering
    #. ``None`` = undeclared (treated as NOT read — safe-deny).
    # Distinct from ``metadata["capabilities"]`` (a permission-subsystem tag):
    # this is the coarse read/write/execute privilege a spawn's
    # ``capability_mode`` gates on. A ``read_only`` tool needs no explicit
    # value — ``capability_tier`` reads it as ``read`` — so only write/execute
    # tools declare one here.
    capability: Literal["read", "write", "execute"] | None = None

    # The knobs a BUILT-IN registration DECLARES and a manifest entry merely
    # CACHES. Each has a class default that is a decision made by silence — so
    # an entry simply LACKING the key is indistinguishable from one deliberately
    # choosing the default, and reads as the default either way.
    #
    # Deliberately excluded: identity (``tool_id``/``name``/``prompt_path``),
    # ``description`` (the two sides word one tool differently on purpose),
    # ``kind``/``metadata`` (the manifest carries discovery's schema), and
    # ``enabled`` — a manifest records a tool discovery found UNREACHABLE, and
    # overlaying that would resurrect it.
    DECLARED_KNOBS: ClassVar[tuple[str, ...]] = (
        "max_result_chars",
        "timeout",
        "concurrency_safe",
        "interrupt_behavior",
        "poll",
        "poll_when_args",
        "read_only",
        "capability",
    )

    def knob_mismatches(self, declared: ToolSpec) -> dict[str, tuple[object, object]]:
        """Fields where THIS spec disagrees with its *declared* counterpart.

        Maps each differing field to ``(mine, declared)``. Pure comparison, no
        I/O — the two specs arrive as values so a caller can compare a
        manifest-loaded spec against the built-in registration that authored it.

        This exists because the defect class is *cached data outliving the
        declaration*, which nothing else can see: a manifest written before a
        field existed carries no key for it, the reader defaults it, and the
        result is a spec that is internally consistent, loads clean, and is
        wrong. Only a comparison against the declaration catches that.
        """
        return {
            name: (mine, theirs)
            for name in self.DECLARED_KNOBS
            if (mine := getattr(self, name)) != (theirs := getattr(declared, name))
        }

    def with_declared_knobs(self, declared: ToolSpec) -> ToolSpec:
        """Return this spec with *declared*'s knob values overlaid.

        The built-in registration is the AUTHORITY for these fields and the
        manifest entry is a cache of it, so the declaration wins — which is what
        lets an already-written manifest self-heal on the next load with no
        operator step and no regeneration. Everything else on the manifest entry
        (identity, schema, enablement, MCP wiring) is untouched.
        """
        return replace(
            self, **{name: getattr(declared, name) for name in self.DECLARED_KNOBS}
        )

    def capability_tier(self) -> str | None:
        """Resolve this tool's privilege tier for ``capability_mode`` filtering.

        Returns ``"read"`` / ``"write"`` / ``"execute"``, or ``None`` when the
        tool makes no declaration. A ``read_only`` tool is read-tier by
        construction (no side effects), so the fallback keeps ``read_only`` and
        ``capability`` in lockstep — a new read-only tool is correctly admitted
        under ``read_only`` mode without a second annotation, and the failure
        mode for a *forgotten* declaration is safe-deny, not silent grant.
        ``None`` is deliberately NOT treated as read: an undeclared tool is
        withheld from a ``read_only`` child.
        """
        if self.capability is not None:
            return self.capability
        if self.read_only:
            return "read"
        return None

capability_tier() -> str | None

Resolve this tool's privilege tier for capability_mode filtering.

Returns "read" / "write" / "execute", or None when the tool makes no declaration. A read_only tool is read-tier by construction (no side effects), so the fallback keeps read_only and capability in lockstep — a new read-only tool is correctly admitted under read_only mode without a second annotation, and the failure mode for a forgotten declaration is safe-deny, not silent grant. None is deliberately NOT treated as read: an undeclared tool is withheld from a read_only child.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
def capability_tier(self) -> str | None:
    """Resolve this tool's privilege tier for ``capability_mode`` filtering.

    Returns ``"read"`` / ``"write"`` / ``"execute"``, or ``None`` when the
    tool makes no declaration. A ``read_only`` tool is read-tier by
    construction (no side effects), so the fallback keeps ``read_only`` and
    ``capability`` in lockstep — a new read-only tool is correctly admitted
    under ``read_only`` mode without a second annotation, and the failure
    mode for a *forgotten* declaration is safe-deny, not silent grant.
    ``None`` is deliberately NOT treated as read: an undeclared tool is
    withheld from a ``read_only`` child.
    """
    if self.capability is not None:
        return self.capability
    if self.read_only:
        return "read"
    return None

knob_mismatches(declared: ToolSpec) -> dict[str, tuple[object, object]]

Fields where THIS spec disagrees with its declared counterpart.

Maps each differing field to (mine, declared). Pure comparison, no I/O — the two specs arrive as values so a caller can compare a manifest-loaded spec against the built-in registration that authored it.

This exists because the defect class is cached data outliving the declaration, which nothing else can see: a manifest written before a field existed carries no key for it, the reader defaults it, and the result is a spec that is internally consistent, loads clean, and is wrong. Only a comparison against the declaration catches that.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
def knob_mismatches(self, declared: ToolSpec) -> dict[str, tuple[object, object]]:
    """Fields where THIS spec disagrees with its *declared* counterpart.

    Maps each differing field to ``(mine, declared)``. Pure comparison, no
    I/O — the two specs arrive as values so a caller can compare a
    manifest-loaded spec against the built-in registration that authored it.

    This exists because the defect class is *cached data outliving the
    declaration*, which nothing else can see: a manifest written before a
    field existed carries no key for it, the reader defaults it, and the
    result is a spec that is internally consistent, loads clean, and is
    wrong. Only a comparison against the declaration catches that.
    """
    return {
        name: (mine, theirs)
        for name in self.DECLARED_KNOBS
        if (mine := getattr(self, name)) != (theirs := getattr(declared, name))
    }

with_declared_knobs(declared: ToolSpec) -> ToolSpec

Return this spec with declared's knob values overlaid.

The built-in registration is the AUTHORITY for these fields and the manifest entry is a cache of it, so the declaration wins — which is what lets an already-written manifest self-heal on the next load with no operator step and no regeneration. Everything else on the manifest entry (identity, schema, enablement, MCP wiring) is untouched.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
173
174
175
176
177
178
179
180
181
182
183
184
def with_declared_knobs(self, declared: ToolSpec) -> ToolSpec:
    """Return this spec with *declared*'s knob values overlaid.

    The built-in registration is the AUTHORITY for these fields and the
    manifest entry is a cache of it, so the declaration wins — which is what
    lets an already-written manifest self-heal on the next load with no
    operator step and no regeneration. Everything else on the manifest entry
    (identity, schema, enablement, MCP wiring) is untouched.
    """
    return replace(
        self, **{name: getattr(declared, name) for name in self.DECLARED_KNOBS}
    )

capability_mode_admits(capability_mode: str, tier: str | None) -> bool

True if a tool of privilege tier survives capability_mode.

The ONE home of the mode→tier law, shared by the registry filter (:func:filter_specs) AND the session-tool build (SessionToolRegistry.ids_for) so the two enforcement surfaces can never drift (the two-surface trap). all — and any unrecognised mode, since the authoritative validation is the Literal at the SpawnAgentTask boundary — admits everything (the gate is skipped). An undeclared tier (None) is admitted ONLY by all: safe-deny under any restrictive mode.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
def capability_mode_admits(capability_mode: str, tier: str | None) -> bool:
    """True if a tool of privilege *tier* survives *capability_mode*.

    The ONE home of the mode→tier law, shared by the registry filter
    (:func:`filter_specs`) AND the session-tool build
    (``SessionToolRegistry.ids_for``) so the two enforcement surfaces can never
    drift (the two-surface trap). ``all`` — and any unrecognised mode, since
    the authoritative validation is the ``Literal`` at the ``SpawnAgentTask``
    boundary — admits everything (the gate is skipped). An undeclared tier
    (``None``) is admitted ONLY by ``all``: safe-deny under any restrictive mode.
    """
    allowed_tiers = _CAPABILITY_MODE_TIERS.get(capability_mode)
    if allowed_tiers is None:
        return True
    return tier in allowed_tiers

classify_tool_scope(spec: ToolSpec, *, global_servers: set[str], plugin_servers: set[str]) -> str

Classify a tool spec into its deployment scope.

The four real scope categories:

  • builtin — a core built-in Python tool (spec.kind != "mcp"), not an MCP server at all.
  • system — an MCP tool whose server is configured in the shared, deployed-instance mcp.json (global_servers).
  • plugin — an MCP tool contributed by an installed plugin (plugin_servers).
  • project — an MCP tool whose server is configured only in the current project's local MCP config (neither of the above).

A genuine user tier — a personal ~/.mewbo config distinct from the deployed $MEWBO_HOME system instance — is deliberately NOT implemented: no current infra distinguishes the two in a way worth surfacing yet.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
def classify_tool_scope(
    spec: ToolSpec, *, global_servers: set[str], plugin_servers: set[str]
) -> str:
    """Classify a tool spec into its deployment scope.

    The four real scope categories:

    - ``builtin`` — a core built-in Python tool (``spec.kind != "mcp"``),
      not an MCP server at all.
    - ``system`` — an MCP tool whose server is configured in the shared,
      deployed-instance ``mcp.json`` (``global_servers``).
    - ``plugin`` — an MCP tool contributed by an installed plugin
      (``plugin_servers``).
    - ``project`` — an MCP tool whose server is configured only in the
      current project's local MCP config (neither of the above).

    A genuine ``user`` tier — a personal ``~/.mewbo`` config distinct from
    the deployed ``$MEWBO_HOME`` system instance — is deliberately NOT
    implemented: no current infra distinguishes the two in a way worth
    surfacing yet.
    """
    if spec.kind != "mcp":
        return "builtin"
    server = spec.metadata.get("server", "")
    if server in plugin_servers:
        return "plugin"
    if server in global_servers:
        return "system"
    return "project"

filter_specs(specs: list[ToolSpec], *, allowed: list[str] | None = None, denied: list[str] | None = None, capability_mode: str = 'all') -> list[ToolSpec]

Filter tool specs by allowlist, capability mode, and/or denylist.

allowed is THREE-STATE and tested with is None, never truthiness: None is unrestricted (no allowlist gate), [] grants NOTHING, and a non-empty list grants exactly those ids. Collapsing [] into None would turn "this principal gets no tools" into "this principal gets every tool" — the fail-open direction, on the gate that binds an agent's whole tool surface.

When the gate applies, only specs whose tool_id is in the list are kept — EXCEPT always_load specs (the tool_search tool), which are exempt from the allowlist gate so a scoped sub-agent never loses the means to fetch its deferred MCP tools.

capability_mode is a coarse privilege pre-filter applied AFTER the allowlist and layered UNDER it: it can only remove more tools, never resurrect one the allowlist dropped. Only specs whose :meth:ToolSpec.capability_tier is admitted by the mode survive — see CapabilityMode for the tier law. always_load is exempt here too (harmless discovery), mirroring the allowlist gate; "all" (the default) and any unrecognised mode skip the gate entirely.

Then any spec whose tool_id appears in denied (merged with the config agent.default_denied_tools) is removed. Deny always takes precedence over allow AND capability_mode — an explicit deny removes even an always_load tool.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
def filter_specs(
    specs: list[ToolSpec],
    *,
    allowed: list[str] | None = None,
    denied: list[str] | None = None,
    capability_mode: str = "all",
) -> list[ToolSpec]:
    """Filter tool specs by allowlist, capability mode, and/or denylist.

    *allowed* is THREE-STATE and tested with ``is None``, never truthiness:
    ``None`` is unrestricted (no allowlist gate), ``[]`` grants NOTHING, and a
    non-empty list grants exactly those ids. Collapsing ``[]`` into ``None``
    would turn "this principal gets no tools" into "this principal gets every
    tool" — the fail-open direction, on the gate that binds an agent's whole
    tool surface.

    When the gate applies, only specs whose ``tool_id`` is in the list are
    kept — EXCEPT ``always_load`` specs (the ``tool_search`` tool),
    which are exempt from the allowlist gate so a scoped sub-agent never
    loses the means to fetch its deferred MCP tools.

    *capability_mode* is a coarse privilege pre-filter
    applied AFTER the allowlist and layered UNDER it: it can only remove more
    tools, never resurrect one the allowlist dropped. Only specs whose
    :meth:`ToolSpec.capability_tier` is admitted by the mode survive — see
    ``CapabilityMode`` for the tier law. ``always_load`` is exempt here too
    (harmless discovery), mirroring the allowlist gate; ``"all"`` (the default)
    and any unrecognised mode skip the gate entirely.

    Then any spec whose ``tool_id`` appears in *denied* (merged with the config
    ``agent.default_denied_tools``) is removed.  Deny always takes precedence
    over allow AND capability_mode — an explicit deny removes even an
    ``always_load`` tool.
    """
    if allowed is not None:
        allowed_set = set(allowed)
        specs = [s for s in specs if s.tool_id in allowed_set or is_always_load(s)]

    if _CAPABILITY_MODE_TIERS.get(capability_mode) is not None:
        specs = [
            s
            for s in specs
            if is_always_load(s)
            or capability_mode_admits(capability_mode, s.capability_tier())
        ]

    denied_set: set[str] = set(denied or [])
    config_denied_raw = get_config_value("agent", "default_denied_tools", default=[])
    if isinstance(config_denied_raw, str):
        config_denied_raw = [s.strip() for s in config_denied_raw.split(",") if s.strip()]
    denied_set |= set(config_denied_raw or [])

    if denied_set:
        specs = [s for s in specs if s.tool_id not in denied_set]

    return specs

get_or_build_registry(*, cwd: str | None = None, extra_mcp_servers: dict[str, dict] | None = None, trust_cwd: bool = True) -> ToolRegistry

Return a cached ToolRegistry for these inputs, building once per scope.

The single seam Orchestrator uses to avoid rebuilding the registry on every run in a session. Falls through to :func:load_registry on a cache miss. See :class:ToolRegistryCache for the keying contract.

Pass trust_cwd=False when cwd holds content this deployment did not author. A caller that does not know can leave it alone and register the directory with :data:mewbo_core.config.register_untrusted_cwd instead — that is read here too, and overrides an affirmative argument.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
def get_or_build_registry(
    *,
    cwd: str | None = None,
    extra_mcp_servers: dict[str, dict] | None = None,
    trust_cwd: bool = True,
) -> ToolRegistry:
    """Return a cached ``ToolRegistry`` for these inputs, building once per scope.

    The single seam ``Orchestrator`` uses to avoid rebuilding the registry on
    every run in a session. Falls through to :func:`load_registry` on a cache
    miss. See :class:`ToolRegistryCache` for the keying contract.

    Pass ``trust_cwd=False`` when *cwd* holds content this deployment did not
    author. A caller that does not know can leave it alone and register the
    directory with :data:`mewbo_core.config.register_untrusted_cwd` instead —
    that is read here too, and overrides an affirmative argument.
    """
    return _REGISTRY_CACHE.get_or_build(
        cwd=cwd, extra_mcp_servers=extra_mcp_servers, trust_cwd=trust_cwd
    )

is_always_load(spec: ToolSpec) -> bool

Return True if the tool's full schema must always be in the bound list.

Marked via metadata.always_load=True. An opt-out from deferral — used by tools that the model needs immediately (the search tool itself, or any tool whose absence would block the model from making progress).

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
406
407
408
409
410
411
412
413
414
def is_always_load(spec: ToolSpec) -> bool:
    """Return True if the tool's full schema must always be in the bound list.

    Marked via ``metadata.always_load=True``. An opt-out from deferral —
    used by tools that the model needs immediately
    (the search tool itself, or any tool whose absence would block the
    model from making progress).
    """
    return bool(spec.metadata.get("always_load"))

is_deferred(spec: ToolSpec) -> bool

Return True if the tool's schema should be omitted from the initial bind.

Deferred tools surface as names only via <available-deferred-tools> — the model fetches their schemas on demand via tool_search. The deferral rule: always_load wins, the search tool itself never defers, all MCP tools defer, and other tools opt-in via metadata.deferred=True.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
def is_deferred(spec: ToolSpec) -> bool:
    """Return True if the tool's schema should be omitted from the initial bind.

    Deferred tools surface as names only via ``<available-deferred-tools>`` —
    the model fetches their schemas on demand via ``tool_search``. The
    deferral rule: ``always_load`` wins, the search tool
    itself never defers, all MCP tools defer, and other tools opt-in via
    ``metadata.deferred=True``.
    """
    if is_always_load(spec):
        return False
    if spec.tool_id == TOOL_SEARCH_TOOL_ID:
        return False
    if spec.kind == "mcp":
        return True
    return bool(spec.metadata.get("deferred"))

load_registry(manifest_path: str | None = None, *, cwd: str | None = None, extra_mcp_servers: dict[str, dict] | None = None, trust_cwd: bool = True) -> ToolRegistry

Load tool registry, auto-discovering MCP tools when configured.

trust_cwd is the caller's explicit statement about cwd: False means the directory holds content this deployment did not author, so neither its own .mcp.json nor any beneath it may name a server. The decision is carried onto every MCPToolRunner built here, because a runner re-resolves the config at invocation time — covering the boundary at build time alone would leak at call time.

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
def load_registry(
    manifest_path: str | None = None,
    *,
    cwd: str | None = None,
    extra_mcp_servers: dict[str, dict] | None = None,
    trust_cwd: bool = True,
) -> ToolRegistry:
    """Load tool registry, auto-discovering MCP tools when configured.

    *trust_cwd* is the caller's explicit statement about *cwd*: ``False`` means
    the directory holds content this deployment did not author, so neither its
    own ``.mcp.json`` nor any beneath it may name a server. The decision is
    carried onto every ``MCPToolRunner`` built here, because a runner
    re-resolves the config at invocation time — covering the boundary at build
    time alone would leak at call time.
    """
    from mewbo_core.config import is_untrusted_cwd

    # A directory registered as untrusted overrides an affirmative argument: the
    # caller that CREATED it knows more than the caller that passed it along.
    trust_cwd = trust_cwd and not is_untrusted_cwd(cwd)
    if manifest_path is None:
        mcp_config_path = get_mcp_config_path()
        # Check for CWD .mcp.json even when no global config exists
        has_cwd_mcp = False
        if cwd and trust_cwd:
            from pathlib import Path as _Path

            has_cwd_mcp = (_Path(cwd) / ".mcp.json").is_file()
        # Also check for subtree .mcp.json files
        has_subtree_mcp = False
        if cwd and trust_cwd and not has_cwd_mcp:
            from mewbo_core.config import _discover_subtree_mcp_json

            has_subtree_mcp = bool(_discover_subtree_mcp_json(cwd))
        if (
            (mcp_config_path and os.path.exists(mcp_config_path))
            or has_cwd_mcp
            or has_subtree_mcp
            or extra_mcp_servers
        ):
            manifest_path = _ensure_auto_manifest(
                mcp_config_path or "",
                cwd=cwd,
                extra_mcp_servers=extra_mcp_servers,
                trust_cwd=trust_cwd,
            )

    if not manifest_path:
        return _default_registry()

    manifest_path = os.path.abspath(manifest_path)
    if not os.path.exists(manifest_path):
        logging.warning("Tool manifest not found: {}", manifest_path)
        return _default_registry()

    try:
        with open(manifest_path, encoding="utf-8") as handle:
            manifest = json.load(handle)
    except Exception as exc:  # pragma: no cover - defensive
        logging.error("Failed to load tool manifest: {}", exc)
        return _default_registry()

    registry = ToolRegistry()
    # Built once, up front, and reused for BOTH the declared-limit overlay below
    # and the merge of un-manifested built-ins further down — the same object
    # that used to be built only for the merge, so this costs no extra probe.
    builtin_registry = _default_registry()
    declared_by_id = {
        spec.tool_id: spec for spec in builtin_registry.list_specs(include_disabled=True)
    }
    for tool in manifest.get("tools", []):
        kind = tool.get("kind", "local")
        prompt_path = tool.get("prompt")
        if kind == "local":
            module_path = tool.get("module")
            class_name = tool.get("class")
            if not module_path or not class_name:
                logging.warning("Skipping tool with missing module/class: {}", tool)
                continue
            factory = _import_factory(module_path, class_name)
        else:
            mcp_module = _load_mcp_support()
            if mcp_module is None:
                logging.warning(
                    "Skipping MCP tool because MCP support is not installed: {}",
                    tool,
                )
                continue
            MCPToolRunner = mcp_module.MCPToolRunner

            server_name = tool.get("server")
            tool_name = tool.get("tool")
            if not server_name or not tool_name:
                logging.warning("Skipping MCP tool with missing server/tool: {}", tool)
                continue

            def _mcp_factory(
                server_name: str = server_name,
                tool_name: str = tool_name,
                _cwd: str | None = cwd,
                _trust_cwd: bool = trust_cwd,
            ) -> ToolRunner:
                return MCPToolRunner(
                    server_name=server_name,
                    tool_name=tool_name,
                    cwd=_cwd,
                    trust_cwd=_trust_cwd,
                )

            factory = _mcp_factory

        spec = ToolSpec(
            tool_id=tool.get("tool_id", ""),
            name=tool.get("name", tool.get("tool_id", "")),
            description=tool.get("description", ""),
            factory=factory,
            enabled=tool.get("enabled", True),
            kind=kind,
            prompt_path=prompt_path,
            read_only=bool(tool.get("read_only", False)),
            capability=tool.get("capability"),
            poll=bool(tool.get("poll", False)),
            poll_when_args=tuple(tool.get("poll_when_args") or ()),
            # A field this rebuild forgets is not defaulted deliberately — it is
            # LOST, and it takes the reason someone declared it with it. These
            # four were dropped: every deployment that reaches the registry
            # through a manifest (any deployment with an MCP config, which is
            # the normal one) silently ran all 159 tools at the 2000-char
            # default, so the caps declared on shell and on the file reader were
            # dead code exactly where they mattered. The symptom was a model
            # given the middle of a file removed, arguing with a tool that had
            # told it there was no such limit.
            max_result_chars=int(
                tool.get("max_result_chars", ToolSpec.max_result_chars)
            ),
            timeout=float(tool.get("timeout", ToolSpec.timeout)),
            concurrency_safe=bool(
                tool.get("concurrency_safe", ToolSpec.concurrency_safe)
            ),
            interrupt_behavior=str(
                tool.get("interrupt_behavior", ToolSpec.interrupt_behavior)
            ),
            metadata={
                key: value
                for key, value in tool.items()
                if key
                not in {
                    "tool_id",
                    "name",
                    "description",
                    "module",
                    "class",
                    "enabled",
                    "kind",
                    "prompt",
                    "read_only",
                    "capability",
                    "poll",
                    "poll_when_args",
                    "max_result_chars",
                    "timeout",
                    "concurrency_safe",
                    "interrupt_behavior",
                }
            },
        )
        if not spec.tool_id:
            logging.warning("Skipping tool with empty tool_id: {}", tool)
            continue

        # THE PARITY GUARD, and the repair, in one comparison. A manifest is a
        # CACHE of the built-in declarations and can outlive them: a file written
        # before a field existed carries no key for it, so the reader above
        # defaults it and the declaration dies exactly where it is needed. That
        # is not hypothetical — a live deployment ran all 159 tools at the
        # 2000-char default while `read_file` declared 200_000, because its
        # cached manifest predated the field and the config hash still matched,
        # so no regeneration was ever triggered.
        #
        # Overlaying at LOAD rather than bumping a cache-invalidation hash is
        # deliberate: an already-written manifest must self-heal with no operator
        # step, and a hash bump repairs nothing until something happens to
        # rewrite the file. Loud, because a silent repair leaves the stale file
        # in place and the next reader re-derives the same surprise.
        declared = declared_by_id.get(spec.tool_id)
        if declared is not None:
            mismatches = spec.knob_mismatches(declared)
            if mismatches:
                logging.warning(
                    "Tool manifest is stale for {}: {} — using the built-in "
                    "declaration. Regenerate the manifest to silence this.",
                    spec.tool_id,
                    ", ".join(
                        f"{name} manifest={mine!r} declared={theirs!r}"
                        for name, (mine, theirs) in sorted(mismatches.items())
                    ),
                )
                spec = spec.with_declared_knobs(declared)
        registry.register(spec)

    if not registry.list_specs(include_disabled=True):
        return builtin_registry

    existing_ids = {spec.tool_id for spec in registry.list_specs(include_disabled=True)}
    for spec in builtin_registry.list_specs(include_disabled=True):
        if spec.tool_id == TOOL_SEARCH_TOOL_ID:
            # Skip — re-registered below so its factory binds to ``registry``,
            # not the throwaway ``builtin_registry`` instance.
            continue
        if spec.tool_id in existing_ids:
            continue
        registry.register(spec)
        existing_ids.add(spec.tool_id)

    # Bind tool_search to the final, merged registry so it can search all specs.
    _register_tool_search(registry)

    set_available_tools([spec.tool_id for spec in registry.list_specs()])
    return registry

mcp_tool_id(server_name: str, tool_name: str) -> str

Public alias for the canonical mcp_<server>_<tool> id convention.

Anything that must NAME an executable MCP tool outside the registry (the SCG route projection emitting a probe's allowed_tools, allowlist builders, trace labels) must derive the id HERE — never by string-mangling a source_key. A graph source_key (<source>#<Capability>) is a graph address, not a tool id; passing it to allowed_tools silently grants nothing (the run-c52e9597 probe failure).

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
889
890
891
892
893
894
895
896
897
898
899
def mcp_tool_id(server_name: str, tool_name: str) -> str:
    """Public alias for the canonical ``mcp_<server>_<tool>`` id convention.

    Anything that must NAME an executable MCP tool outside the registry (the
    SCG route projection emitting a probe's ``allowed_tools``, allowlist
    builders, trace labels) must derive the id HERE — never by string-mangling
    a ``source_key``. A graph ``source_key`` (``<source>#<Capability>``) is a
    graph address, not a tool id; passing it to ``allowed_tools`` silently
    grants nothing (the run-c52e9597 probe failure).
    """
    return _sanitize_tool_id(server_name, tool_name)

reset_registry_cache() -> None

Drop all cached registries (tests; an explicit /mcp refresh).

Source code in packages/mewbo_core/src/mewbo_core/tooling/tool_registry.py
1604
1605
1606
def reset_registry_cache() -> None:
    """Drop all cached registries (tests; an explicit ``/mcp`` refresh)."""
    _REGISTRY_CACHE.clear()

mewbo_core.classes

Core data models and tool abstractions for Mewbo orchestration.

AbstractTool

Bases: ABC

Base tool with shared initialization helpers.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
class AbstractTool(abc.ABC):
    """Base tool with shared initialization helpers."""

    def __init__(
        self,
        name: str,
        description: str,
        model_name: str | None = None,
        use_llm: bool = True,
    ) -> None:
        """Initialize tool configuration."""
        tool_model = get_config_value("llm", "tool_model")
        default_model = get_config_value("llm", "default_model", default="gpt-5.2")
        self.model_name = cast(
            str,
            model_name or tool_model or default_model,
        )
        self.name = name
        self.description = description
        self.use_llm = use_llm
        self._id = f"{name.lower().replace(' ', '_')}_tool"
        session_id = f"{self._id}-tool-id-{get_unique_timestamp()}"
        logging.info(f"Tool created <name={name}; session_id={session_id};>")
        self.langfuse_handler = build_langfuse_handler(
            user_id=f"mewbo-{name}",
            session_id=session_id,
            trace_name=f"mewbo-{self._id}",
            version=get_version(),
            release=get_config_value("runtime", "envmode", default="Not Specified"),
        )
        self.model = None
        if self.use_llm:
            # Imported at the call site, not at module top: this module sits
            # below `llm` in the layering everywhere else, and a module-top
            # import here is the edge that pins the whole LLM stack to the
            # package root.
            #
            # Safe HERE specifically, and the argument is per-site rather than
            # general (a lazy import is only safe when the imported module's
            # import-time side effects cannot re-enter the caller). `llm` pulls
            # in LiteLLM, whose module body reads a `.env` and populates
            # `os.environ`; this line runs long after import, and the config
            # this constructor depends on is already built by the time control
            # reaches it — `get_config_value` above and `get_logger` at module
            # scope both force it.
            from mewbo_core.llm.llm import build_chat_model

            self.model = build_chat_model(model_name=self.model_name)
        root_cache_dir = get_config_value("runtime", "cache_dir", default=".cache")
        if not root_cache_dir:
            raise ValueError("runtime.cache_dir is not set.")
        self.cache_dir = os.path.abspath(os.path.join(str(root_cache_dir), self._id))
        logging.debug("{} cache directory is {}.", self._id, self.cache_dir)

    def _save_json(self, data: object, filename: str) -> None:
        """Persist JSON data under the cache directory."""
        if not os.path.exists(self.cache_dir):
            os.makedirs(self.cache_dir)
        filename = os.path.join(self.cache_dir, filename)
        with open(filename, "w", encoding="utf-8") as f:
            json.dump(data, f, indent=4)
        logging.info(f"Data saved to {filename}.")

    def _load_rag_json(self, filename: str) -> list[Document]:
        """Load JSON content as documents."""
        logging.debug("RAG directory is {}.", self.cache_dir)
        logging.info(f"Loading `{filename}` as JSON.")
        filename = os.path.join(self.cache_dir, filename)
        filename = os.path.abspath(filename)
        loader = JSONLoader(file_path=filename, jq_schema=".", text_content=False)
        data = loader.load()
        return data

    def _load_rag_documents(self, filenames: list[str]) -> list[Document]:
        """Load and concatenate multiple JSON files."""
        rag_documents: list[Document] = []
        for rag_file in filenames:
            data = self._load_rag_json(rag_file)
            rag_documents.extend(data)
        return rag_documents

    def set_state(self, action_step: ActionStep | None = None) -> MockSpeaker:
        """Perform a state-changing action."""
        MockSpeaker = get_mock_speaker()
        return MockSpeaker(content="Not implemented yet.")

    def get_state(self, action_step: ActionStep | None = None) -> MockSpeaker:
        """Perform a read-only action."""
        MockSpeaker = get_mock_speaker()
        return MockSpeaker(content="Not implemented yet.")

    def run(self, action_step: ActionStep) -> MockSpeaker:
        """Execute the action based on the operation."""
        if action_step.operation == "set":
            return self.set_state(action_step)
        if action_step.operation == "get":
            return self.get_state(action_step)
        raise ValueError(f"Invalid operation: {action_step.operation}")

__init__(name: str, description: str, model_name: str | None = None, use_llm: bool = True) -> None

Initialize tool configuration.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
def __init__(
    self,
    name: str,
    description: str,
    model_name: str | None = None,
    use_llm: bool = True,
) -> None:
    """Initialize tool configuration."""
    tool_model = get_config_value("llm", "tool_model")
    default_model = get_config_value("llm", "default_model", default="gpt-5.2")
    self.model_name = cast(
        str,
        model_name or tool_model or default_model,
    )
    self.name = name
    self.description = description
    self.use_llm = use_llm
    self._id = f"{name.lower().replace(' ', '_')}_tool"
    session_id = f"{self._id}-tool-id-{get_unique_timestamp()}"
    logging.info(f"Tool created <name={name}; session_id={session_id};>")
    self.langfuse_handler = build_langfuse_handler(
        user_id=f"mewbo-{name}",
        session_id=session_id,
        trace_name=f"mewbo-{self._id}",
        version=get_version(),
        release=get_config_value("runtime", "envmode", default="Not Specified"),
    )
    self.model = None
    if self.use_llm:
        # Imported at the call site, not at module top: this module sits
        # below `llm` in the layering everywhere else, and a module-top
        # import here is the edge that pins the whole LLM stack to the
        # package root.
        #
        # Safe HERE specifically, and the argument is per-site rather than
        # general (a lazy import is only safe when the imported module's
        # import-time side effects cannot re-enter the caller). `llm` pulls
        # in LiteLLM, whose module body reads a `.env` and populates
        # `os.environ`; this line runs long after import, and the config
        # this constructor depends on is already built by the time control
        # reaches it — `get_config_value` above and `get_logger` at module
        # scope both force it.
        from mewbo_core.llm.llm import build_chat_model

        self.model = build_chat_model(model_name=self.model_name)
    root_cache_dir = get_config_value("runtime", "cache_dir", default=".cache")
    if not root_cache_dir:
        raise ValueError("runtime.cache_dir is not set.")
    self.cache_dir = os.path.abspath(os.path.join(str(root_cache_dir), self._id))
    logging.debug("{} cache directory is {}.", self._id, self.cache_dir)

get_state(action_step: ActionStep | None = None) -> MockSpeaker

Perform a read-only action.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
375
376
377
378
def get_state(self, action_step: ActionStep | None = None) -> MockSpeaker:
    """Perform a read-only action."""
    MockSpeaker = get_mock_speaker()
    return MockSpeaker(content="Not implemented yet.")

run(action_step: ActionStep) -> MockSpeaker

Execute the action based on the operation.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
380
381
382
383
384
385
386
def run(self, action_step: ActionStep) -> MockSpeaker:
    """Execute the action based on the operation."""
    if action_step.operation == "set":
        return self.set_state(action_step)
    if action_step.operation == "get":
        return self.get_state(action_step)
    raise ValueError(f"Invalid operation: {action_step.operation}")

set_state(action_step: ActionStep | None = None) -> MockSpeaker

Perform a state-changing action.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
370
371
372
373
def set_state(self, action_step: ActionStep | None = None) -> MockSpeaker:
    """Perform a state-changing action."""
    MockSpeaker = get_mock_speaker()
    return MockSpeaker(content="Not implemented yet.")

ActionStep

Bases: BaseModel

Action step with validation metadata.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
class ActionStep(BaseModel):
    """Action step with validation metadata."""

    title: str | None = Field(
        default=None,
        description="Short header summarizing the task for this step.",
    )
    objective: str | None = Field(
        default=None,
        description="Brief objective explaining why this step is needed.",
    )
    execution_checklist: list[str] = Field(
        default_factory=list,
        description="Short checklist of execution details for this step.",
    )
    expected_output: str | None = Field(
        default=None,
        description="Optional description of what success looks like.",
    )
    tool_id: str = Field(
        description=(
            "Specify the tool_id that should execute the action. "
            "Use only tool IDs listed under Available tools."
        )
    )
    operation: str = Field(description="Specify the execution type (get/set or execute).")
    tool_input: ToolInput = Field(
        description=(
            "Provide details for the action. If 'task', specify the task to perform. "
            "If 'talk', include the message to speak to the user."
        )
    )
    result: object | None = Field(
        alias="_result",
        default=None,
        description="Private field to persist the action status and other data.",
    )

    model_config = ConfigDict(populate_by_name=True, extra="forbid")

OrchestrationState

Bases: BaseModel

State for the orchestration loop.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
class OrchestrationState(BaseModel):
    """State for the orchestration loop."""

    goal: str
    session_id: str | None = None
    plan: list[PlanStep] = Field(default_factory=list)
    tool_results: list[str] = Field(default_factory=list)
    open_questions: list[str] = Field(default_factory=list)
    done: bool = False
    done_reason: str | None = None
    summary: str | None = None
    plan_approved: bool = False
    plan_path: str | None = None
    # Verifier-gated completion outcome. ``None`` = the gate never ran (no
    # spec, or inactive); ``True``/``False`` = a
    # ground-truth check passed/exhausted its retries. ``verify_attempts``
    # counts how many verifier runs this task drove (0 when the gate was
    # never active). Both survive to the ``AgentResult`` so a spawner sees an
    # honest done-claim rather than an invisible null.
    verified: bool | None = None
    verify_attempts: int = 0
    # The tool-envelope error code of a blocked-class condition the run never
    # recovered from — credentials, reachability, permission, quota. ``None``
    # (the overwhelming majority) means no such condition stood at the end.
    # Deliberately NOT folded into ``done_reason``: that vocabulary is a wire
    # contract every client already switches on, whereas this is an additive
    # fact the status layer reads to tell a run that FAILED from one that is
    # BLOCKED on something a user can go fix. Set by the loop, projected onto
    # the completion event, mapped to a status downstream.
    blocked_code: str | None = None

    def terminal_status(self) -> Literal["completed", "failed", "cancelled"]:
        """Project this settled state onto the terminal a spawner is told.

        ``done`` answers "did the loop stop", never "did the task succeed", and
        reading it as the latter is what let a doom-looped or budget-spent child
        report ``completed`` to its parent. A run that stopped short is
        ``failed``; only a genuine natural completion whose ground-truth check
        did not fail is ``completed``.

        ``verified is False`` is checked in its own right rather than trusted to
        the reason: it is the authoritative record that a ground-truth check ran
        and did not pass, and a claim contradicted by ground truth must not
        depend on a second field spelling it the same way. It is checked BEFORE
        cancellation because a contradicted claim is a substantive failure,
        whereas a stop is only a stop.

        A cancelled run is neither: nobody claimed the goal was reached and
        nothing contradicted a claim, so folding it into ``failed`` would report
        an error that never happened while ``completed`` reports a success that
        never happened. It gets the ``AgentStatus`` member that means what
        occurred. A ``completed``/``failed`` projection cannot express the one
        terminal a user causes directly, which is why there is a third member.

        The returned values are members of the hypervisor's ``AgentStatus``
        vocabulary — this is a NARROWING of that authority, not a vocabulary of
        its own, and it is exactly ``SettledStatus``. A halt reports as
        ``failed`` because ``AgentStatus`` has no "stopped short" arm; the
        precise reason is not lost — it rides the ``stop`` event's ``detail``
        and the attestation's ``done_reason``.

        Lives on the model rather than on whichever service happens to settle a
        run: the projection reads nothing but this state's own fields, so a
        second copy at another call site could only ever drift from this one.
        """
        if not self.done:
            return "failed"
        if self.verified is False:
            return "failed"
        reason = self.done_reason
        if isinstance(reason, str) and reason in CANCELLED_DONE_REASONS:
            return "cancelled"
        if isinstance(reason, str) and reason in UNACHIEVED_DONE_REASONS:
            return "failed"
        return "completed"

terminal_status() -> Literal['completed', 'failed', 'cancelled']

Project this settled state onto the terminal a spawner is told.

done answers "did the loop stop", never "did the task succeed", and reading it as the latter is what let a doom-looped or budget-spent child report completed to its parent. A run that stopped short is failed; only a genuine natural completion whose ground-truth check did not fail is completed.

verified is False is checked in its own right rather than trusted to the reason: it is the authoritative record that a ground-truth check ran and did not pass, and a claim contradicted by ground truth must not depend on a second field spelling it the same way. It is checked BEFORE cancellation because a contradicted claim is a substantive failure, whereas a stop is only a stop.

A cancelled run is neither: nobody claimed the goal was reached and nothing contradicted a claim, so folding it into failed would report an error that never happened while completed reports a success that never happened. It gets the AgentStatus member that means what occurred. A completed/failed projection cannot express the one terminal a user causes directly, which is why there is a third member.

The returned values are members of the hypervisor's AgentStatus vocabulary — this is a NARROWING of that authority, not a vocabulary of its own, and it is exactly SettledStatus. A halt reports as failed because AgentStatus has no "stopped short" arm; the precise reason is not lost — it rides the stop event's detail and the attestation's done_reason.

Lives on the model rather than on whichever service happens to settle a run: the projection reads nothing but this state's own fields, so a second copy at another call site could only ever drift from this one.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
def terminal_status(self) -> Literal["completed", "failed", "cancelled"]:
    """Project this settled state onto the terminal a spawner is told.

    ``done`` answers "did the loop stop", never "did the task succeed", and
    reading it as the latter is what let a doom-looped or budget-spent child
    report ``completed`` to its parent. A run that stopped short is
    ``failed``; only a genuine natural completion whose ground-truth check
    did not fail is ``completed``.

    ``verified is False`` is checked in its own right rather than trusted to
    the reason: it is the authoritative record that a ground-truth check ran
    and did not pass, and a claim contradicted by ground truth must not
    depend on a second field spelling it the same way. It is checked BEFORE
    cancellation because a contradicted claim is a substantive failure,
    whereas a stop is only a stop.

    A cancelled run is neither: nobody claimed the goal was reached and
    nothing contradicted a claim, so folding it into ``failed`` would report
    an error that never happened while ``completed`` reports a success that
    never happened. It gets the ``AgentStatus`` member that means what
    occurred. A ``completed``/``failed`` projection cannot express the one
    terminal a user causes directly, which is why there is a third member.

    The returned values are members of the hypervisor's ``AgentStatus``
    vocabulary — this is a NARROWING of that authority, not a vocabulary of
    its own, and it is exactly ``SettledStatus``. A halt reports as
    ``failed`` because ``AgentStatus`` has no "stopped short" arm; the
    precise reason is not lost — it rides the ``stop`` event's ``detail``
    and the attestation's ``done_reason``.

    Lives on the model rather than on whichever service happens to settle a
    run: the projection reads nothing but this state's own fields, so a
    second copy at another call site could only ever drift from this one.
    """
    if not self.done:
        return "failed"
    if self.verified is False:
        return "failed"
    reason = self.done_reason
    if isinstance(reason, str) and reason in CANCELLED_DONE_REASONS:
        return "cancelled"
    if isinstance(reason, str) and reason in UNACHIEVED_DONE_REASONS:
        return "failed"
    return "completed"

Plan

Bases: BaseModel

Plan with human-readable steps.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
116
117
118
119
120
121
122
123
124
class Plan(BaseModel):
    """Plan with human-readable steps."""

    human_message: str | None = Field(
        alias="_human_message",
        default=None,
        description="Human message associated with the plan.",
    )
    steps: list[PlanStep] = Field(default_factory=list)

PlanStep

Bases: BaseModel

High-level plan step produced by the planner.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
109
110
111
112
113
class PlanStep(BaseModel):
    """High-level plan step produced by the planner."""

    title: str = Field(description="Short title for the step.")
    description: str = Field(description="One-paragraph description of the step.")

TaskQueue

Bases: BaseModel

Queue of executed tool steps and results.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
class TaskQueue(BaseModel):
    """Queue of executed tool steps and results."""

    human_message: str | None = Field(
        alias="_human_message",
        default=None,
        description="Human message associated with the task queue.",
    )
    plan_steps: list[PlanStep] = Field(default_factory=list)
    action_steps: list[ActionStep] = Field(default_factory=list)
    task_result: str | None = Field(
        alias="_task_result", default=None, description="Store the result for the entire task queue"
    )
    last_error: str | None = Field(
        alias="_last_error",
        default=None,
        description="Short description of the most recent tool failure.",
    )

    @field_validator("action_steps")
    @classmethod
    def validate_actions(cls, field: list[ActionStep]) -> list[ActionStep]:
        """Normalize and validate action steps."""
        for action in field:
            action.tool_id = action.tool_id.lower()
            action.operation = action.operation.lower()
            error_msg_list = []

            if action.tool_id not in AVAILABLE_TOOLS and action.tool_id not in INTERNAL_TOOL_IDS:
                error_msg_list.append(f"`{action.tool_id}` is not a valid Assistant tool.")

            if action.operation not in ["get", "set", "execute"]:
                error_msg = f"`{action.operation}` is not a valid operation."
                error_msg_list.append(error_msg)

            if action.tool_input is None:
                error_msg_list.append("Tool input cannot be None.")

            if error_msg_list:
                for msg in error_msg_list:
                    logging.error(msg)

        return field

validate_actions(field: list[ActionStep]) -> list[ActionStep] classmethod

Normalize and validate action steps.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
@field_validator("action_steps")
@classmethod
def validate_actions(cls, field: list[ActionStep]) -> list[ActionStep]:
    """Normalize and validate action steps."""
    for action in field:
        action.tool_id = action.tool_id.lower()
        action.operation = action.operation.lower()
        error_msg_list = []

        if action.tool_id not in AVAILABLE_TOOLS and action.tool_id not in INTERNAL_TOOL_IDS:
            error_msg_list.append(f"`{action.tool_id}` is not a valid Assistant tool.")

        if action.operation not in ["get", "set", "execute"]:
            error_msg = f"`{action.operation}` is not a valid operation."
            error_msg_list.append(error_msg)

        if action.tool_input is None:
            error_msg_list.append("Tool input cannot be None.")

        if error_msg_list:
            for msg in error_msg_list:
                logging.error(msg)

    return field

ToolResult dataclass

Structured tool execution result.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
51
52
53
54
55
56
57
58
59
@dataclass
class ToolResult:
    """Structured tool execution result."""

    content: str
    success: bool = True
    error: str | None = None
    truncated: bool = False
    original_length: int | None = None

create_plan(step_data: list[dict[str, str]] | None = None, is_example: bool = True) -> Plan

Create a Plan from serialized step data.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
404
405
406
407
408
409
410
411
412
413
414
415
def create_plan(
    step_data: list[dict[str, str]] | None = None,
    is_example: bool = True,
) -> Plan:
    """Create a Plan from serialized step data."""
    if step_data is None:
        raise ValueError("Step data cannot be None.")
    steps = [PlanStep(**step) for step in step_data]
    plan = Plan(steps=steps)
    if is_example:
        del plan.human_message
    return plan

create_task_queue(action_data: list[ActionStepPayload] | None = None, is_example: bool = True) -> TaskQueue

Create a TaskQueue from serialized action data.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
389
390
391
392
393
394
395
396
397
398
399
400
401
def create_task_queue(
    action_data: list[ActionStepPayload] | None = None,
    is_example: bool = True,
) -> TaskQueue:
    """Create a TaskQueue from serialized action data."""
    if action_data is None:
        raise ValueError("Action data cannot be None.")

    action_steps = [ActionStep(**action) for action in action_data]
    task_queue = TaskQueue(action_steps=action_steps)
    if is_example:
        del task_queue.human_message
    return task_queue

get_task_master_examples(example_id: int = 0, available_tools: Sequence[str] | None = None) -> str

Return serialized example plan data.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
def get_task_master_examples(
    example_id: int = 0,
    available_tools: Sequence[str] | None = None,
) -> str:
    """Return serialized example plan data."""
    if available_tools is None:
        available_tools = AVAILABLE_TOOLS
    include_home_assistant = "home_assistant_tool" in available_tools
    if include_home_assistant:
        examples: list[list[dict[str, str]]] = [
            [
                {
                    "title": "Turn on strip lights",
                    "description": "Use Home Assistant to switch on the strip lights.",
                },
                {
                    "title": "Turn on heater",
                    "description": "Use Home Assistant to switch on the heater.",
                },
            ],
            [
                {
                    "title": "Check weather",
                    "description": "Use Home Assistant to retrieve today's weather details.",
                },
            ],
        ]
    else:
        examples = [[], []]
    if example_id not in range(0, len(examples)):
        raise ValueError(f"Invalid example ID: {example_id}")

    return create_plan(step_data=examples[example_id], is_example=True).model_dump_json()

set_available_tools(tool_ids: list[str]) -> None

Update available tool IDs for validation.

Source code in packages/mewbo_core/src/mewbo_core/classes.py
62
63
64
65
def set_available_tools(tool_ids: list[str]) -> None:
    """Update available tool IDs for validation."""
    global AVAILABLE_TOOLS
    AVAILABLE_TOOLS = tool_ids

mewbo_core.contracts.types

Shared type definitions for core components.

ActionPlanPayload

Bases: TypedDict

Payload describing an action plan.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
21
22
23
24
class ActionPlanPayload(TypedDict):
    """Payload describing an action plan."""

    steps: list[PlanStepPayload]

ActionStepPayload

Bases: TypedDict

Serialized tool call data sent to/from execution.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
27
28
29
30
31
32
33
34
35
36
class ActionStepPayload(TypedDict):
    """Serialized tool call data sent to/from execution."""

    tool_id: str
    operation: str
    tool_input: ToolInput
    title: NotRequired[str]
    objective: NotRequired[str]
    execution_checklist: NotRequired[list[str]]
    expected_output: NotRequired[str]

AgentMessagePayload

Bases: TypedDict

Payload describing an intermediate agent text message.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
161
162
163
164
165
166
class AgentMessagePayload(TypedDict):
    """Payload describing an intermediate agent text message."""

    text: str
    agent_id: str
    depth: int

AssistantPayload

Bases: TypedDict

Payload describing an assistant response.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
95
96
97
98
class AssistantPayload(TypedDict):
    """Payload describing an assistant response."""

    text: str

CompletionPayload

Bases: TypedDict

Payload describing overall completion state.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
class CompletionPayload(TypedDict):
    """Payload describing overall completion state."""

    done: bool
    done_reason: str | None
    task_result: str | None
    # Flat, bounded blurbs — Aura's error card and the CLI read these.
    # Written from ``RunError.brief()``, so never longer than 500
    # characters no matter what the provider stack emitted.
    error: NotRequired[str]
    last_error: NotRequired[str]
    # Additive structured failure record — a serialized ``RunError``
    # (``mewbo_core.contracts.run_error``), stored via ``model_dump(mode="json")`` so the
    # value is plain JSON types for Mongo and the wire. Present only alongside
    # ``error``; a client that only knows the flat keys is unaffected.
    error_detail: NotRequired[dict[str, Any]]
    # The last UNRECOVERED tool-envelope error code, and only when it is one a
    # user can act on: ``repo_access`` / ``network`` / ``forbidden`` /
    # ``quota_exceeded``. Projected from ``OrchestrationState.blocked_code``,
    # which the loop clears per tool_id as soon as that tool succeeds again.
    #
    # This is the sole input to the derived ``blocked`` status
    # (``session_runtime``): ``done_reason`` stays ``"completed"`` for a blocked
    # run at the loop layer, so without this key the fact that a run died
    # against a credential or a network path is unreachable from the record.
    # NotRequired, so every payload written before it existed still validates.
    blocked_code: NotRequired[str]

DeviceToolCallPayload

Bases: TypedDict

Payload for a client-declared device-tool invocation (device_tool_call).

Emitted when a session-bound ClientDeclaredTool (see client_tools.py) is invoked; delivered to the client over the session's existing SSE stream. The client fulfils the call by POSTing its result back, presenting call_token (single-use, hmac.compare_digest-compared). Honest threat model: this proves the responder had SESSION-STREAM READ ACCESS (received the SSE event) and prevents replay (single-use, consumed-once) — it does NOT prove the response came from the physical device the call was dispatched to. Any concurrent viewer of the same session's stream (e.g. a console tab) receives the same token and could answer on the device's behalf; this is the accepted threat model, distinct from the session's own API key. expires_at is an epoch-seconds deadline after which the server gives up waiting and resolves the tool call with a device_timeout error.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
class DeviceToolCallPayload(TypedDict):
    """Payload for a client-declared device-tool invocation (``device_tool_call``).

    Emitted when a session-bound ``ClientDeclaredTool`` (see ``client_tools.py``)
    is invoked; delivered to the client over the session's existing SSE stream.
    The client fulfils the call by POSTing its result back, presenting
    ``call_token`` (single-use, ``hmac.compare_digest``-compared). Honest
    threat model: this proves the responder had SESSION-STREAM READ ACCESS
    (received the SSE event) and prevents replay (single-use, consumed-once)
    — it does NOT prove the response came from the physical device the call
    was dispatched to. Any concurrent viewer of the same session's stream
    (e.g. a console tab) receives the same token and could answer on the
    device's behalf; this is the accepted threat model, distinct from the
    session's own API key.
    ``expires_at`` is an epoch-seconds deadline after which the server gives
    up waiting and resolves the tool call with a ``device_timeout`` error.
    """

    call_id: str
    call_token: str
    tool_id: str
    args: dict[str, object]
    expires_at: float

Event

Bases: TypedDict

Base event payload stored in transcripts.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
443
444
445
446
447
class Event(TypedDict):
    """Base event payload stored in transcripts."""

    type: str
    payload: EventPayload

EventRecord

Bases: Event

Event payload with a persisted timestamp.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
450
451
452
453
class EventRecord(Event):
    """Event payload with a persisted timestamp."""

    ts: str

LlmCallEndPayload

Bases: TypedDict

Payload emitted for a SUCCESSFUL llm_call_end event.

Brackets the whole RetryStrategy logical call — retries and fallback attempts included, not just the final attempt — so duration_ms is the wall time a caller actually waited, not the cheapest leg of it. The failed variant of this event (success: False) carries error_type/reason instead and is a separate, untyped payload — it never reaches this arm.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
class LlmCallEndPayload(TypedDict):
    """Payload emitted for a SUCCESSFUL ``llm_call_end`` event.

    Brackets the whole ``RetryStrategy`` logical call — retries and fallback
    attempts included, not just the final attempt — so ``duration_ms`` is the
    wall time a caller actually waited, not the cheapest leg of it. The failed
    variant of this event (``success: False``) carries ``error_type``/``reason``
    instead and is a separate, untyped payload — it never reaches this arm.
    """

    agent_id: str
    depth: int
    step: int
    success: Literal[True]
    model: str
    input_tokens: int
    output_tokens: int
    cache_creation_input_tokens: int
    cache_read_input_tokens: int
    reasoning_output_tokens: int
    cumulative_input_tokens: int
    cumulative_output_tokens: int
    duration_ms: int

LlmFallbackPayload

Bases: TypedDict

Payload emitted when the run advances to another model (llm_fallback).

reason is either a classifier reason (quota_exhausted, no_deployments, context_window, auth) for a switch_model decision, or retries_exhausted when the per-model retry cap tripped on a transient error. sticky is true when the destination model is pinned for the rest of the run (always true under the escalation policy).

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
class LlmFallbackPayload(TypedDict):
    """Payload emitted when the run advances to another model (``llm_fallback``).

    ``reason`` is either a classifier reason (``quota_exhausted``,
    ``no_deployments``, ``context_window``, ``auth``) for a ``switch_model``
    decision, or ``retries_exhausted`` when the per-model retry cap tripped on a
    transient error. ``sticky`` is true when the destination model is pinned for
    the rest of the run (always true under the escalation policy).
    """

    agent_id: str
    depth: int
    step: int
    from_model: str
    to_model: str
    reason: str
    previous_error_type: str
    sticky: NotRequired[bool]

LlmRetryPayload

Bases: TypedDict

Payload emitted before a same-model LLM retry (llm_retry event).

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
198
199
200
201
202
203
204
205
206
207
208
209
210
class LlmRetryPayload(TypedDict):
    """Payload emitted before a same-model LLM retry (``llm_retry`` event)."""

    agent_id: str
    depth: int
    step: int
    model: str
    attempt: int
    max_attempts: int
    error: str
    error_type: str
    delay: float
    retryable: bool

PermissionPayload

Bases: TypedDict

Payload emitted for permission decisions.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
39
40
41
42
43
44
45
class PermissionPayload(TypedDict):
    """Payload emitted for permission decisions."""

    tool_id: str
    operation: str
    tool_input: str
    decision: str

PlanApprovedPayload

Bases: TypedDict

Payload emitted when the user approves a proposed plan.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
178
179
180
181
182
class PlanApprovedPayload(TypedDict):
    """Payload emitted when the user approves a proposed plan."""

    plan_path: str
    revision: int

PlanProposedPayload

Bases: TypedDict

Payload emitted when the LLM calls exit_plan_mode.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
169
170
171
172
173
174
175
class PlanProposedPayload(TypedDict):
    """Payload emitted when the LLM calls ``exit_plan_mode``."""

    plan_path: str
    revision: int
    content: str
    summary: NotRequired[str]

PlanRejectedPayload

Bases: TypedDict

Payload emitted when the user rejects a proposed plan.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
185
186
187
188
189
class PlanRejectedPayload(TypedDict):
    """Payload emitted when the user rejects a proposed plan."""

    plan_path: str
    revision: int

PlanStepPayload

Bases: TypedDict

Payload describing a single plan step.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
14
15
16
17
18
class PlanStepPayload(TypedDict):
    """Payload describing a single plan step."""

    title: str
    description: str

RecoveryHaltPayload

Bases: TypedDict

Payload emitted when the doom-loop guard halts a no-progress run.

The recovery event with action == "halt_no_progress" — distinct from :class:RecoveryPayload (user-triggered retry/continue).

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
258
259
260
261
262
263
264
265
266
267
268
269
class RecoveryHaltPayload(TypedDict):
    """Payload emitted when the doom-loop guard halts a no-progress run.

    The ``recovery`` event with ``action == "halt_no_progress"`` — distinct from
    :class:`RecoveryPayload` (user-triggered retry/continue).
    """

    action: Literal["halt_no_progress"]
    agent_id: str
    depth: int
    step: int
    tool: str

RecoveryPayload

Bases: TypedDict

Payload emitted when the user triggers retry/continue after a failure.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
192
193
194
195
class RecoveryPayload(TypedDict):
    """Payload emitted when the user triggers retry/continue after a failure."""

    action: Literal["retry", "continue"]

SubAgentPayload

Bases: TypedDict

Payload describing a sub-agent lifecycle event.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
class SubAgentPayload(TypedDict):
    """Payload describing a sub-agent lifecycle event."""

    action: Literal["start", "stop"]
    agent_id: str
    parent_id: str | None
    depth: int
    model: str
    detail: str
    status: NotRequired[str]
    steps_completed: NotRequired[int]
    input_tokens: NotRequired[int]
    output_tokens: NotRequired[int]
    # The spawned AgentDef name (e.g. ``scg-path-probe``) — additive, present
    # only for an agent_type spawn so the trace projection can label the lane by
    # its definition rather than the model name.
    agent_type: NotRequired[str]
    # The child's compressed result (set only on the terminal ``stop``).
    summary: NotRequired[str]

TodoItemPayload

Bases: TypedDict

One authoritative todo item: a label plus its lifecycle status.

status is one of pending / in_progress / completed (see update_todos.TODO_STATUSES).

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
272
273
274
275
276
277
278
279
280
class TodoItemPayload(TypedDict):
    """One authoritative todo item: a label plus its lifecycle status.

    ``status`` is one of ``pending`` / ``in_progress`` / ``completed`` (see
    ``update_todos.TODO_STATUSES``).
    """

    label: str
    status: str

TodosPayload

Bases: TypedDict

Payload for the authoritative live todo list (todos event).

Re-emitted in FULL on every update_todos call (compaction-resilient). source discriminates the agent's live working set (agent) from an approved plan's roadmap (plan); agent_id attributes it to the emitting agent (the root, in practice).

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
283
284
285
286
287
288
289
290
291
292
293
294
class TodosPayload(TypedDict):
    """Payload for the authoritative live todo list (``todos`` event).

    Re-emitted in FULL on every ``update_todos`` call (compaction-resilient).
    ``source`` discriminates the agent's live working set (``agent``) from an
    approved plan's roadmap (``plan``); ``agent_id`` attributes it to the
    emitting agent (the root, in practice).
    """

    items: list[TodoItemPayload]
    source: str
    agent_id: str | None

ToolCallPayload

Bases: TypedDict

Payload describing a tool invocation that is about to be dispatched.

The initiation half of a tool call, emitted BEFORE the tool is awaited so a client can show the step while it runs rather than only once it finished. tool_call_id is the provider's call id and the only correlation key to the matching :class:ToolResultPayload; it is "" when the provider supplied none, which means "not correlatable" and never "shares a key with the others".

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
class ToolCallPayload(TypedDict):
    """Payload describing a tool invocation that is about to be dispatched.

    The initiation half of a tool call, emitted BEFORE the tool is awaited so a
    client can show the step while it runs rather than only once it finished.
    ``tool_call_id`` is the provider's call id and the only correlation key to the
    matching :class:`ToolResultPayload`; it is ``""`` when the provider supplied
    none, which means "not correlatable" and never "shares a key with the others".
    """

    tool_call_id: str
    tool_id: str
    operation: str
    tool_input: ToolInput
    agent_id: NotRequired[str]
    depth: NotRequired[int]
    model: NotRequired[str]

ToolResultPayload

Bases: TypedDict

Payload describing the outcome of a tool invocation.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
67
68
69
70
71
72
73
74
75
76
77
78
79
80
class ToolResultPayload(TypedDict):
    """Payload describing the outcome of a tool invocation."""

    tool_id: str
    operation: str
    tool_input: ToolInput
    result: str | None
    # Pairs this outcome with its :class:`ToolCallPayload`. Absent on every
    # transcript written before the initiation event existed, and ``""`` when the
    # provider supplied no id — both mean "not correlatable".
    tool_call_id: NotRequired[str]
    success: NotRequired[bool]
    summary: NotRequired[str]
    error: NotRequired[str]

UserPayload

Bases: TypedDict

Payload describing a user message.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
83
84
85
86
87
88
89
90
91
92
class UserPayload(TypedDict):
    """Payload describing a user message."""

    text: str
    # Present only when the turn carried attachments — the same
    # AttachmentDescriptor dicts riding the sibling ``context`` event's
    # ``payload.attachments`` (see ``mewbo_core.session.context._iter_attachments``),
    # additively duplicated here so clients can render attachment cards
    # above the user turn without cross-referencing an adjacent event.
    attachments: NotRequired[list[dict[str, object]]]

UserQuestionAnswerItemPayload

Bases: TypedDict

One delivered answer (indexes XOR text) on the answered event.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
365
366
367
368
369
class UserQuestionAnswerItemPayload(TypedDict):
    """One delivered answer (indexes XOR text) on the answered event."""

    selected_indexes: list[int] | None
    text: str | None

UserQuestionAnsweredPayload

Bases: TypedDict

Resolution record for a question group (user_question_answered).

outcome is answered / declined (user sent a message instead) / interrupted / cancelled / timed_out; answers and notes are present only when answered. Every surface — not just the one that answered — folds this onto its pending card.

Only answered settles a card. The other four record that the RUN stopped waiting, which is not the same as the question being resolved: the user may still answer, and that answer reaches the session as a new turn. A surface that greys out its card on any terminal outcome would be hiding the affordance precisely when the user finally came back to use it. delivery distinguishes the two landing paths for an answered outcome — run (the blocked tool call took it) vs message (the run had moved on, so it arrived as a new turn).

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
class UserQuestionAnsweredPayload(TypedDict):
    """Resolution record for a question group (``user_question_answered``).

    ``outcome`` is ``answered`` / ``declined`` (user sent a message instead)
    / ``interrupted`` / ``cancelled`` / ``timed_out``; ``answers`` and
    ``notes`` are present only when answered. Every surface — not just the one
    that answered — folds this onto its pending card.

    **Only ``answered`` settles a card.** The other four record that the RUN
    stopped waiting, which is not the same as the question being resolved: the
    user may still answer, and that answer reaches the session as a new turn.
    A surface that greys out its card on any terminal outcome would be hiding
    the affordance precisely when the user finally came back to use it.
    ``delivery`` distinguishes the two landing paths for an ``answered``
    outcome — ``run`` (the blocked tool call took it) vs ``message`` (the run
    had moved on, so it arrived as a new turn).
    """

    call_id: str
    outcome: str
    answered_via: str | None
    answers: list[UserQuestionAnswerItemPayload] | None
    notes: str | None
    delivery: str | None

UserQuestionItemPayload

Bases: TypedDict

One question in a user_question event's group.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
329
330
331
332
333
334
335
class UserQuestionItemPayload(TypedDict):
    """One question in a ``user_question`` event's group."""

    header: str
    question: str
    options: list[UserQuestionOptionPayload]
    multi_select: bool

UserQuestionOptionPayload

Bases: TypedDict

One selectable option on a user_question event.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
322
323
324
325
326
class UserQuestionOptionPayload(TypedDict):
    """One selectable option on a ``user_question`` event."""

    label: str
    description: str | None

UserQuestionPayload

Bases: TypedDict

Payload announcing a pending ask-user question group (user_question).

Emitted by the api's question dispatcher when the root agent calls ask_user_question (see ask_user.py); rides the session's SSE stream + backlog replay so every attached surface renders the card. call_token follows the device_tool_call threat model verbatim: it proves session-stream read access and prevents replay — not which surface answered.

This event is the DURABLE record of the question, and that is what makes a late answer possible. The in-process waiter registry is a rendezvous, not a store; it dies with the run. A client rendering this card long after the run moved on can still answer, because the answer route recovers the questions and the token from THIS payload. So the two additive fields are not decoration: timeout_seconds is what a surface needs to show how long the run will wait, and notes_placeholder is what makes it render the free-text box at all. Both are None when the run declared neither.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
class UserQuestionPayload(TypedDict):
    """Payload announcing a pending ask-user question group (``user_question``).

    Emitted by the api's question dispatcher when the root agent calls
    ``ask_user_question`` (see ``ask_user.py``); rides the session's SSE
    stream + backlog replay so every attached surface renders the card.
    ``call_token`` follows the ``device_tool_call`` threat model verbatim:
    it proves session-stream read access and prevents replay — not which
    surface answered.

    **This event is the DURABLE record of the question, and that is what makes
    a late answer possible.** The in-process waiter registry is a rendezvous,
    not a store; it dies with the run. A client rendering this card long after
    the run moved on can still answer, because the answer route recovers the
    questions and the token from THIS payload. So the two additive fields are
    not decoration: ``timeout_seconds`` is what a surface needs to show how
    long the run will wait, and ``notes_placeholder`` is what makes it render
    the free-text box at all. Both are ``None`` when the run declared neither.
    """

    call_id: str
    call_token: str
    questions: list[UserQuestionItemPayload]
    timeout_seconds: int | None
    notes_placeholder: str | None

VerificationPayload

Bases: TypedDict

One verifier-gate check verdict (verification event).

Bounded scalars ONLY: the grounded verifier stdout/stderr is injected into the model's own context (a SystemMessage), never onto the wire — so a chatty command can't bloat transcripts. attempt is 1-based; passed is the interpreted verdict; exit_code/timed_out explain a failure.

Source code in packages/mewbo_core/src/mewbo_core/contracts/types.py
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
class VerificationPayload(TypedDict):
    """One verifier-gate check verdict (``verification`` event).

    Bounded scalars ONLY: the grounded verifier stdout/stderr is injected into
    the model's own context (a SystemMessage), never onto the wire — so a
    chatty command can't bloat transcripts. ``attempt`` is 1-based; ``passed``
    is the interpreted verdict; ``exit_code``/``timed_out`` explain a failure.
    """

    agent_id: str
    depth: int
    step: int
    passed: bool
    attempt: int
    exit_code: int
    timed_out: bool

mewbo_core.config

Central JSON configuration for Mewbo.

APIAuthConfig

Bases: BaseModel

Identity & access management for the REST API (opt-in).

Off by default: with no auth block — or enabled: false — every request resolves to the built-in full-power identity and the server behaves exactly as it did before IAM existed. The union-shaped fields (authenticators, the group mappings, bootstrap) are carried here as open objects and validated in full against the identity kernel's typed models at server startup, which refuses to boot on an invalid block. These settings are documented in full in docs/authentication.md.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
class APIAuthConfig(BaseModel):
    """Identity & access management for the REST API (opt-in).

    Off by default: with no ``auth`` block — or ``enabled: false`` — every
    request resolves to the built-in full-power identity and the server behaves
    exactly as it did before IAM existed. The union-shaped fields
    (``authenticators``, the group mappings, ``bootstrap``) are carried here as
    open objects and validated in full against the identity kernel's typed models
    at server startup, which refuses to boot on an invalid block. These settings
    are documented in full in ``docs/authentication.md``.
    """

    model_config = ConfigDict(extra="forbid", json_schema_extra={"title": "Authentication"})

    enabled: bool = Field(
        False,
        description=(
            "Master switch for identity & access management. When off (the "
            "default), every request resolves to the built-in full-power identity "
            "and the server behaves exactly as it did before IAM. Turn on only "
            "after configuring at least one authenticator."
        ),
    )
    authenticators: list[AuthenticatorEntry] = Field(
        default_factory=list,
        description=(
            "Ordered list of identity sources (local API keys, OIDC, trusted "
            "reverse-proxy headers, LDAP, SAML). Each entry is an object whose "
            "`kind` field selects the authenticator type, plus that type's own "
            "settings. Validated in full at server startup; an invalid entry stops "
            "the server from booting. Each authenticator type's own settings are "
            "documented in `docs/authentication.md`."
        ),
    )
    role_mappings: dict[str, Any] | None = Field(
        None,
        description=(
            "Rules mapping identity-provider group names to Mewbo roles at login: "
            "an object with an ordered `rules` list and a `default_role` applied "
            "when no rule matches. The rule format is documented in "
            "`docs/authentication.md`."
        ),
    )
    team_mappings: dict[str, Any] | None = Field(
        None,
        description=(
            "Rules mapping identity-provider group names to team slugs at login: "
            "an object with an ordered `rules` list. The rule format is documented in "
            "`docs/authentication.md`."
        ),
    )
    bootstrap: dict[str, Any] | None = Field(
        None,
        description=(
            "Cold-start admin rule: grants the admin role to the first users "
            "matching an identity-provider group or an explicit subject allowlist, "
            "so an administrator exists before any role has been assigned. Documented "
            "in `docs/authentication.md`."
        ),
    )
    session: APIAuthSessionConfig = Field(
        default_factory=lambda: APIAuthSessionConfig.model_validate({}),
        description="Browser session/cookie settings for federated logins.",
    )
    avatars: dict[str, Any] | None = Field(
        None,
        description=(
            "Avatar-resolution policy: whether to fall back to Gravatar for users "
            "without a profile picture, and the default image style. Documented in "
            "`docs/authentication.md`."
        ),
    )
    scim: APIAuthScimConfig = Field(
        default_factory=lambda: APIAuthScimConfig.model_validate({}),
        description="SCIM 2.0 provisioning settings.",
    )
    audit: dict[str, Any] | None = Field(
        None,
        description=(
            "Auth audit-trail settings: an object with an `enabled` flag; on by "
            "default once IAM is enabled. The events recorded are listed in "
            "`docs/authentication.md`."
        ),
    )

APIAuthScimConfig

Bases: BaseModel

SCIM 2.0 provisioning settings.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
class APIAuthScimConfig(BaseModel):
    """SCIM 2.0 provisioning settings."""

    model_config = ConfigDict(extra="forbid", json_schema_extra={"title": "SCIM"})

    enabled: bool = Field(
        False,
        description="Whether the SCIM 2.0 provisioning endpoint is served.",
    )
    secret: str = Field(
        "",
        description=(
            "Bearer secret an identity provider presents to the SCIM endpoint. "
            "Write-only: never returned by the config API."
        ),
        json_schema_extra={"x-secret": True},
    )

APIAuthSessionConfig

Bases: BaseModel

Browser session/cookie settings for federated logins.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
class APIAuthSessionConfig(BaseModel):
    """Browser session/cookie settings for federated logins."""

    model_config = ConfigDict(extra="forbid", json_schema_extra={"title": "Session"})

    cookie_name: str = Field(
        "mewbo_session",
        description="Name of the browser session cookie issued after a federated login.",
    )
    ttl_seconds: int = Field(
        28800,
        ge=1,
        description="Lifetime of a browser session, in seconds.",
    )
    secret: str = Field(
        "",
        description=(
            "Signing secret for the browser session cookie. Required once any "
            "non-API-key authenticator is configured. Write-only: never returned "
            "by the config API."
        ),
        json_schema_extra={"x-secret": True},
    )

APIConfig

Bases: BaseModel

REST API authentication.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
class APIConfig(BaseModel):
    """REST API authentication."""

    model_config = ConfigDict(
        json_schema_extra={"title": "API Server", "x-group": "server", "x-order": 1},
    )

    master_token: str = Field(
        "msk-strong-password",
        description=(
            "Bearer token required for all REST API requests. Change from the "
            "default before deploying. "
            "The `MEWBO_MASTER_API_TOKEN` environment variable OVERRIDES whatever "
            "is set here: when it is present, the API and the MCP server both "
            "use it and this value is never consulted — which is the case in "
            "every containerised deployment. To keep the token out of this file "
            "and say so plainly, set this to `${MEWBO_MASTER_API_TOKEN}`. Any "
            "value may name an environment variable that way, and a name that is "
            "not set in the environment is refused at startup rather than read "
            "as empty."
        ),
        examples=["${MEWBO_MASTER_API_TOKEN}"],
        json_schema_extra={"x-protected": True},
    )
    allow_external_cwd: bool = Field(
        False,
        description=(
            "Allow callers to anchor sessions in arbitrary host paths via the `cwd` "
            "field on POST /api/sessions and POST /api/sessions/{id}/query. A "
            "directory belonging to a configured project, managed project or worktree, "
            "or registered repository checkout is always accepted, as is a re-send of "
            "the session's own bound directory; this flag governs host paths a caller "
            "names that the server does not already own."
        ),
    )
    max_concurrent_streams: int = Field(
        6,
        ge=0,
        description=(
            "How many Server-Sent Events streams the server keeps open at once. "
            "A stream holds one of the server's request threads for as long as "
            "it stays open rather than for the work it does, so without a bound "
            "enough of them starve every other endpoint and the server stops "
            "answering at all. Past this many, a new stream is refused with a "
            "retryable 503 and the rest of the API keeps serving. Keep it below "
            "the worker's thread count so ordinary requests always have "
            "headroom; 0 removes the bound."
        ),
    )
    apps_token_secret: str = Field(
        "",
        description=(
            "Signing secret for Mewbo Apps render tokens — the short-lived, "
            "app-scoped read tokens the served app frontend presents on the "
            "read-only data/system endpoints. Set this to sign (and rotate) app "
            "tokens independently of the master token; when left empty it falls "
            "back to the master token, logging one startup warning."
        ),
        json_schema_extra={"x-secret": True},
    )
    apps_exec_binaries: list[str] = Field(
        default_factory=lambda: ["git", "tea", "gh"],
        description=(
            "Command-line programs a Mewbo App's code pipeline may run. A pipeline "
            "must still declare the ones it needs, so this is a ceiling the "
            "deployment sets rather than a grant: a program absent here cannot be "
            "reached however the pipeline is written. Widening it is an operator "
            "decision — pipeline code runs in-process, so a program added here runs "
            "with the API server's own file and network access, and a program that "
            "can be steered into running other programs effectively grants a shell."
        ),
    )
    apps_max_concurrent_pipelines: int = Field(
        4,
        ge=0,
        description=(
            "How many Mewbo App pipelines may execute at once. A pipeline runs "
            "synchronously and holds one of the server's request threads for its "
            "whole duration, so without a bound enough concurrent invocations "
            "starve every other endpoint. Past this many, an invocation is refused "
            "with a retryable 429 and the rest of the API keeps serving. 0 removes "
            "the bound."
        ),
    )
    auth: APIAuthConfig = Field(
        default_factory=lambda: APIAuthConfig.model_validate({}),
        description=(
            "Identity & access management (opt-in; off by default). Configures "
            "authenticators, roles, and sessions. Documented in "
            "`docs/authentication.md`."
        ),
    )

AgentConfig

Bases: EnvOverridable

Sub-agent hypervisor settings.

Source code in packages/mewbo_core/src/mewbo_core/config.py
2274
2275
2276
2277
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
2291
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
2308
2309
2310
2311
2312
2313
2314
2315
2316
2317
2318
2319
2320
2321
2322
2323
2324
2325
2326
2327
2328
2329
2330
2331
2332
2333
2334
2335
2336
2337
2338
2339
2340
2341
2342
2343
2344
2345
2346
2347
2348
2349
2350
2351
2352
2353
2354
2355
2356
2357
2358
2359
2360
2361
2362
2363
2364
2365
2366
2367
2368
2369
2370
2371
2372
2373
2374
2375
2376
2377
2378
2379
2380
2381
2382
2383
2384
2385
2386
2387
2388
2389
2390
2391
2392
2393
2394
2395
2396
2397
2398
2399
2400
2401
2402
2403
2404
2405
2406
2407
2408
2409
2410
2411
2412
2413
2414
2415
2416
2417
2418
2419
2420
2421
2422
2423
2424
2425
2426
2427
2428
2429
2430
2431
2432
2433
2434
2435
2436
2437
2438
2439
2440
2441
2442
2443
2444
2445
2446
2447
2448
2449
2450
2451
2452
2453
2454
2455
2456
2457
2458
2459
2460
2461
2462
2463
2464
2465
2466
2467
2468
2469
2470
2471
2472
2473
2474
2475
2476
2477
2478
2479
2480
2481
2482
2483
2484
2485
2486
2487
2488
2489
2490
2491
2492
2493
2494
2495
2496
2497
2498
2499
2500
2501
2502
2503
2504
2505
2506
2507
2508
2509
2510
2511
2512
2513
2514
2515
2516
2517
2518
2519
2520
2521
2522
2523
2524
2525
2526
2527
2528
2529
2530
2531
2532
2533
2534
2535
2536
2537
2538
2539
2540
2541
2542
2543
2544
2545
2546
2547
2548
2549
2550
2551
2552
2553
2554
2555
2556
2557
2558
2559
2560
2561
2562
2563
2564
2565
2566
2567
2568
2569
2570
2571
2572
2573
2574
2575
2576
2577
2578
2579
2580
2581
2582
2583
2584
2585
2586
2587
2588
2589
2590
2591
2592
2593
2594
2595
2596
2597
2598
2599
2600
2601
2602
2603
2604
2605
2606
2607
2608
2609
2610
2611
2612
2613
2614
2615
2616
2617
2618
2619
2620
2621
2622
2623
2624
2625
2626
2627
2628
2629
2630
2631
2632
2633
2634
2635
2636
2637
2638
2639
2640
2641
2642
2643
2644
2645
2646
2647
2648
2649
2650
2651
2652
2653
2654
2655
2656
2657
2658
2659
2660
2661
2662
2663
2664
2665
2666
2667
2668
2669
2670
2671
2672
2673
2674
2675
2676
2677
2678
2679
2680
2681
2682
2683
2684
2685
2686
2687
2688
2689
2690
2691
2692
2693
2694
2695
2696
2697
2698
2699
2700
2701
2702
2703
2704
2705
2706
2707
2708
2709
2710
2711
2712
2713
2714
2715
2716
2717
2718
2719
2720
2721
2722
2723
2724
2725
2726
2727
2728
2729
2730
2731
2732
2733
2734
2735
2736
2737
2738
2739
2740
2741
2742
2743
2744
2745
2746
2747
2748
2749
2750
2751
2752
2753
2754
2755
2756
2757
2758
2759
2760
2761
2762
2763
2764
2765
2766
2767
2768
2769
2770
2771
2772
2773
2774
2775
2776
2777
2778
2779
2780
2781
2782
2783
2784
2785
2786
2787
2788
2789
2790
2791
2792
2793
2794
2795
2796
2797
2798
2799
2800
2801
2802
2803
2804
2805
2806
2807
2808
2809
2810
2811
2812
2813
2814
2815
2816
2817
2818
2819
2820
2821
2822
2823
2824
2825
2826
2827
2828
2829
2830
2831
2832
2833
2834
2835
2836
2837
2838
2839
2840
2841
2842
2843
2844
2845
2846
2847
2848
2849
2850
2851
2852
2853
2854
2855
2856
2857
2858
2859
2860
2861
2862
2863
2864
2865
2866
2867
2868
2869
2870
2871
2872
2873
2874
2875
2876
2877
2878
2879
2880
2881
2882
2883
2884
2885
2886
2887
2888
2889
2890
2891
2892
2893
2894
2895
2896
2897
2898
2899
2900
2901
2902
2903
2904
2905
2906
2907
class AgentConfig(EnvOverridable):
    """Sub-agent hypervisor settings."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Agent", "x-group": "agent", "x-order": 1},
    )

    max_depth: int = Field(
        5, description="Maximum nesting depth for sub-agent delegation (1 = no sub-agents)."
    )
    max_concurrent: int = Field(
        20, description="Maximum number of sub-agents allowed to run concurrently."
    )
    default_sub_model: str = Field(
        "",
        description=(
            "Default LLM model for sub-agents. Falls back to the root agent's model when empty."
        ),
        examples=["anthropic/claude-haiku-4-5"],
    )
    allowed_models: list[str] = Field(
        default_factory=list,
        description=(
            "Allowlist of model names sub-agents may use. Empty means all models are allowed."
        ),
    )
    model_tiers: dict[str, str] = Field(
        default_factory=dict,
        description=(
            "Coarse model-cost tier -> concrete model id map (e.g. "
            "{'economy': 'anthropic/claude-haiku-4-5', 'frontier': "
            "'anthropic/claude-opus-4-6'}), resolved by a spawned agent's "
            "DelegationContract.model_tier. An explicit spawn "
            "model arg or an agent_type's configured model always wins; a "
            "declared tier with no map entry is a silent no-op fallthrough."
        ),
    )
    max_iters: int = Field(
        30,
        deprecated=True,
        description=(
            "Deprecated. The tool-use loop now runs until natural completion "
            "(model returns text without tool calls). This field is retained "
            "for API backward compatibility but is not enforced."
        ),
    )
    session_step_budget: int = Field(
        0,
        description=(
            "Ceiling on total tool-execution steps across every agent in a "
            "session (root + all sub-agents combined), enforced by the "
            "hypervisor: a warning is injected as the budget nears, and the "
            "run hard-stops at exhaustion. 0 = unlimited."
        ),
    )
    attestation_enabled: bool = Field(
        True,
        description=(
            "Record a best-effort provenance hash chain over "
            "every spawn/terminal transition in a session's agent tree — "
            "bounded scalars + a contract snapshot + a summary FINGERPRINT "
            "only, never task text or raw summary content. Additive and "
            "never fatal: a failed or absent chain simply records no "
            "attestation. Default ON; this is the kill switch."
        ),
    )
    default_workspace_mode: str = Field(
        "full_access",
        description=(
            "Root filesystem-containment tier every session starts at, "
            "narrowed per sub-agent by spawn_agent's workspace_mode. One "
            "of 'read_only' (reads confined to the workspace, no writes), "
            "'workspace_write' (reads + writes confined to the workspace), or "
            "'full_access' (no path restriction). The root stays 'full_access' "
            "because containment attenuates privilege ACROSS A SPAWN: the root "
            "agent acts directly for the operator, while a sub-agent defaults to "
            "'workspace_write' and can only ever narrow further. Set this to "
            "confine the root session itself as well."
        ),
        examples=["full_access", "workspace_write", "read_only"],
    )
    workspace_enforcement: bool = Field(
        True,
        description=(
            "Master switch for workspace_mode filesystem containment. While on, "
            "an agent whose tier is narrower than 'full_access' resolves every "
            "path against its own workspace root plus the Mewbo-owned scratch "
            "roots, and the workspace root it was handed is authoritative — a "
            "wider root supplied in tool arguments is ignored rather than "
            "honoured. Turn it off to restore the older behaviour, where a path "
            "resolves against the union of every configured project root and an "
            "agent's tier governs nothing."
        ),
    )
    shell_sandbox: bool = Field(
        True,
        description=(
            "Confine shell subprocesses using the kernel's Landlock LSM, so a "
            "command cannot read another task's files no matter which binary "
            "it runs. This is a DENY-list, not an allowlist: every configured "
            "project other than the session's own active one is denied, "
            "together with anything listed in agent.shell_denied_paths. "
            "Everything else — the interpreter, system libraries, the CLI "
            "toolbox and HOME — stays reachable, so nothing has to be "
            "enumerated to keep the shell working. Normal filesystem "
            "permissions still apply on top of this; Landlock only removes "
            "access, never grants it. On a kernel without Landlock this logs "
            "once and changes nothing."
        ),
    )
    shell_denied_paths: list[str] = Field(
        default_factory=list,
        description=(
            "Extra absolute directories the shell tool may never read or "
            "write, on top of the other-projects denial shell_sandbox already "
            "applies — for example the harness's own source and config "
            "directory on a deployment where those sit outside every "
            "configured project. A path that does not exist is ignored. This "
            "only affects commands run through the shell tool; every other "
            "tool is already argument-validated."
        ),
    )
    harness_self_deny: bool = Field(
        True,
        description=(
            "Hide Mewbo's own source trees and the directory holding app.json "
            "from every session, so a command cannot read the file that holds "
            "the API keys. The directories are derived from where Mewbo is "
            "installed rather than configured, and a directory containing the "
            "Python runtime is never denied — that would stop every command "
            "instead of confining it. Turn this off to develop Mewbo itself, "
            "where its packages are the work rather than internals to hide; the "
            "session then reaches them like any other project. It denies "
            "unconditionally by default, including when the session's own "
            "project contains the installation, because the alternative was to "
            "guess when to make an exception: the guess is invisible on a "
            "deployment, where an unnoticed exception exposes the keys, while "
            "an unwanted denial is immediate and local on a workstation, where "
            "it is fixed by turning this off."
        ),
    )
    path_scope_to_active_project: bool = Field(
        True,
        description=(
            "Scope the path-taking tools — file read, file edit, directory "
            "listing and LSP — to the session's active project, so a path "
            "argument resolves only under that project plus the extra "
            "directories its allowed_paths re-admits. While off, a path "
            "resolves against the union of every configured project root and "
            "the API host's own working directory, which lets one session read "
            "another project's files by naming them. This is the "
            "argument-validated half of the boundary agent.shell_sandbox "
            "enforces in the kernel for shell subprocesses, and both halves "
            "read the same active project. It is a separate switch because "
            "shell_sandbox also decides whether a kernel mechanism is applied "
            "at all, so a deployment that turns that one off — an older kernel, "
            "or a ruleset that broke a command — should not silently lose this "
            "check too. The Mewbo-owned scratch roots stay reachable either "
            "way, and a call with no active project, such as a direct library "
            "caller, is unaffected."
        ),
    )
    server_sandbox: bool = Field(
        False,
        description=(
            "Extend the shell sandbox to the MCP and language servers Mewbo "
            "configures but does not spawn itself. Their command line is "
            "prefixed with a small launcher that applies the same Landlock "
            "ruleset to itself and is then replaced by the real server, which "
            "inherits the confinement. Each server is scoped to the workspace "
            "it serves — a language server to the project root it resolves, an "
            "MCP server to its own configured working directory — and denied "
            "every other configured project, so an MCP server that names no "
            "working directory reaches none of them. Servers reached over HTTP "
            "spawn no process here and are unaffected. This needs "
            "agent.shell_sandbox as well, because it applies the same kernel "
            "mechanism: a deployment that turned that one off must not have it "
            "reappear underneath its language servers. Off by default because "
            "a denied path surfaces inside a server as a missing file rather "
            "than as a refusal, and which servers a deployment runs is not "
            "knowable in advance."
        ),
    )
    stall_threshold_s: float = Field(
        120.0,
        description=(
            "Seconds of no tool-execution progress before the watchdog "
            "flags an agent as stalled and injects an NL warning into its "
            "message queue."
        ),
    )
    stall_check_interval_s: float = Field(
        30.0,
        description="Seconds between watchdog stall-detection sweeps.",
    )
    write_progress_signal_step_threshold: int = Field(
        DEFAULT_WRITE_PROGRESS_THRESHOLD,
        description=(
            "Consecutive non-write tool-execution steps before the "
            "write-progress signal fires telemetry for a write-capable "
            "agent. 0 disables."
        ),
    )
    write_progress_signal_event_interval: int = Field(
        DEFAULT_WRITE_PROGRESS_EVENT_INTERVAL,
        description=(
            "Steps between repeat write-progress signal events once the threshold is crossed."
        ),
    )
    write_progress_signal_max_events: int = Field(
        DEFAULT_WRITE_PROGRESS_MAX_EVENTS,
        description="Maximum write-progress signal events emitted before it goes quiet.",
    )
    write_progress_signal_reminder_enabled: bool = Field(
        False,
        description=(
            "When the write-progress signal fires, also inject a "
            "criterion-blind objective-restatement reminder (states the "
            "task goal only — never the signal or its criteria). Default "
            "off: telemetry alone is the observe-only default."
        ),
    )
    verification_enabled: bool = Field(
        False,
        description=(
            "Master switch for verifier-gated completion. OFF by default "
            "(staged): while off, a spawn's verification spec is carried but "
            "never run, so every natural completion is accepted unchanged. "
            "Flip on to gate a write-capable agent's claimed "
            "completion behind a ground-truth command check before its text "
            "is accepted."
        ),
    )
    verification_max_retries: int = Field(
        2,
        description=(
            "Maximum times a failed completion verifier re-drives the agent "
            "before its text is accepted, honestly flagged verification_failed. "
            "Clamped to [0, 10] and bounded ALSO by the step/wall budget, "
            "whichever is tighter. 0 = one check, no retry."
        ),
    )
    verification_timeout_s: float = Field(
        DEFAULT_VERIFICATION_TIMEOUT,
        description=(
            "Ceiling in seconds on a single verifier subprocess; a spawn's "
            "per-spec timeout_s is clamped down to this at run. Clamped to "
            "[1, 600]."
        ),
    )
    llm_call_timeout: float = Field(
        DEFAULT_TIMEOUT,
        description=(
            "Ceiling in seconds for a single model.ainvoke() call. "
            "Covers extended-thinking models (raised from 60s because bare "
            "timeouts were the largest single failure class). On timeout, the call is "
            "retried up to llm_call_retries times before cascading to "
            "fallback models. Deployments where one call legitimately runs long "
            "(slow local inference, a saturated proxy) can lift the ceiling "
            "without editing this file via MEWBO_AGENT_LLM_CALL_TIMEOUT."
        ),
        json_schema_extra={"x-env-var": "MEWBO_AGENT_LLM_CALL_TIMEOUT"},
    )
    llm_first_token_timeout: float = Field(
        DEFAULT_FIRST_TOKEN_TIMEOUT,
        description=(
            "Ceiling in seconds on the wait for the FIRST chunk of a streamed "
            "model call. 0 (the default) defers to llm_call_timeout, so one "
            "bound covers the whole call. A large prompt on an "
            "extended-thinking model can legitimately take minutes to emit its "
            "first token, so a tighter value here manufactures failures unless "
            "a deployment has measured its own first-token latency."
        ),
        json_schema_extra={"x-advanced": True},
    )
    llm_stream_idle_timeout: float = Field(
        DEFAULT_STREAM_IDLE_TIMEOUT,
        description=(
            "Maximum seconds between chunks once a streamed model call has "
            "started producing. This is the bound llm_call_timeout cannot "
            "express: a provider that returns 200 and then goes silent is "
            "otherwise unbounded until the total ceiling, and a total ceiling "
            "loose enough for a long healthy generation is far too loose for a "
            "dead one. 0 disables it, leaving llm_call_timeout as the only "
            "bound."
        ),
        json_schema_extra={"x-advanced": True},
    )
    llm_call_liveness_s: float = Field(
        DEFAULT_LLM_CALL_LIVENESS_S,
        description=(
            "Seconds one outstanding model call may go silent before the loop emits "
            "an llm_call_stalled event. This DETECTS a wedged call; it does not "
            "recover one — asyncio's attempt cap can only cancel a coroutine that "
            "reaches a cancellation point, and a provider read wedged below the "
            "event loop never does. Keep it above llm_call_timeout, or a normally "
            "timing-out call reports as stalled."
        ),
        json_schema_extra={"x-advanced": True},
    )
    llm_call_retries: int = Field(
        DEFAULT_PRIMARY_RETRIES,
        description=(
            "Maximum attempts for the primary model before cascading to "
            "fallback models (default 2 = one try + one retry). Each fallback "
            "model gets retry.fallback_retries attempts. A rescue model that "
            "wins is pinned for the rest of the run. "
            "Backoff/budget/circuit-breaker live under agent.retry."
        ),
    )

    @field_validator("attestation_enabled", mode="before")
    @classmethod
    def _normalize_attestation_enabled(cls, value: Any) -> bool:
        return _coerce_bool(value, default=True)

    @field_validator("default_workspace_mode", mode="before")
    @classmethod
    def _normalize_default_workspace_mode(cls, value: Any) -> str:
        # Unknown / malformed collapses to the widest (no-op) tier — the same
        # unknown → widest rule the AgentContext narrowing applies, so a typo
        # never silently CONTAINS a deployment that meant full access.
        if value is None:
            return "full_access"
        normalized = str(value).strip().lower()
        if normalized in {"read_only", "workspace_write", "full_access"}:
            return normalized
        return "full_access"

    @field_validator("workspace_enforcement", mode="before")
    @classmethod
    def _normalize_workspace_enforcement(cls, value: Any) -> bool:
        return _coerce_bool(value, default=False)

    @field_validator("shell_denied_paths", mode="before")
    @classmethod
    def _normalize_shell_denied_paths(cls, value: Any) -> list[str]:
        raw_list = value if isinstance(value, list) else [value] if value else []
        normalized: list[str] = []
        for raw in raw_list:
            text = str(raw).strip() if raw else ""
            if text:
                normalized.append(str(Path(text).expanduser().resolve()))
        return normalized

    @field_validator("verification_enabled", mode="before")
    @classmethod
    def _normalize_verification_enabled(cls, value: Any) -> bool:
        return _coerce_bool(value, default=False)

    @field_validator("verification_max_retries", mode="before")
    @classmethod
    def _clamp_verification_max_retries(cls, value: Any) -> int:
        # Coerce-and-clamp: a malformed value collapses to the default rather
        # than failing config load; the bound keeps a runaway retry count from
        # starving the step/wall budget.
        try:
            return max(0, min(10, int(value)))
        except (TypeError, ValueError):
            return 2

    @field_validator("verification_timeout_s", mode="before")
    @classmethod
    def _clamp_verification_timeout(cls, value: Any) -> float:
        try:
            return max(1.0, min(600.0, float(value)))
        except (TypeError, ValueError):
            return DEFAULT_VERIFICATION_TIMEOUT

    retry: RetryConfig = Field(
        default_factory=lambda: RetryConfig.model_validate({}),
        description="Automatic LLM-call retry / fallback resilience knobs.",
    )

    def ladder_budget_seconds(self, fallback_count: int) -> float:
        """Worst-case seconds the configured model ladder can spend in one turn.

        The primary rung may spend ``llm_call_timeout x llm_call_retries``; every
        fallback rung may spend ``llm_call_timeout x retry.fallback_retries``.
        The rung count arrives as an ARGUMENT because the ladder is declared
        under ``llm``, not here — this model must never reach across the config
        tree to read it.
        """
        rungs = self.llm_call_timeout * max(self.llm_call_retries, 1)
        per_fallback = self.llm_call_timeout * max(self.retry.fallback_retries, 1)
        return rungs + per_fallback * max(fallback_count, 0)

    def unreachable_rung(self, fallback_count: int) -> tuple[float, float] | None:
        """``(needed, deadline)`` when the ladder cannot fit, else ``None``.

        THE BUDGET LAW. ``retry.turn_deadline`` bounds ONE logical LLM call
        across every retry and every fallback, and it is checked between
        attempts. So rungs that can together consume the whole budget leave
        every rung below them unreachable — the chain advances, finds the
        deadline spent, and stops. Cross-model fallback then silently never
        runs, which reads as a provider-wide outage rather than as the tuning
        mistake it is.
        """
        deadline = self.retry.turn_deadline
        if deadline <= 0:
            return None  # terminator disabled — nothing bounds the ladder
        needed = self.ladder_budget_seconds(fallback_count)
        return (needed, deadline) if needed > deadline else None

    @model_validator(mode="after")
    def _warn_when_fallback_is_unreachable(self) -> AgentConfig:
        """Report a budget too small for even a SINGLE fallback rung.

        This is the floor check, and it is all this model can perform on its own:
        the real ladder length lives in ``llm.fallback``, so ``AppConfig`` runs
        the full-ladder check once both halves are present. Raising
        ``llm_call_timeout`` without raising ``turn_deadline`` is the way
        deployments walk into the floor case.

        This reports instead of refusing: the single-model path still works, and a
        hard failure at config load would take down a running deployment over a
        knob it can keep serving traffic with (same reasoning as the
        ``system_instructions`` template-compilation carve-out).
        """
        breach = self.unreachable_rung(1)
        if breach is not None:
            needed, deadline = breach
            _logger.warning(
                "agent.retry.turn_deadline=%.0fs cannot reach a fallback model: the "
                "primary rung alone may spend %.0fs (llm_call_timeout=%.0fs x "
                "llm_call_retries=%d) and one fallback rung needs %.0fs more. "
                "Cross-model fallback will not run. Raise turn_deadline to >= %.0fs.",
                deadline,
                self.llm_call_timeout * max(self.llm_call_retries, 1),
                self.llm_call_timeout,
                self.llm_call_retries,
                self.llm_call_timeout * max(self.retry.fallback_retries, 1),
                needed,
            )
        return self

    default_denied_tools: list[str] = Field(
        default_factory=list,
        description="Tool IDs denied to all sub-agents by default (e.g. spawn_agent).",
    )
    edit_tool: str = Field(
        "",
        description=(
            "File editing mechanism override: 'search_replace_block' (Aider-style "
            "SEARCH/REPLACE blocks) or 'structured_patch' (per-file exact "
            "string replacement). Leave empty (default) to auto-select based on "
            "the active model via llm.structured_patch_models."
        ),
        examples=["", "search_replace_block", "structured_patch"],
    )
    plan_mode_shell_allowlist: list[str] = Field(
        default_factory=lambda: [
            # Filesystem inspection
            "ls",
            "pwd",
            "cat",
            "head",
            "tail",
            "wc",
            "file",
            "stat",
            "tree",
            # Searching
            "find",
            "grep",
            "rg",
            "ag",
            "ack",
            # Environment / process / system introspection
            "echo",
            "which",
            "whereis",
            "env",
            "printenv",
            "ps",
            "uname",
            "date",
            # Disk usage
            "du",
            "df",
            # Git read-only subcommands (prefix-matched; all flags/args allowed)
            "git status",
            "git log",
            "git diff",
            "git show",
            "git blame",
            "git branch",
            "git tag",
            "git remote",
            "git config --get",
            "git rev-parse",
            "git ls-files",
            "git describe",
            "git reflog",
        ],
        description=(
            "Shell command prefixes allowed during plan mode. Each entry "
            "matches a command at a word boundary (e.g. 'git log' matches "
            "'git log --oneline' but not 'git logger'). Commands containing "
            "pipes, redirects, variable expansion, command substitution, or "
            "chaining (|, >, <, &, ;, $, backtick) are always rejected. "
            "Set to an empty list to disable shell in plan mode entirely."
        ),
    )
    web_ide: WebIdeConfig | None = Field(
        default=None,
        description="Optional 'Open in Web IDE' feature config (code-server containers).",
    )
    lsp: LSPConfig = Field(
        default_factory=lambda: LSPConfig.model_validate({}),
        description="Language Server Protocol integration settings.",
    )
    tool_search: ToolSearchConfig = Field(
        default_factory=lambda: ToolSearchConfig.model_validate({}),
        description="Deferred tool loading via on-demand schema fetching.",
    )

    @field_validator("edit_tool", mode="before")
    @classmethod
    def _normalize_edit_tool(cls, value: Any) -> str:
        if value is None:
            return ""
        normalized = str(value).strip().lower()
        if normalized in {"", "search_replace_block", "structured_patch"}:
            return normalized
        return ""

    @field_validator("max_depth", mode="before")
    @classmethod
    def _normalize_max_depth(cls, value: Any) -> int:
        try:
            parsed = int(value)
        except (TypeError, ValueError):
            return 5
        return max(parsed, 1)

    @field_validator("max_concurrent", mode="before")
    @classmethod
    def _normalize_max_concurrent(cls, value: Any) -> int:
        try:
            parsed = int(value)
        except (TypeError, ValueError):
            return 20
        return max(parsed, 1)

    @field_validator("max_iters", mode="before")
    @classmethod
    def _normalize_max_iters(cls, value: Any) -> int:
        try:
            parsed = int(value)
        except (TypeError, ValueError):
            return 30
        return max(parsed, 1)

    @field_validator("session_step_budget", mode="before")
    @classmethod
    def _normalize_session_step_budget(cls, value: Any) -> int:
        try:
            parsed = int(value)
        except (TypeError, ValueError):
            return 0
        return max(parsed, 0)

    @field_validator("stall_threshold_s", mode="before")
    @classmethod
    def _normalize_stall_threshold_s(cls, value: Any) -> float:
        try:
            parsed = float(value)
        except (TypeError, ValueError):
            return 120.0
        return parsed if parsed > 0 else 120.0

    @field_validator("stall_check_interval_s", mode="before")
    @classmethod
    def _normalize_stall_check_interval_s(cls, value: Any) -> float:
        try:
            parsed = float(value)
        except (TypeError, ValueError):
            return 30.0
        return parsed if parsed > 0 else 30.0

    @field_validator("write_progress_signal_step_threshold", mode="before")
    @classmethod
    def _normalize_write_progress_step_threshold(cls, value: Any) -> int:
        try:
            parsed = int(value)
        except (TypeError, ValueError):
            return DEFAULT_WRITE_PROGRESS_THRESHOLD
        return max(parsed, 0)

    @field_validator("write_progress_signal_event_interval", mode="before")
    @classmethod
    def _normalize_write_progress_event_interval(cls, value: Any) -> int:
        try:
            parsed = int(value)
        except (TypeError, ValueError):
            return DEFAULT_WRITE_PROGRESS_EVENT_INTERVAL
        return max(parsed, 1)

    @field_validator("write_progress_signal_max_events", mode="before")
    @classmethod
    def _normalize_write_progress_max_events(cls, value: Any) -> int:
        try:
            parsed = int(value)
        except (TypeError, ValueError):
            return DEFAULT_WRITE_PROGRESS_MAX_EVENTS
        return max(parsed, 0)

    @field_validator(
        "allowed_models",
        "default_denied_tools",
        "plan_mode_shell_allowlist",
        mode="before",
    )
    @classmethod
    def _normalize_string_lists(cls, value: Any) -> list[str]:
        return _coerce_list(value)

    @field_validator("model_tiers", mode="before")
    @classmethod
    def _normalize_model_tiers(cls, value: Any) -> dict[str, str]:
        if not isinstance(value, dict):
            return {}
        normalized: dict[str, str] = {}
        dropped: list[Any] = []
        for k, v in value.items():
            if isinstance(k, str) and isinstance(v, str) and k.strip() and v.strip():
                normalized[k.strip()] = v.strip()
            else:
                dropped.append(k)
        if dropped:
            _logger.warning("agent.model_tiers: dropping malformed entries %r", dropped)
        return normalized

ladder_budget_seconds(fallback_count: int) -> float

Worst-case seconds the configured model ladder can spend in one turn.

The primary rung may spend llm_call_timeout x llm_call_retries; every fallback rung may spend llm_call_timeout x retry.fallback_retries. The rung count arrives as an ARGUMENT because the ladder is declared under llm, not here — this model must never reach across the config tree to read it.

Source code in packages/mewbo_core/src/mewbo_core/config.py
2649
2650
2651
2652
2653
2654
2655
2656
2657
2658
2659
2660
def ladder_budget_seconds(self, fallback_count: int) -> float:
    """Worst-case seconds the configured model ladder can spend in one turn.

    The primary rung may spend ``llm_call_timeout x llm_call_retries``; every
    fallback rung may spend ``llm_call_timeout x retry.fallback_retries``.
    The rung count arrives as an ARGUMENT because the ladder is declared
    under ``llm``, not here — this model must never reach across the config
    tree to read it.
    """
    rungs = self.llm_call_timeout * max(self.llm_call_retries, 1)
    per_fallback = self.llm_call_timeout * max(self.retry.fallback_retries, 1)
    return rungs + per_fallback * max(fallback_count, 0)

unreachable_rung(fallback_count: int) -> tuple[float, float] | None

(needed, deadline) when the ladder cannot fit, else None.

THE BUDGET LAW. retry.turn_deadline bounds ONE logical LLM call across every retry and every fallback, and it is checked between attempts. So rungs that can together consume the whole budget leave every rung below them unreachable — the chain advances, finds the deadline spent, and stops. Cross-model fallback then silently never runs, which reads as a provider-wide outage rather than as the tuning mistake it is.

Source code in packages/mewbo_core/src/mewbo_core/config.py
2662
2663
2664
2665
2666
2667
2668
2669
2670
2671
2672
2673
2674
2675
2676
2677
def unreachable_rung(self, fallback_count: int) -> tuple[float, float] | None:
    """``(needed, deadline)`` when the ladder cannot fit, else ``None``.

    THE BUDGET LAW. ``retry.turn_deadline`` bounds ONE logical LLM call
    across every retry and every fallback, and it is checked between
    attempts. So rungs that can together consume the whole budget leave
    every rung below them unreachable — the chain advances, finds the
    deadline spent, and stops. Cross-model fallback then silently never
    runs, which reads as a provider-wide outage rather than as the tuning
    mistake it is.
    """
    deadline = self.retry.turn_deadline
    if deadline <= 0:
        return None  # terminator disabled — nothing bounds the ladder
    needed = self.ladder_budget_seconds(fallback_count)
    return (needed, deadline) if needed > deadline else None

AppConfig

Bases: BaseModel

Typed configuration for the Mewbo runtime.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3751
3752
3753
3754
3755
3756
3757
3758
3759
3760
3761
3762
3763
3764
3765
3766
3767
3768
3769
3770
3771
3772
3773
3774
3775
3776
3777
3778
3779
3780
3781
3782
3783
3784
3785
3786
3787
3788
3789
3790
3791
3792
3793
3794
3795
3796
3797
3798
3799
3800
3801
3802
3803
3804
3805
3806
3807
3808
3809
3810
3811
3812
3813
3814
3815
3816
3817
3818
3819
3820
3821
3822
3823
3824
3825
3826
3827
3828
3829
3830
3831
3832
3833
3834
3835
3836
3837
3838
3839
3840
3841
3842
3843
3844
3845
3846
3847
3848
3849
3850
3851
3852
3853
3854
3855
3856
3857
3858
3859
3860
3861
3862
3863
3864
3865
3866
3867
3868
3869
3870
3871
3872
3873
3874
3875
3876
3877
3878
3879
3880
3881
3882
3883
3884
3885
3886
3887
3888
3889
3890
3891
3892
3893
3894
3895
3896
3897
3898
3899
3900
3901
3902
3903
3904
3905
3906
3907
3908
3909
3910
3911
3912
3913
3914
3915
3916
3917
3918
3919
3920
3921
3922
3923
3924
3925
3926
3927
3928
3929
3930
3931
3932
3933
3934
3935
3936
3937
3938
3939
3940
3941
3942
3943
3944
3945
3946
3947
3948
3949
3950
3951
3952
3953
3954
3955
3956
3957
3958
3959
3960
3961
3962
3963
3964
3965
3966
3967
3968
3969
3970
3971
3972
3973
3974
3975
3976
3977
3978
3979
3980
3981
3982
3983
3984
3985
3986
3987
3988
3989
3990
3991
3992
3993
3994
3995
3996
3997
3998
3999
4000
4001
4002
4003
4004
4005
4006
4007
4008
4009
4010
4011
4012
4013
4014
4015
4016
4017
4018
4019
4020
4021
4022
4023
4024
4025
4026
4027
4028
4029
4030
4031
4032
4033
4034
4035
4036
4037
4038
4039
4040
4041
4042
4043
4044
4045
4046
4047
4048
4049
4050
4051
4052
4053
4054
4055
4056
4057
4058
4059
4060
4061
4062
4063
4064
4065
4066
4067
4068
4069
4070
4071
4072
4073
4074
4075
4076
4077
4078
4079
4080
4081
4082
4083
4084
4085
4086
4087
4088
4089
4090
4091
4092
4093
4094
4095
4096
4097
4098
4099
4100
4101
4102
4103
4104
4105
4106
4107
4108
4109
4110
4111
4112
4113
4114
4115
4116
4117
4118
4119
4120
4121
4122
4123
4124
4125
4126
4127
4128
4129
4130
4131
4132
4133
4134
4135
4136
4137
4138
4139
4140
4141
4142
4143
4144
4145
4146
4147
4148
4149
4150
4151
4152
4153
4154
4155
4156
4157
4158
4159
4160
4161
4162
4163
4164
4165
4166
4167
4168
4169
4170
4171
4172
4173
4174
4175
4176
4177
4178
4179
4180
4181
4182
4183
4184
4185
4186
4187
4188
4189
4190
4191
4192
4193
4194
4195
4196
4197
4198
4199
4200
4201
4202
4203
4204
4205
4206
4207
4208
4209
4210
4211
4212
4213
4214
4215
4216
4217
4218
4219
4220
4221
4222
4223
4224
4225
4226
4227
4228
4229
4230
4231
4232
4233
4234
4235
4236
4237
4238
4239
4240
4241
4242
4243
4244
class AppConfig(BaseModel):
    """Typed configuration for the Mewbo runtime."""

    model_config = ConfigDict(extra="ignore", validate_default=True)

    #: Set once the first non-atomic save degrades, so the warning that
    #: explains it is logged one time per process rather than per save.
    _warned_non_atomic: ClassVar[bool] = False

    #: Retired keys already announced, so a config validated on every settings
    #: PATCH repeats itself once per key rather than once per request.
    _warned_retired: ClassVar[set[str]] = set()

    runtime: RuntimeConfig = Field(
        default_factory=_runtime_config_default, description="Runtime environment settings."
    )
    storage: StorageConfig = Field(
        default_factory=_storage_config_default,
        description="Session storage backend (json or mongodb).",
    )
    llm: LLMConfig = Field(
        default_factory=_llm_config_default,
        description="LLM provider connection and model selection.",
    )
    context: ContextConfig = Field(
        default_factory=_context_config_default,
        description="Context window selection and event filtering.",
    )
    speech: SpeechConfig = Field(
        default_factory=_speech_config_default,
        description="Speech models for reading answers aloud and for dictation.",
    )
    token_budget: TokenBudgetConfig = Field(
        default_factory=_token_budget_config_default,
        description="Token budget and auto-compaction thresholds.",
    )
    compaction: CompactionConfig = Field(
        default_factory=_compaction_config_default,
        description="Conversation compaction prompt selection (caveman mode).",
    )
    langfuse: LangfuseConfig = Field(
        default_factory=_langfuse_config_default,
        description="Langfuse LLM observability integration.",
    )
    home_assistant: HomeAssistantConfig = Field(
        default_factory=_home_assistant_config_default,
        description="Home Assistant smart-home integration.",
    )
    permissions: PermissionsConfig = Field(
        default_factory=_permissions_config_default,
        description="Tool execution permission policy.",
    )
    safety: SafetyConfig = Field(
        default_factory=_safety_config_default,
        description="Operator-owned tool-call gate and session observer. Off by default.",
    )
    cli: CLIConfig = Field(
        default_factory=_cli_config_default,
        description="Terminal CLI display and interaction settings.",
    )
    api: APIConfig = Field(
        default_factory=_api_config_default, description="REST API authentication."
    )
    agent: AgentConfig = Field(
        default_factory=_agent_config_default, description="Sub-agent hypervisor settings."
    )
    wiki: WikiConfig = Field(
        default_factory=_wiki_config_default,
        description="Wiki subsystem defaults and embedding behavior.",
    )
    scg: ScgConfig = Field(
        default_factory=_scg_config_default,
        description="Source Capability Graph feature gate and search-tier default.",
    )
    hooks: HooksConfig = Field(
        default_factory=_hooks_config_default,
        description="External hooks fired during the session lifecycle (command or http).",
    )
    plugins: PluginsConfig = Field(
        default_factory=_plugins_config_default,
        description=(
            "How Mewbo finds, installs, and enables plugins.\n\n"
            "A plugin extends what the agent can do. It brings its own skills, agent "
            "definitions, lifecycle hooks, and MCP tools, and those hooks run on this "
            "machine, so install only what you trust. Marketplaces are the catalogs "
            "Mewbo looks in for plugins to offer you; the list above shows what is "
            "installed today and what is available to install."
        ),
    )
    triggers: TriggersConfig = Field(
        default_factory=_triggers_config_default,
        description=(
            "Ceilings on the triggers that start a session without you.\n\n"
            "A trigger is a reverse invocation. Instead of you opening a session, a "
            "session asks to be woken later and Mewbo wakes it, on its own, with "
            "nobody watching. There are five kinds: `time.at` fires once at a set "
            "moment, `time.cron` fires on a repeating schedule, `ci.workflow` fires "
            "when a CI run finishes, `forge.pr` fires when a pull request changes, "
            "and `webhook` fires when something outside calls in. The agent arms a "
            "trigger from inside a session, and the triggers list above is where you "
            "watch, pause, and cancel what it armed. The settings here are the limits "
            "all of that has to stay inside."
        ),
    )
    channels: dict[str, dict[str, Any]] = Field(
        default_factory=dict,
        description="Channel adapters, keyed by name (nextcloud-talk, email).",
        json_schema_extra={"x-group": "integrations", "x-order": 4},
    )
    projects: dict[str, ProjectConfig] = Field(
        default_factory=_projects_config_default,
        description=(
            "Directories you've already created, registered here by hand under "
            'a short name; sessions reference them by that name (e.g. `"'
            'project": "<key>"`). Distinct from Mewbo-managed ("virtual") '
            "projects, which the API creates and owns itself and which sessions "
            "reference as `managed:<project_id>`: entries here are never "
            "created, modified, or deleted by Mewbo, only pointed at."
        ),
        json_schema_extra={"x-group": "workspace", "x-order": 1},
    )

    @model_validator(mode="before")
    @classmethod
    def _prepare_document(cls, payload: Any) -> Any:
        """Drop retired keys, then resolve environment references — in that order.

        Both steps are here, sequenced by hand, rather than in a validator each.
        Pydantic runs ``mode="before"`` model validators bottom-up, so two of
        them would put this ordering in the DECLARATION ORDER of two methods —
        invisible, and silently reversed by anyone who moves one. The order is
        load-bearing: a retired key holding a reference to a variable nobody
        sets any more must not refuse the boot, when the whole point of that key
        is that it is ignored.
        """
        if not isinstance(payload, dict):
            return payload
        return EnvRef.resolve_document(cls._without_retired_keys(payload), os.environ)

    @classmethod
    def _without_retired_keys(cls, payload: dict[str, Any]) -> dict[str, Any]:
        """Return ``payload`` minus every retired key, announcing each one once.

        Prunes into a copy, and that is load-bearing rather than defensive
        habit. The settings PATCH path validates the operator's merged document
        and then persists THAT SAME dict — so a validator pruning in place
        would delete lines out of somebody's ``app.json`` as a side effect of
        an unrelated save, silently and with no record that the key was ever
        set. The pruning exists to let the process start, not to edit the file;
        removing the line stays the operator's own act.
        """
        if not _RETIRED_KEYS:
            return payload
        pruned = payload
        for retired in _RETIRED_KEYS:
            pruned, was_set = retired.prune(pruned)
            if not was_set:
                continue
            if retired.dotted in cls._warned_retired:
                continue
            cls._warned_retired.add(retired.dotted)
            _logger.warning(
                "Configuration key '%s' has been retired and is ignored. %s "
                "Remove it from your app.json to silence this warning.",
                retired.dotted,
                retired.guidance,
            )
        return pruned

    @field_validator("projects", mode="before")
    @classmethod
    def _normalize_projects(cls, value: Any) -> dict[str, ProjectConfig]:
        if not isinstance(value, dict):
            return {}
        # Mirrors ``mewbo_core.workspaces.project_catalog.AUTO_PROJECT`` by VALUE, not
        # import: ``project_catalog`` imports ``ProjectConfig`` from this
        # module, and its own import chain calls back into ``get_config()``
        # (``project_store`` -> ``worktree``'s module-level logger setup ->
        # ``common._configure_logging``) before it finishes loading — even a
        # lazy import here deadlocks on the FIRST ``AppConfig()`` built in a
        # fresh process (``validate_default=True`` runs this on every
        # construction, including the default). Kept in lockstep by
        # ``TestAppConfigNormalizeProjects.test_a_project_named_auto_is_refused``,
        # which asserts this literal against the canonical constant directly.
        reserved_auto_project_name = "auto"
        result: dict[str, ProjectConfig] = {}
        for name, cfg in value.items():
            key = str(name)
            if key == reserved_auto_project_name:
                raise ValueError(
                    f"A configured project cannot be named '{reserved_auto_project_name}' "
                    "— that name is reserved for the auto-select sentinel. Choose a "
                    "different project name."
                )
            if isinstance(cfg, dict):
                result[key] = ProjectConfig.model_validate(cfg)
            elif isinstance(cfg, ProjectConfig):
                result[key] = cfg
        return result

    @model_validator(mode="after")
    def _warn_when_the_ladder_is_unreachable(self) -> AppConfig:
        """Apply the budget law to the REAL ladder, which only this level can see.

        ``AgentConfig`` can check one fallback rung and no more — the ladder is
        declared under ``llm``, and a model that reached across the config tree
        to read it would be the coupling the layering rules exist to prevent. So
        the floor check lives there and the full check lives here, where both
        halves are in hand.

        The distinction is not academic: a deployment can size ``turn_deadline``
        so the primary plus ONE fallback fits, pass that check, and still have
        every rung after the second be dead configuration. The chain reaches
        them, finds the wall clock spent, and stops — silently, since an
        unreachable rung produces no event of its own.

        Only fires when the floor check passed, so one misconfiguration never
        logs twice. Reports rather than refuses, for the reason ``AgentConfig``
        gives.
        """
        fallbacks = len(self.llm.effective_fallback_models())
        if fallbacks < 2 or self.agent.unreachable_rung(1) is not None:
            return self  # nothing more to say than the floor check already said
        breach = self.agent.unreachable_rung(fallbacks)
        if breach is not None:
            needed, deadline = breach
            _logger.warning(
                "agent.retry.turn_deadline=%.0fs cannot reach the whole model ladder: "
                "%d fallback model(s) after the primary may spend %.0fs in total "
                "(llm_call_timeout=%.0fs, llm_call_retries=%d, "
                "retry.fallback_retries=%d). The last rungs are unreachable and "
                "cross-model fallback stops short of them. Raise turn_deadline to "
                ">= %.0fs, or shorten the ladder.",
                deadline,
                fallbacks,
                needed,
                self.agent.llm_call_timeout,
                self.agent.llm_call_retries,
                self.agent.retry.fallback_retries,
                needed,
            )
        return self

    @classmethod
    def load(cls, path: str | Path) -> AppConfig:
        """Load configuration from a JSON file."""
        payload = _load_json(path)
        return cls.model_validate(payload)

    def to_json(self, *, indent: int = 2) -> str:
        """Serialize config to JSON."""
        return self.model_dump_json(indent=indent, exclude_none=True)

    def write(self, path: str | Path, *, indent: int = 2) -> None:
        """Atomically persist THIS MODEL as the whole config file.

        Renders every declared field, so an unset one is written out at its
        resolved default. That suits a caller scaffolding a fresh config; it is
        the wrong tool for saving an edit to a file someone already maintains —
        use :meth:`write_document` for that and see the warning there.
        """
        self.write_document(path, json.loads(self.to_json(indent=indent)), indent=indent)

    @classmethod
    def write_document(
        cls,
        path: str | Path,
        document: Mapping[str, Any],
        *,
        indent: int = 2,
    ) -> None:
        """Atomically persist an already-shaped config document.

        Editing a config means writing back the operator's OWN document with
        their change applied — never a re-rendering of the validated model.
        Re-rendering looks equivalent and is not, in two ways that both corrupt
        a shared file. It converts every unset field into an explicit default,
        so "leave this to the runtime" silently becomes "pin this forever". And
        because several defaults are derived from the environment of whichever
        process happens to save (``MEWBO_HOME`` feeding the ``runtime.*``
        directories, the storage driver's URI), a save from inside a container
        rewrites the shared file with container-only absolute paths and
        hostnames, which then breaks every other consumer of it. The model is
        still the validator — nothing unvalidated reaches disk — it is just not
        the thing serialized. This also keeps a key the model does not declare
        (``$schema``, a block for a feature with no typed field yet) instead of
        dropping it, since ``extra="ignore"`` discards those at validation.
        """
        cls._write_atomically(Path(path), json.dumps(document, indent=indent) + "\n")

    @classmethod
    def _write_atomically(cls, target: Path, payload: str) -> None:
        """Stage ``payload`` beside ``target`` and swap it in, or raise ``ConfigWriteError``.

        A plain ``write_text`` truncates the target before it writes, so a
        failure part-way through (out of space, killed process) leaves the live
        config truncated or half-written. The payload is staged in a temporary
        file in the SAME directory — same filesystem, which is what makes
        ``os.replace`` atomic — and only swapped in once it is fully on disk,
        so a failed save leaves the previous config byte-for-byte intact.

        Some deployments make that impossible; those degrade to a non-atomic
        in-place write rather than failing. See ``_in_place_fallback_applies``.
        """
        tmp_path: Path | None = None
        try:
            target.parent.mkdir(parents=True, exist_ok=True)
            handle, name = tempfile.mkstemp(
                dir=target.parent, prefix=f".{target.name}.", suffix=".tmp"
            )
            tmp_path = Path(name)
            with os.fdopen(handle, "w", encoding="utf-8") as stream:
                stream.write(payload)
                stream.flush()
                os.fsync(stream.fileno())
            # ``mkstemp`` creates owner-only; the deployed stack has more than
            # one process reading this file, so keep whatever mode the operator
            # already has on it and fall back to the umask-default of the
            # ``write_text`` this replaced.
            os.chmod(tmp_path, target.stat().st_mode & 0o777 if target.exists() else 0o644)
            os.replace(tmp_path, target)
        except OSError as exc:
            if tmp_path is not None:
                # Cleanup failures must not mask the real error, and must not
                # turn a save the fallback below completed into a reported one.
                with contextlib.suppress(OSError):
                    tmp_path.unlink(missing_ok=True)
            if cls._in_place_fallback_applies(target, exc):
                cls._write_in_place(target, payload)
                return
            raise ConfigWriteError.from_oserror(target, exc) from exc

    @classmethod
    def _in_place_fallback_applies(cls, target: Path, exc: OSError) -> bool:
        """Whether a failed atomic write may degrade to writing ``target`` directly.

        WHY: a rename cannot replace a path that is ITSELF a bind-mount point —
        it fails with EBUSY — and a single-file mount can equally leave the
        containing directory read-only while the file stays writable. In both
        shapes the staged-file dance can never work however healthy the
        filesystem is, and writing the target directly is the only way to save
        at all. Both gates are load-bearing: the errno gate keeps a full or
        failing disk OUT (an in-place write truncates first, so degrading there
        would destroy the config on the way to failing anyway), and the
        writability gate keeps a genuinely read-only deployment out, so that
        one still gets its honest ``read_only`` refusal.
        """
        return ConfigWriteError.is_structural(exc) and os.access(target, os.W_OK)

    @classmethod
    def _write_in_place(cls, target: Path, payload: str) -> None:
        """Write ``payload`` straight onto ``target`` — the non-atomic last resort."""
        try:
            with open(target, "w", encoding="utf-8") as stream:
                stream.write(payload)
                stream.flush()
                os.fsync(stream.fileno())
        except OSError as exc:
            raise ConfigWriteError.from_oserror(target, exc) from exc
        if not cls._warned_non_atomic:
            # Once per process: the shape of the deployment does not change
            # between saves, so repeating this per save is pure log noise.
            AppConfig._warned_non_atomic = True
            _logger.warning(
                "Configuration saved without atomic replacement: %s cannot be replaced by "
                "rename, which is how a single-file bind mount of the config presents. A "
                "crash mid-write can truncate it; mount the containing directory rather "
                "than the file to restore atomic saves.",
                target,
            )

    @classmethod
    def probe_write_access(cls, path: str | Path) -> ConfigWriteAccess:
        """Report whether ``path`` could be persisted to, without modifying it.

        The probe creates and removes a temporary file in the directory an
        atomic :meth:`write` would stage into — exactly the permission that
        write needs — so it never opens, truncates or replaces an existing
        config. A missing parent is probed at its nearest existing ancestor
        (where ``mkdir`` would have to write) rather than being created, so the
        probe itself has no side effects at all.

        It answers for the SAME rungs :meth:`write` will try, fallback
        included. Reporting a directory-only verdict would say "not writable"
        for a single-file-mounted config that saves perfectly well in place,
        and a console that disables its Save button on this would then be
        refusing a write the server would have accepted.
        """
        target = Path(path)
        probe_dir = target.parent
        while not probe_dir.exists() and probe_dir.parent != probe_dir:
            probe_dir = probe_dir.parent
        try:
            handle, name = tempfile.mkstemp(
                dir=probe_dir, prefix=f".{target.name}.", suffix=".probe"
            )
            os.close(handle)
            os.unlink(name)
        except OSError as exc:
            if cls._in_place_fallback_applies(target, exc):
                return ConfigWriteAccess(writable=True)
            code, reason = ConfigWriteError.classify(exc)
            return ConfigWriteAccess(writable=False, code=code, reason=reason)
        except Exception:
            # A probe that raises is a probe nobody can call from a request
            # handler; an unexpected failure is still "cannot persist".
            _logger.warning("Config write probe failed for %s", target, exc_info=True)
            return ConfigWriteAccess(
                writable=False, code="io_error", reason=ConfigWriteError.reason_for("io_error")
            )
        return ConfigWriteAccess(writable=True)

    async def preflight(self, *, disable_on_failure: bool = True) -> dict[str, dict[str, Any]]:
        """Run async validation checks for optional integrations."""
        results: dict[str, ConfigCheck] = {}

        async def _llm_check() -> ConfigCheck:
            return await asyncio.to_thread(self.llm.validate_models)

        async def _langfuse_check() -> ConfigCheck:
            enabled, reason, metadata = self.langfuse.evaluate()
            if not enabled:
                return ConfigCheck(
                    name="langfuse",
                    enabled=False,
                    ok=True,
                    reason=reason,
                    metadata=metadata,
                )
            try:
                host = self.langfuse.host.rstrip("/")
                if host:
                    await asyncio.to_thread(_probe_http, f"{host}/api/public/health")
                return ConfigCheck(name="langfuse", enabled=True, ok=True)
            except ValueError as exc:
                return ConfigCheck(name="langfuse", enabled=True, ok=False, reason=str(exc))

        async def _ha_check() -> ConfigCheck:
            enabled, reason, metadata = self.home_assistant.evaluate()
            if not enabled:
                return ConfigCheck(
                    name="home_assistant",
                    enabled=False,
                    ok=True,
                    reason=reason,
                    metadata=metadata,
                )
            try:
                url = self.home_assistant.url.rstrip("/")
                headers = {"Authorization": f"Bearer {self.home_assistant.token}"}
                await asyncio.to_thread(_probe_http, f"{url}/api/config", headers=headers)
                return ConfigCheck(name="home_assistant", enabled=True, ok=True)
            except ValueError as exc:
                return ConfigCheck(name="home_assistant", enabled=True, ok=False, reason=str(exc))

        async def _mcp_check() -> ConfigCheck:
            config_path = get_mcp_config_path()
            if not config_path:
                return ConfigCheck(name="mcp", enabled=False, ok=True, reason="mcp config disabled")
            try:
                from mewbo_tools.integration import mcp as mcp_module

                config = mcp_module._load_mcp_config(config_path)
                tools, failures = await asyncio.to_thread(
                    mcp_module.discover_mcp_tool_details_with_failures, config
                )
                if failures:
                    return ConfigCheck(
                        name="mcp",
                        enabled=True,
                        ok=False,
                        reason="mcp discovery failed",
                        metadata={"failures": {k: str(v) for k, v in failures.items()}},
                    )
                return ConfigCheck(
                    name="mcp",
                    enabled=True,
                    ok=True,
                    metadata={"servers": list(tools.keys())},
                )
            except Exception as exc:
                return ConfigCheck(name="mcp", enabled=True, ok=False, reason=str(exc))

        checks = await asyncio.gather(_llm_check(), _langfuse_check(), _ha_check(), _mcp_check())
        for check in checks:
            results[check.name] = check
        if disable_on_failure:
            langfuse_check = results.get("langfuse")
            if langfuse_check and not langfuse_check.ok and self.langfuse.enabled:
                self.langfuse.enabled = False
            ha_check = results.get("home_assistant")
            if ha_check and not ha_check.ok and self.home_assistant.enabled:
                self.home_assistant.enabled = False
        return {name: check.to_dict() for name, check in results.items()}

load(path: str | Path) -> AppConfig classmethod

Load configuration from a JSON file.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3994
3995
3996
3997
3998
@classmethod
def load(cls, path: str | Path) -> AppConfig:
    """Load configuration from a JSON file."""
    payload = _load_json(path)
    return cls.model_validate(payload)

preflight(*, disable_on_failure: bool = True) -> dict[str, dict[str, Any]] async

Run async validation checks for optional integrations.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4163
4164
4165
4166
4167
4168
4169
4170
4171
4172
4173
4174
4175
4176
4177
4178
4179
4180
4181
4182
4183
4184
4185
4186
4187
4188
4189
4190
4191
4192
4193
4194
4195
4196
4197
4198
4199
4200
4201
4202
4203
4204
4205
4206
4207
4208
4209
4210
4211
4212
4213
4214
4215
4216
4217
4218
4219
4220
4221
4222
4223
4224
4225
4226
4227
4228
4229
4230
4231
4232
4233
4234
4235
4236
4237
4238
4239
4240
4241
4242
4243
4244
async def preflight(self, *, disable_on_failure: bool = True) -> dict[str, dict[str, Any]]:
    """Run async validation checks for optional integrations."""
    results: dict[str, ConfigCheck] = {}

    async def _llm_check() -> ConfigCheck:
        return await asyncio.to_thread(self.llm.validate_models)

    async def _langfuse_check() -> ConfigCheck:
        enabled, reason, metadata = self.langfuse.evaluate()
        if not enabled:
            return ConfigCheck(
                name="langfuse",
                enabled=False,
                ok=True,
                reason=reason,
                metadata=metadata,
            )
        try:
            host = self.langfuse.host.rstrip("/")
            if host:
                await asyncio.to_thread(_probe_http, f"{host}/api/public/health")
            return ConfigCheck(name="langfuse", enabled=True, ok=True)
        except ValueError as exc:
            return ConfigCheck(name="langfuse", enabled=True, ok=False, reason=str(exc))

    async def _ha_check() -> ConfigCheck:
        enabled, reason, metadata = self.home_assistant.evaluate()
        if not enabled:
            return ConfigCheck(
                name="home_assistant",
                enabled=False,
                ok=True,
                reason=reason,
                metadata=metadata,
            )
        try:
            url = self.home_assistant.url.rstrip("/")
            headers = {"Authorization": f"Bearer {self.home_assistant.token}"}
            await asyncio.to_thread(_probe_http, f"{url}/api/config", headers=headers)
            return ConfigCheck(name="home_assistant", enabled=True, ok=True)
        except ValueError as exc:
            return ConfigCheck(name="home_assistant", enabled=True, ok=False, reason=str(exc))

    async def _mcp_check() -> ConfigCheck:
        config_path = get_mcp_config_path()
        if not config_path:
            return ConfigCheck(name="mcp", enabled=False, ok=True, reason="mcp config disabled")
        try:
            from mewbo_tools.integration import mcp as mcp_module

            config = mcp_module._load_mcp_config(config_path)
            tools, failures = await asyncio.to_thread(
                mcp_module.discover_mcp_tool_details_with_failures, config
            )
            if failures:
                return ConfigCheck(
                    name="mcp",
                    enabled=True,
                    ok=False,
                    reason="mcp discovery failed",
                    metadata={"failures": {k: str(v) for k, v in failures.items()}},
                )
            return ConfigCheck(
                name="mcp",
                enabled=True,
                ok=True,
                metadata={"servers": list(tools.keys())},
            )
        except Exception as exc:
            return ConfigCheck(name="mcp", enabled=True, ok=False, reason=str(exc))

    checks = await asyncio.gather(_llm_check(), _langfuse_check(), _ha_check(), _mcp_check())
    for check in checks:
        results[check.name] = check
    if disable_on_failure:
        langfuse_check = results.get("langfuse")
        if langfuse_check and not langfuse_check.ok and self.langfuse.enabled:
            self.langfuse.enabled = False
        ha_check = results.get("home_assistant")
        if ha_check and not ha_check.ok and self.home_assistant.enabled:
            self.home_assistant.enabled = False
    return {name: check.to_dict() for name, check in results.items()}

probe_write_access(path: str | Path) -> ConfigWriteAccess classmethod

Report whether path could be persisted to, without modifying it.

The probe creates and removes a temporary file in the directory an atomic :meth:write would stage into — exactly the permission that write needs — so it never opens, truncates or replaces an existing config. A missing parent is probed at its nearest existing ancestor (where mkdir would have to write) rather than being created, so the probe itself has no side effects at all.

It answers for the SAME rungs :meth:write will try, fallback included. Reporting a directory-only verdict would say "not writable" for a single-file-mounted config that saves perfectly well in place, and a console that disables its Save button on this would then be refusing a write the server would have accepted.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4122
4123
4124
4125
4126
4127
4128
4129
4130
4131
4132
4133
4134
4135
4136
4137
4138
4139
4140
4141
4142
4143
4144
4145
4146
4147
4148
4149
4150
4151
4152
4153
4154
4155
4156
4157
4158
4159
4160
4161
@classmethod
def probe_write_access(cls, path: str | Path) -> ConfigWriteAccess:
    """Report whether ``path`` could be persisted to, without modifying it.

    The probe creates and removes a temporary file in the directory an
    atomic :meth:`write` would stage into — exactly the permission that
    write needs — so it never opens, truncates or replaces an existing
    config. A missing parent is probed at its nearest existing ancestor
    (where ``mkdir`` would have to write) rather than being created, so the
    probe itself has no side effects at all.

    It answers for the SAME rungs :meth:`write` will try, fallback
    included. Reporting a directory-only verdict would say "not writable"
    for a single-file-mounted config that saves perfectly well in place,
    and a console that disables its Save button on this would then be
    refusing a write the server would have accepted.
    """
    target = Path(path)
    probe_dir = target.parent
    while not probe_dir.exists() and probe_dir.parent != probe_dir:
        probe_dir = probe_dir.parent
    try:
        handle, name = tempfile.mkstemp(
            dir=probe_dir, prefix=f".{target.name}.", suffix=".probe"
        )
        os.close(handle)
        os.unlink(name)
    except OSError as exc:
        if cls._in_place_fallback_applies(target, exc):
            return ConfigWriteAccess(writable=True)
        code, reason = ConfigWriteError.classify(exc)
        return ConfigWriteAccess(writable=False, code=code, reason=reason)
    except Exception:
        # A probe that raises is a probe nobody can call from a request
        # handler; an unexpected failure is still "cannot persist".
        _logger.warning("Config write probe failed for %s", target, exc_info=True)
        return ConfigWriteAccess(
            writable=False, code="io_error", reason=ConfigWriteError.reason_for("io_error")
        )
    return ConfigWriteAccess(writable=True)

to_json(*, indent: int = 2) -> str

Serialize config to JSON.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4000
4001
4002
def to_json(self, *, indent: int = 2) -> str:
    """Serialize config to JSON."""
    return self.model_dump_json(indent=indent, exclude_none=True)

write(path: str | Path, *, indent: int = 2) -> None

Atomically persist THIS MODEL as the whole config file.

Renders every declared field, so an unset one is written out at its resolved default. That suits a caller scaffolding a fresh config; it is the wrong tool for saving an edit to a file someone already maintains — use :meth:write_document for that and see the warning there.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4004
4005
4006
4007
4008
4009
4010
4011
4012
def write(self, path: str | Path, *, indent: int = 2) -> None:
    """Atomically persist THIS MODEL as the whole config file.

    Renders every declared field, so an unset one is written out at its
    resolved default. That suits a caller scaffolding a fresh config; it is
    the wrong tool for saving an edit to a file someone already maintains —
    use :meth:`write_document` for that and see the warning there.
    """
    self.write_document(path, json.loads(self.to_json(indent=indent)), indent=indent)

write_document(path: str | Path, document: Mapping[str, Any], *, indent: int = 2) -> None classmethod

Atomically persist an already-shaped config document.

Editing a config means writing back the operator's OWN document with their change applied — never a re-rendering of the validated model. Re-rendering looks equivalent and is not, in two ways that both corrupt a shared file. It converts every unset field into an explicit default, so "leave this to the runtime" silently becomes "pin this forever". And because several defaults are derived from the environment of whichever process happens to save (MEWBO_HOME feeding the runtime.* directories, the storage driver's URI), a save from inside a container rewrites the shared file with container-only absolute paths and hostnames, which then breaks every other consumer of it. The model is still the validator — nothing unvalidated reaches disk — it is just not the thing serialized. This also keeps a key the model does not declare ($schema, a block for a feature with no typed field yet) instead of dropping it, since extra="ignore" discards those at validation.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4014
4015
4016
4017
4018
4019
4020
4021
4022
4023
4024
4025
4026
4027
4028
4029
4030
4031
4032
4033
4034
4035
4036
4037
4038
4039
@classmethod
def write_document(
    cls,
    path: str | Path,
    document: Mapping[str, Any],
    *,
    indent: int = 2,
) -> None:
    """Atomically persist an already-shaped config document.

    Editing a config means writing back the operator's OWN document with
    their change applied — never a re-rendering of the validated model.
    Re-rendering looks equivalent and is not, in two ways that both corrupt
    a shared file. It converts every unset field into an explicit default,
    so "leave this to the runtime" silently becomes "pin this forever". And
    because several defaults are derived from the environment of whichever
    process happens to save (``MEWBO_HOME`` feeding the ``runtime.*``
    directories, the storage driver's URI), a save from inside a container
    rewrites the shared file with container-only absolute paths and
    hostnames, which then breaks every other consumer of it. The model is
    still the validator — nothing unvalidated reaches disk — it is just not
    the thing serialized. This also keeps a key the model does not declare
    (``$schema``, a block for a feature with no typed field yet) instead of
    dropping it, since ``extra="ignore"`` discards those at validation.
    """
    cls._write_atomically(Path(path), json.dumps(document, indent=indent) + "\n")

AuthenticatorEntry

Bases: BaseModel

One identity source in api.auth.authenticators.

Deliberately PERMISSIVE (extra="allow"): the authoritative model is the identity kernel's discriminated authenticator union, which sits a layer above core and must never be imported down into it. Every kind-specific setting therefore rides through here untyped and is re-validated STRICTLY — per-kind, extra="forbid" — when the server builds its auth settings at startup, which refuses to boot on an invalid entry. So this model is not a second validator and must not grow into one.

What it DOES declare is the plaintext credentials, because the config API redacts by SCHEMA: a field carries x-secret or its value is returned to every caller holding config.read. An untyped dict contributes no schema, so nothing marked these and they were served in the clear. Core knows these two NAMES — a stable wire contract it shares with the identity kernel — without knowing which kind each belongs to or what it means, which is precisely the sliver of knowledge redaction needs and no more. kind/name stay typed so an entry is self-describing.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
class AuthenticatorEntry(BaseModel):
    """One identity source in ``api.auth.authenticators``.

    Deliberately PERMISSIVE (``extra="allow"``): the authoritative model is the
    identity kernel's discriminated authenticator union, which sits a layer
    above core and must never be imported down into it. Every kind-specific
    setting therefore rides through here untyped and is re-validated STRICTLY —
    per-kind, ``extra="forbid"`` — when the server builds its auth settings at
    startup, which refuses to boot on an invalid entry. So this model is not a
    second validator and must not grow into one.

    What it DOES declare is the plaintext credentials, because the config API
    redacts by SCHEMA: a field carries ``x-secret`` or its value is returned to
    every caller holding ``config.read``. An untyped ``dict`` contributes no
    schema, so nothing marked these and they were served in the clear. Core
    knows these two NAMES — a stable wire contract it shares with the identity
    kernel — without knowing which kind each belongs to or what it means, which
    is precisely the sliver of knowledge redaction needs and no more.
    ``kind``/``name`` stay typed so an entry is self-describing.
    """

    model_config = ConfigDict(extra="allow", json_schema_extra={"title": "Authenticator"})

    name: str = Field(
        ...,
        description="Unique label for this identity source, used in logs and the admin UI.",
    )
    kind: str = Field(
        ...,
        description=(
            "Which authenticator type this entry configures: `api_key`, `oidc`, "
            "`trusted_header`, `ldap`, or `saml`. Selects which further settings "
            "the entry must carry."
        ),
    )
    client_secret: str | None = Field(
        None,
        description=(
            "OIDC client secret issued by the identity provider. "
            "Write-only: never returned by the config API."
        ),
        json_schema_extra={"x-secret": True},
    )
    bind_password: str | None = Field(
        None,
        description=(
            "Password for the LDAP service account used to search the directory. "
            "Write-only: never returned by the config API."
        ),
        json_schema_extra={"x-secret": True},
    )

    @model_serializer(mode="wrap")
    def _drop_unset_credentials(self, handler: SerializerFunctionWrapHandler) -> dict[str, Any]:
        """Serialize, omitting the credential fields this entry's kind lacks.

        Declaring both credentials on one model means a plain dump stamps
        ``bind_password: None`` onto an OIDC entry and ``client_secret: None``
        onto an LDAP one. That dump is fed straight back into the strict
        per-kind union at startup, where an unexpected key is a hard boot
        failure — so a null here is not cosmetic noise, it would stop the
        server. Absent and null must stay distinguishable for these two.
        """
        data = handler(self)
        for credential in ("client_secret", "bind_password"):
            if data.get(credential) is None:
                data.pop(credential, None)
        return data

CLIConfig

Bases: BaseModel

Terminal CLI display and interaction settings.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
class CLIConfig(BaseModel):
    """Terminal CLI display and interaction settings."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "CLI", "x-group": "interface", "x-order": 1},
    )

    disable_textual: bool = Field(
        False,
        description="Disable the Textual TUI and fall back to plain Rich output.",
    )
    remote: CliRemoteConfig = Field(
        default_factory=lambda: CliRemoteConfig.model_validate({}),
        description="Opt-in remote session sync + product tools (CLI-only).",
    )
    @field_validator("disable_textual", mode="before")
    @classmethod
    def _normalize_disable_textual(cls, value: Any) -> bool:
        return _coerce_bool(value, default=False)

CliRemoteConfig

Bases: BaseModel

Opt-in remote endpoint for the terminal CLI; CLI-scoped ONLY.

Every other surface ignores this block. When base_url is set the CLI is still a strictly-local engine (the run loop + authoritative JSONL transcript stay on this host), but it additionally (a) mirrors each session event to the remote REST API fire-and-forget for cross-device visibility and (b) auto-registers the Mewbo MCP server so the product tools (ask_wiki/search/structured_query + wiki-graph reads) appear in the CLI registry and execute remotely (compute offload). Empty ⇒ fully local.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
class CliRemoteConfig(BaseModel):
    """Opt-in remote endpoint for the terminal CLI; CLI-scoped ONLY.

    Every other surface ignores this block. When ``base_url`` is set the CLI is
    still a strictly-local engine (the run loop + authoritative JSONL transcript
    stay on this host), but it additionally (a) mirrors each session event to the
    remote REST API fire-and-forget for cross-device visibility and (b)
    auto-registers the Mewbo MCP server so the product tools
    (``ask_wiki``/``search``/``structured_query`` + wiki-graph reads) appear in
    the CLI registry and execute remotely (compute offload). Empty ⇒ fully local.
    """

    model_config = ConfigDict(
        extra="forbid",
        validate_default=True,
        json_schema_extra={"title": "CLI Remote"},
    )

    base_url: str = Field(
        "",
        description=(
            "Base URL of the remote Mewbo deployment, a reverse-proxy root that "
            "serves the REST API under ``/api`` and the Mewbo MCP server under "
            "``/mcp``. Empty (default) ⇒ the CLI is fully local."
        ),
        examples=["https://mewbo.example.com"],
    )
    token: str = Field(
        "",
        description=(
            "API token presented to the remote deployment: sent as ``X-API-Key`` "
            "to the REST API for transcript sync and as a ``Bearer`` token to the "
            "Mewbo MCP server (which forwards it to the REST API). ``${ENV_VAR}`` "
            "references are expanded by the CLI at use time."
        ),
        json_schema_extra={"x-secret": True},
    )

    @property
    def enabled(self) -> bool:
        """True when a remote base URL is configured (sync + product tools on)."""
        return bool(self.base_url.strip())

enabled: bool property

True when a remote base URL is configured (sync + product tools on).

CompactionConfig

Bases: BaseModel

Summarization prompt selection for conversation compaction.

caveman_mode enables a rule-augmented "caveman" prompt that instructs the summarizer LLM to drop articles, filler, pleasantries, and hedging while preserving code, paths, URLs, and error strings verbatim. Reduces output tokens in the compaction summary without changing the <analysis>/<summary> response structure downstream parsers expect.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
class CompactionConfig(BaseModel):
    """Summarization prompt selection for conversation compaction.

    ``caveman_mode`` enables a rule-augmented "caveman" prompt that
    instructs the summarizer LLM to drop articles, filler, pleasantries,
    and hedging while preserving code, paths, URLs, and error strings
    verbatim. Reduces output tokens in the compaction summary without
    changing the
    ``<analysis>/<summary>`` response structure downstream parsers expect.
    """

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Compaction", "x-group": "models", "x-order": 4},
    )

    caveman_mode: bool = Field(
        False,
        description=(
            "Enable caveman-style terse summarization prompt. Drops articles, "
            "filler, pleasantries, and hedging in the compacted summary while "
            "preserving code, file paths, URLs, and error strings verbatim. "
            "Reduces compaction output tokens without changing the response "
            "structure downstream parsers expect."
        ),
        examples=[False],
    )

    @field_validator("caveman_mode", mode="before")
    @classmethod
    def _normalize_caveman_mode(cls, value: Any) -> bool:
        return _coerce_bool(value, default=False)

ConfigCheck dataclass

Result of a configuration preflight check.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4292
4293
4294
4295
4296
4297
4298
4299
4300
4301
4302
4303
4304
4305
4306
4307
4308
4309
4310
@dataclass
class ConfigCheck:
    """Result of a configuration preflight check."""

    name: str
    enabled: bool
    ok: bool
    reason: str | None = None
    metadata: dict[str, Any] = field(default_factory=dict)

    def to_dict(self) -> dict[str, Any]:
        """Serialize the check result to a dictionary."""
        return {
            "name": self.name,
            "enabled": self.enabled,
            "ok": self.ok,
            "reason": self.reason,
            "metadata": self.metadata,
        }

to_dict() -> dict[str, Any]

Serialize the check result to a dictionary.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4302
4303
4304
4305
4306
4307
4308
4309
4310
def to_dict(self) -> dict[str, Any]:
    """Serialize the check result to a dictionary."""
    return {
        "name": self.name,
        "enabled": self.enabled,
        "ok": self.ok,
        "reason": self.reason,
        "metadata": self.metadata,
    }

ConfigWriteAccess

Bases: BaseModel

Whether the configuration store can be persisted to.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3500
3501
3502
3503
3504
3505
3506
3507
class ConfigWriteAccess(BaseModel):
    """Whether the configuration store can be persisted to."""

    model_config = ConfigDict(extra="forbid")

    writable: bool
    code: str | None = None
    reason: str | None = None

ConfigWriteError

Bases: RuntimeError

Raised when the configuration file cannot be persisted.

A bare OSError escaping the persistence path reaches an HTTP surface as an opaque 500 with a traceback and nothing a deployment can act on. This carries the two things a caller needs instead: a stable machine code to branch on and an operator-actionable reason to render. The absolute path stays on the exception (and therefore in the logs) rather than in the prose, because reason is user-facing copy and a server path is not.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3415
3416
3417
3418
3419
3420
3421
3422
3423
3424
3425
3426
3427
3428
3429
3430
3431
3432
3433
3434
3435
3436
3437
3438
3439
3440
3441
3442
3443
3444
3445
3446
3447
3448
3449
3450
3451
3452
3453
3454
3455
3456
3457
3458
3459
3460
3461
3462
3463
3464
3465
3466
3467
3468
3469
3470
3471
3472
3473
3474
3475
3476
3477
3478
3479
3480
3481
3482
3483
3484
3485
3486
3487
3488
3489
3490
3491
3492
3493
3494
3495
3496
3497
class ConfigWriteError(RuntimeError):
    """Raised when the configuration file cannot be persisted.

    A bare ``OSError`` escaping the persistence path reaches an HTTP surface as
    an opaque 500 with a traceback and nothing a deployment can act on. This
    carries the two things a caller needs instead: a stable machine ``code`` to
    branch on and an operator-actionable ``reason`` to render. The absolute
    ``path`` stays on the exception (and therefore in the logs) rather than in
    the prose, because ``reason`` is user-facing copy and a server path is not.
    """

    #: ``errno`` -> stable machine code. Anything unlisted is ``io_error``.
    #: The mapping is by errno alone because that is the only unambiguous
    #: signal, and because the two denial codes call for DIFFERENT operator
    #: actions: ``read_only`` means the mount itself refuses writes (fix the
    #: deployment), ``permission_denied`` means the filesystem allows writes
    #: but not to this process (fix ownership/mode). Collapsing them would
    #: hand the operator prose for the wrong repair.
    _CODES_BY_ERRNO: ClassVar[dict[int, str]] = {
        errno.EROFS: "read_only",
        errno.EACCES: "permission_denied",
        errno.EPERM: "permission_denied",
        errno.ENOSPC: "no_space",
        errno.EDQUOT: "no_space",
    }

    #: Failures meaning atomic replacement is STRUCTURALLY impossible at this
    #: path rather than transiently broken — the only ones a caller may degrade
    #: to a non-atomic in-place write. ENOSPC/EDQUOT/EIO are excluded on
    #: purpose: an in-place write truncates the live config first, so degrading
    #: on a full or failing disk would destroy the config on its way to failing
    #: anyway. See ``AppConfig._in_place_fallback_applies``.
    _STRUCTURAL_ERRNOS: ClassVar[frozenset[int]] = frozenset(
        {errno.EBUSY, errno.EROFS, errno.EACCES, errno.EPERM}
    )

    _REASONS: ClassVar[dict[str, str]] = {
        "read_only": (
            "The configuration directory is mounted read-only; settings cannot be saved "
            "until the deployment grants write access."
        ),
        "permission_denied": (
            "The server process is not allowed to write the configuration directory; "
            "check its ownership and mode for the user the server runs as."
        ),
        "no_space": (
            "The filesystem holding the configuration is out of space; free space or grow "
            "the volume, then save again."
        ),
        "io_error": (
            "The configuration could not be written because the filesystem rejected the "
            "write; see the server logs for the underlying error."
        ),
    }

    def __init__(self, path: Path, code: str, reason: str) -> None:
        """Build the failure from an already-classified ``code``/``reason``."""
        super().__init__(reason)
        self.path = path
        self.code = code
        self.reason = reason

    @classmethod
    def reason_for(cls, code: str) -> str:
        """Return the operator-facing prose for a stable machine ``code``."""
        return cls._REASONS.get(code, cls._REASONS["io_error"])

    @classmethod
    def classify(cls, exc: OSError) -> tuple[str, str]:
        """Map an ``OSError`` to its stable ``(code, reason)`` pair."""
        code = cls._CODES_BY_ERRNO.get(exc.errno or 0, "io_error")
        return code, cls._REASONS[code]

    @classmethod
    def is_structural(cls, exc: OSError) -> bool:
        """Whether ``exc`` means atomic replacement can never work at this path."""
        return exc.errno in cls._STRUCTURAL_ERRNOS

    @classmethod
    def from_oserror(cls, path: Path, exc: OSError) -> ConfigWriteError:
        """Build the typed failure for an ``OSError`` raised while persisting."""
        code, reason = cls.classify(exc)
        return cls(path, code, reason)

__init__(path: Path, code: str, reason: str) -> None

Build the failure from an already-classified code/reason.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3470
3471
3472
3473
3474
3475
def __init__(self, path: Path, code: str, reason: str) -> None:
    """Build the failure from an already-classified ``code``/``reason``."""
    super().__init__(reason)
    self.path = path
    self.code = code
    self.reason = reason

classify(exc: OSError) -> tuple[str, str] classmethod

Map an OSError to its stable (code, reason) pair.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3482
3483
3484
3485
3486
@classmethod
def classify(cls, exc: OSError) -> tuple[str, str]:
    """Map an ``OSError`` to its stable ``(code, reason)`` pair."""
    code = cls._CODES_BY_ERRNO.get(exc.errno or 0, "io_error")
    return code, cls._REASONS[code]

from_oserror(path: Path, exc: OSError) -> ConfigWriteError classmethod

Build the typed failure for an OSError raised while persisting.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3493
3494
3495
3496
3497
@classmethod
def from_oserror(cls, path: Path, exc: OSError) -> ConfigWriteError:
    """Build the typed failure for an ``OSError`` raised while persisting."""
    code, reason = cls.classify(exc)
    return cls(path, code, reason)

is_structural(exc: OSError) -> bool classmethod

Whether exc means atomic replacement can never work at this path.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3488
3489
3490
3491
@classmethod
def is_structural(cls, exc: OSError) -> bool:
    """Whether ``exc`` means atomic replacement can never work at this path."""
    return exc.errno in cls._STRUCTURAL_ERRNOS

reason_for(code: str) -> str classmethod

Return the operator-facing prose for a stable machine code.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3477
3478
3479
3480
@classmethod
def reason_for(cls, code: str) -> str:
    """Return the operator-facing prose for a stable machine ``code``."""
    return cls._REASONS.get(code, cls._REASONS["io_error"])

ContextConfig

Bases: BaseModel

Context window selection and event filtering.

Source code in packages/mewbo_core/src/mewbo_core/config.py
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
class ContextConfig(BaseModel):
    """Context window selection and event filtering."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Context", "x-group": "models", "x-order": 3},
    )

    recent_event_limit: int = Field(
        8,
        description="Maximum number of recent events injected into the context window.",
        examples=[8],
    )
    selection_threshold: float = Field(
        0.8,
        description=(
            "Relevance score threshold (0.0-1.0) for the context selector to keep an event."
        ),
        examples=[0.8],
    )
    selection_enabled: bool = Field(
        True,
        description=(
            "Enable LLM-based context event selection. When false, all recent events are used."
        ),
        examples=[True],
    )
    context_selector_model: str = Field(
        "",
        description=("Model ID for context selection. Falls back to llm.default_model when empty."),
        examples=["anthropic/claude-sonnet-4-6"],
    )

    @field_validator("recent_event_limit", mode="before")
    @classmethod
    def _normalize_recent_event_limit(cls, value: Any) -> int:
        try:
            parsed = int(value)
        except (TypeError, ValueError):
            return 8
        return max(parsed, 1)

    @field_validator("selection_threshold", mode="before")
    @classmethod
    def _normalize_selection_threshold(cls, value: Any) -> float:
        try:
            parsed = float(value)
        except (TypeError, ValueError):
            return 0.8
        return min(max(parsed, 0.0), 1.0)

    @field_validator("selection_enabled", mode="before")
    @classmethod
    def _normalize_selection_enabled(cls, value: Any) -> bool:
        return _coerce_bool(value, default=True)

EnvOverridable

Bases: BaseModel

Base for a section whose fields may be OVERRIDDEN from the environment.

A field declares the variable that overrides it next to the field itself, and this base applies every such declaration in one place::

driver: str = Field(
    "json",
    json_schema_extra={"x-env-var": "MEWBO_STORAGE_DRIVER"},
)

Adding an override is therefore a declaration, not another hand-written validator that reads os.environ. A hand-written one is invisible to configs/app.schema.json and so to the console and the docs; a declaration reaches both, because x-env-var is emitted on the field's schema node like every other x- annotation.

This is the OPPOSITE direction to :class:EnvRef, and the two are not interchangeable:

  • an :class:EnvRef is written BY the operator in app.json as the value ${VARIABLE}, names a variable the file has chosen to defer to, and is an error at load when that variable is unset;
  • an override is declared BY this module on a field the operator may have set to a perfectly good literal, and takes precedence over it. An unset — or empty — variable is simply "not overridden", never an error, because the file value is still the answer.

The override is applied to the raw payload, so the field's own validators still run over it: a bad MEWBO_STORAGE_DRIVER is rejected exactly as a bad driver in the file is.

Cost: O(1) — one pass over a section's declared fields, at validation.

Source code in packages/mewbo_core/src/mewbo_core/config.py
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
class EnvOverridable(BaseModel):
    """Base for a section whose fields may be OVERRIDDEN from the environment.

    A field declares the variable that overrides it next to the field itself,
    and this base applies every such declaration in one place::

        driver: str = Field(
            "json",
            json_schema_extra={"x-env-var": "MEWBO_STORAGE_DRIVER"},
        )

    Adding an override is therefore a declaration, not another hand-written
    validator that reads ``os.environ``. A hand-written one is invisible to
    ``configs/app.schema.json`` and so to the console and the docs; a
    declaration reaches both, because ``x-env-var`` is emitted on the field's
    schema node like every other ``x-`` annotation.

    This is the OPPOSITE direction to :class:`EnvRef`, and the two are not
    interchangeable:

    * an :class:`EnvRef` is written BY the operator in ``app.json`` as the value
      ``${VARIABLE}``, names a variable the file has chosen to defer to, and is
      an error at load when that variable is unset;
    * an override is declared BY this module on a field the operator may have
      set to a perfectly good literal, and takes precedence over it. An unset —
      or empty — variable is simply "not overridden", never an error, because
      the file value is still the answer.

    The override is applied to the raw payload, so the field's own validators
    still run over it: a bad ``MEWBO_STORAGE_DRIVER`` is rejected exactly as a
    bad ``driver`` in the file is.

    Cost: ``O(1)`` — one pass over a section's declared fields, at validation.
    """

    #: Schema key a field uses to declare its override. See the annotation
    #: contract block above ``AppConfig``.
    ENV_VAR_KEY: ClassVar[str] = "x-env-var"

    @model_validator(mode="before")
    @classmethod
    def _apply_env_overrides(cls, payload: Any) -> Any:
        if not isinstance(payload, dict):
            return payload
        overridden = dict(payload)
        for name, field_info in cls.model_fields.items():
            extra = field_info.json_schema_extra
            variable = extra.get(cls.ENV_VAR_KEY) if isinstance(extra, dict) else None
            if not isinstance(variable, str):
                continue
            # Empty reads as absent on purpose: an exported-but-blank variable is
            # how a shell says nothing, not how an operator blanks a setting.
            value = os.environ.get(variable)
            if value:
                overridden[name] = value
        return overridden if overridden != payload else payload

EnvRef

Bases: BaseModel

A config value that NAMES an environment variable instead of holding one.

Written as a value that is exactly ${VARIABLE}, so a secret can stay in the environment and out of app.json::

{"llm": {"api_key": "${OPENAI_API_KEY}"}}

Two deliberate limits, both there so this stays a naming convention rather than a template language:

  • the reference is the WHOLE value or it is not a reference at all — a value that merely contains ${ is left exactly as written, which is what lets a literal password contain those characters with nothing to escape;
  • a reference to a variable that is not set is an error at load, never an empty string. Substituting empty turns a missing secret into a puzzling 401 much later; naming a variable is a claim that it will be there, and a claim is worth checking. A variable set to the empty string IS set, and resolves to empty — that is the way to say a value is deliberately blank.

Not to be confused with :class:EnvOverridable, which runs the other way round: there the FIELD names a variable that overrides whatever the file says, and an unset variable is a no-op rather than an error.

Cost: O(one record) — one walk of the config document, on load only.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3510
3511
3512
3513
3514
3515
3516
3517
3518
3519
3520
3521
3522
3523
3524
3525
3526
3527
3528
3529
3530
3531
3532
3533
3534
3535
3536
3537
3538
3539
3540
3541
3542
3543
3544
3545
3546
3547
3548
3549
3550
3551
3552
3553
3554
3555
3556
3557
3558
3559
3560
3561
3562
3563
3564
3565
3566
3567
3568
3569
3570
3571
3572
3573
3574
3575
3576
3577
3578
3579
3580
3581
3582
3583
3584
3585
class EnvRef(BaseModel):
    """A config value that NAMES an environment variable instead of holding one.

    Written as a value that is exactly ``${VARIABLE}``, so a secret can stay in
    the environment and out of ``app.json``::

        {"llm": {"api_key": "${OPENAI_API_KEY}"}}

    Two deliberate limits, both there so this stays a naming convention rather
    than a template language:

    * the reference is the WHOLE value or it is not a reference at all — a value
      that merely contains ``${`` is left exactly as written, which is what lets
      a literal password contain those characters with nothing to escape;
    * a reference to a variable that is not set is an error at load, never an
      empty string. Substituting empty turns a missing secret into a puzzling
      401 much later; naming a variable is a claim that it will be there, and a
      claim is worth checking. A variable set to the empty string IS set, and
      resolves to empty — that is the way to say a value is deliberately blank.

    Not to be confused with :class:`EnvOverridable`, which runs the other way
    round: there the FIELD names a variable that overrides whatever the file
    says, and an unset variable is a no-op rather than an error.

    Cost: ``O(one record)`` — one walk of the config document, on load only.
    """

    model_config = ConfigDict(extra="forbid", frozen=True)

    variable: str = Field(..., description="Name of the environment variable to read.")

    #: Anchored on purpose: see the whole-value rule in the class docstring.
    _SYNTAX: ClassVar[re.Pattern[str]] = re.compile(r"^\$\{([A-Za-z_][A-Za-z0-9_]*)\}$")

    @classmethod
    def parse(cls, value: Any) -> EnvRef | None:
        """Return the reference ``value`` spells, or ``None`` if it spells none."""
        if not isinstance(value, str):
            return None
        match = cls._SYNTAX.match(value.strip())
        return cls(variable=match.group(1)) if match else None

    def resolve(self, environ: Mapping[str, str], where: str) -> str:
        """Return the variable's value, or raise naming both it and ``where``."""
        resolved = environ.get(self.variable)
        if resolved is None:
            raise ValueError(
                f"Configuration key '{where}' references environment variable "
                f"'{self.variable}', which is not set. Set it, or replace the "
                f"reference with a literal value."
            )
        return resolved

    @classmethod
    def resolve_document(cls, payload: Any, environ: Mapping[str, str], where: str = "") -> Any:
        """Rebuild ``payload`` with every reference in it replaced by its value.

        Returns the payload unchanged — the same object, not a copy — when it
        holds no reference, so the common case allocates nothing and the
        caller's document is never touched. See
        :meth:`AppConfig._drop_retired_keys` for why not touching it matters.
        """
        if isinstance(payload, dict):
            resolved = {
                key: cls.resolve_document(item, environ, f"{where}.{key}" if where else str(key))
                for key, item in payload.items()
            }
            return payload if resolved == payload else resolved
        if isinstance(payload, list):
            items = [
                cls.resolve_document(item, environ, f"{where}[{index}]")
                for index, item in enumerate(payload)
            ]
            return payload if items == payload else items
        reference = cls.parse(payload)
        return reference.resolve(environ, where) if reference is not None else payload

parse(value: Any) -> EnvRef | None classmethod

Return the reference value spells, or None if it spells none.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3544
3545
3546
3547
3548
3549
3550
@classmethod
def parse(cls, value: Any) -> EnvRef | None:
    """Return the reference ``value`` spells, or ``None`` if it spells none."""
    if not isinstance(value, str):
        return None
    match = cls._SYNTAX.match(value.strip())
    return cls(variable=match.group(1)) if match else None

resolve(environ: Mapping[str, str], where: str) -> str

Return the variable's value, or raise naming both it and where.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3552
3553
3554
3555
3556
3557
3558
3559
3560
3561
def resolve(self, environ: Mapping[str, str], where: str) -> str:
    """Return the variable's value, or raise naming both it and ``where``."""
    resolved = environ.get(self.variable)
    if resolved is None:
        raise ValueError(
            f"Configuration key '{where}' references environment variable "
            f"'{self.variable}', which is not set. Set it, or replace the "
            f"reference with a literal value."
        )
    return resolved

resolve_document(payload: Any, environ: Mapping[str, str], where: str = '') -> Any classmethod

Rebuild payload with every reference in it replaced by its value.

Returns the payload unchanged — the same object, not a copy — when it holds no reference, so the common case allocates nothing and the caller's document is never touched. See :meth:AppConfig._drop_retired_keys for why not touching it matters.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3563
3564
3565
3566
3567
3568
3569
3570
3571
3572
3573
3574
3575
3576
3577
3578
3579
3580
3581
3582
3583
3584
3585
@classmethod
def resolve_document(cls, payload: Any, environ: Mapping[str, str], where: str = "") -> Any:
    """Rebuild ``payload`` with every reference in it replaced by its value.

    Returns the payload unchanged — the same object, not a copy — when it
    holds no reference, so the common case allocates nothing and the
    caller's document is never touched. See
    :meth:`AppConfig._drop_retired_keys` for why not touching it matters.
    """
    if isinstance(payload, dict):
        resolved = {
            key: cls.resolve_document(item, environ, f"{where}.{key}" if where else str(key))
            for key, item in payload.items()
        }
        return payload if resolved == payload else resolved
    if isinstance(payload, list):
        items = [
            cls.resolve_document(item, environ, f"{where}[{index}]")
            for index, item in enumerate(payload)
        ]
        return payload if items == payload else items
    reference = cls.parse(payload)
    return reference.resolve(environ, where) if reference is not None else payload

FallbackConfig

Bases: BaseModel

Opt-in cross-model fallback policy.

Disabled by default so a run never fans out to a different model, with different cost, latency, output style and prompt-cache behaviour, without an explicit opt-in. When disabled, an error that is hopeless on the current model (e.g. quota exhausted) halts cleanly for one-click recovery instead.

Source code in packages/mewbo_core/src/mewbo_core/config.py
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
class FallbackConfig(BaseModel):
    """Opt-in cross-model fallback policy.

    Disabled by default so a run never fans out to a different model, with
    different cost, latency, output style and prompt-cache behaviour, without
    an explicit opt-in. When disabled, an error that is hopeless on the current
    model (e.g. quota exhausted) halts cleanly for one-click recovery instead.
    """

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Fallback"},
    )

    enabled: bool = Field(
        False,
        description=(
            "Enable automatic fallback to other models when the primary is "
            "exhausted or hits a hopeless-here error. Off by default."
        ),
    )
    models: list[str] = Field(
        default_factory=list,
        description="Ordered fallback model IDs, tried after the primary is exhausted.",
        examples=[["gpt-5.4", "gemini-2.5-pro"]],
    )
    self_steering: bool = Field(
        False,
        description=(
            "Allow the agent to steer its own model routing at runtime via the "
            "model_control tool (switch down the declared fallback ladder, or up "
            "when allow_upgrade is set). Off by default; the automatic fallback "
            "ladder still operates regardless. Every deliberate switch is bounded "
            "by max_switches and the existing retry budget / circuit breaker."
        ),
    )
    max_switches: int = Field(
        2,
        ge=0,
        description=(
            "Maximum deliberate model switches the model_control tool may perform "
            "in one run. 0 disables switching while leaving status/list readable."
        ),
    )
    allow_upgrade: bool = Field(
        False,
        description=(
            "Permit model_control switches UP the declared ladder (toward the "
            "primary). Off by default, so self-steering is a one-way ratchet "
            "downward — the direction that heals a failing primary without "
            "re-provoking it."
        ),
    )

HomeAssistantConfig

Bases: BaseModel

Home Assistant smart-home integration.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
class HomeAssistantConfig(BaseModel):
    """Home Assistant smart-home integration."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Home Assistant", "x-group": "integrations", "x-order": 1},
    )

    enabled: bool = Field(
        False,
        description=("Enable the Home Assistant tool for smart-home control."),
    )
    url: str = Field(
        "",
        description="Home Assistant API base URL.",
        examples=["http://homeassistant.local:8123"],
    )
    token: str = Field(
        "",
        description="Long-lived access token for Home Assistant authentication.",
        examples=["ha_token_here"],
        json_schema_extra={"x-secret": True},
    )

    @field_validator("enabled", mode="before")
    @classmethod
    def _normalize_enabled(cls, value: Any) -> bool:
        return _coerce_bool(value, default=False)

    def evaluate(self) -> tuple[bool, str | None, dict[str, Any]]:
        if not self.enabled:
            return False, "disabled via config", {}
        missing: list[str] = []
        if not self.url:
            missing.append("home_assistant.url")
        if not self.token:
            missing.append("home_assistant.token")
        if missing:
            return (
                False,
                "missing home_assistant.url/home_assistant.token",
                {"required_config": missing},
            )
        return True, None, {}

HookEntry

Bases: BaseModel

A single hook configuration entry.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
class HookEntry(BaseModel):
    """A single hook configuration entry."""

    model_config = ConfigDict(json_schema_extra={"title": "Hook"})

    type: Literal["command", "http"] = Field(
        "command", description="Hook type: 'command' (shell) or 'http' (POST to URL)."
    )
    command: str = Field(
        "",
        description=(
            "Shell command run via subprocess with shell=True, on the host "
            "running the API/CLI process and with that process's own "
            "privileges: no sandboxing, no approval prompt. Only point this "
            "at trusted, version-controlled scripts. (type=command)"
        ),
    )
    url: str = Field("", description="Target URL for HTTP POST (type=http).")
    headers: dict[str, str] = Field(
        default_factory=dict, description="Extra HTTP headers (type=http)."
    )
    matcher: str | None = Field(
        None, description="Optional fnmatch pattern to limit which tool IDs trigger this hook."
    )
    timeout: int = Field(30, description="Maximum seconds to wait for the hook to finish.")

    @model_validator(mode="after")
    def _validate_type_fields(self) -> HookEntry:
        if self.type == "http" and not self.url:
            msg = "HookEntry type='http' requires a non-empty 'url'."
            raise ValueError(msg)
        # Allow default empty HookEntry() for schema generation.
        if self.type == "command" and not self.command and self.url:
            msg = "HookEntry type='command' but only 'url' is set; use type='http'."
            raise ValueError(msg)
        return self

HooksConfig

Bases: BaseModel

External shell hooks fired during the session lifecycle.

Command hooks run unsandboxed shell commands with the API/CLI process's own privileges (see HookEntry.command's docstring) — a caller who can PATCH this section can execute arbitrary code on the host. x-protected puts the whole section in the same never-read-never-written-via-API tier as the other host-level settings in this file: settable only by editing the config file directly, never over the network regardless of credential (see ConfigSchemaView in apps/mewbo_api).

Source code in packages/mewbo_core/src/mewbo_core/config.py
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
class HooksConfig(BaseModel):
    """External shell hooks fired during the session lifecycle.

    Command hooks run unsandboxed shell commands with the API/CLI process's
    own privileges (see ``HookEntry.command``'s docstring) — a caller who can
    PATCH this section can execute arbitrary code on the host. ``x-protected``
    puts the whole section in the same never-read-never-written-via-API tier
    as the other host-level settings in this file: settable only by editing
    the config file directly, never over the network regardless of
    credential (see ``ConfigSchemaView`` in ``apps/mewbo_api``).
    """

    model_config = ConfigDict(
        json_schema_extra={
            "title": "Hooks",
            "x-group": "integrations",
            "x-order": 3,
            "x-protected": True,
        },
    )

    pre_tool_use: list[HookEntry] = Field(
        default_factory=list, description="Hooks executed before each tool invocation."
    )
    post_tool_use: list[HookEntry] = Field(
        default_factory=list, description="Hooks executed after each tool invocation."
    )
    on_session_start: list[HookEntry] = Field(
        default_factory=list, description="Hooks executed when a new session begins."
    )
    on_session_end: list[HookEntry] = Field(
        default_factory=list, description="Hooks executed when a session ends."
    )
    on_event: list[HookEntry] = Field(
        default_factory=list,
        description=(
            "Hooks executed (fire-and-forget) for every event appended to a "
            "session transcript. The matcher fnmatches the event type."
        ),
    )

LLMConfig

Bases: BaseModel

LLM provider connection and model selection.

Source code in packages/mewbo_core/src/mewbo_core/config.py
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
class LLMConfig(BaseModel):
    """LLM provider connection and model selection."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Language Model", "x-group": "models", "x-order": 1},
    )

    api_base: str = Field(
        "",
        description=(
            "Optional base URL override. Leave empty for direct "
            "provider access (LiteLLM routes automatically from "
            "the model prefix). Set only when using a proxy "
            "(e.g. LiteLLM, Bifrost)."
        ),
        examples=["", "https://my-litellm-proxy.example.com/v1"],
    )
    api_key: str = Field(
        "",
        description=("API key for the LLM provider (e.g. Anthropic, OpenAI) or proxy master key."),
        examples=["sk-ant-xxxxxxxx"],
        json_schema_extra={"x-secret": True},
    )
    default_model: str = Field(
        "gpt-5.2",
        description=(
            "Model ID using 'provider/model' syntax. LiteLLM "
            "auto-routes to the right API endpoint. When using "
            "a proxy, adjust the prefix to match its routing."
        ),
        examples=["anthropic/claude-sonnet-4-6"],
    )
    action_plan_model: str = Field(
        "",
        description=(
            "Model ID the orchestrator uses to generate a session's initial plan "
            "and, when no explicit model is passed in, as that session's default "
            "model. Falls back to default_model when empty."
        ),
        examples=["anthropic/claude-sonnet-4-6"],
    )
    tool_model: str = Field(
        "",
        description=("Model ID used by individual tools. Falls back to default_model when empty."),
        examples=["anthropic/claude-sonnet-4-6"],
    )
    title_model: str = Field(
        "",
        description=(
            "Model ID for session-title generation. Falls back to default_model when empty."
        ),
        examples=["anthropic/claude-haiku-4-5-20251001"],
    )
    compact_models: list[str] = Field(
        default_factory=lambda: ["default"],
        description=(
            "Priority-ordered list of models for context compaction. "
            "On failure, the next model in the list is tried. "
            'The keyword "default" resolves to the running agent\'s model. '
            'Example: ["anthropic/claude-haiku-4-5-20251001", "default"]'
        ),
        examples=[["anthropic/claude-haiku-4-5-20251001", "default"]],
    )
    fallback_models: list[str] = Field(
        default_factory=list,
        description=(
            "Flat ordered list of fallback model IDs. Prefer 'fallback' "
            "(typed, explicit opt-in). A non-empty value here is still honored, "
            "and is treated as fallback enabled."
        ),
        examples=[["gpt-5.4", "gemini-2.5-pro"]],
    )
    fallback: FallbackConfig = Field(
        default_factory=lambda: FallbackConfig.model_validate({}),
        description="Opt-in cross-model fallback policy (see FallbackConfig).",
    )
    proxy_model_prefix: str = Field(
        "openai",
        description=(
            "LiteLLM provider prefix prepended to model names when api_base is set. "
            "LiteLLM strips this prefix before forwarding the model name to the proxy, "
            "so the proxy receives the model ID it advertises in /v1/models. "
            "Leave as 'openai' for LiteLLM proxy, Bifrost, and OpenRouter. "
            "Only relevant when api_base is configured."
        ),
        examples=["openai", "azure", "vertex_ai"],
        json_schema_extra={"x-advanced": True},
    )
    request_timeout: float = Field(
        DEFAULT_REQUEST_TIMEOUT,
        description=(
            "Seconds handed to the LiteLLM client as its HTTP timeout. Because the "
            "loop STREAMS, this behaves as the maximum gap BETWEEN response chunks, "
            "not as a ceiling on total call duration — a long generation that keeps "
            "emitting chunks never trips it. The total-duration ceiling is "
            "agent.llm_call_timeout. Raise this only when a provider is slow to send "
            "its FIRST chunk."
        ),
        json_schema_extra={"x-advanced": True},
    )
    reasoning_effort: str = Field(
        "",
        description=(
            "Reasoning effort hint for supported models. One of low, medium, high, none, or empty."
        ),
        examples=["medium"],
    )
    reasoning_effort_models: list[str] = Field(
        default_factory=list,
        description=(
            "Additional model IDs (or 'prefix*' patterns) that should receive the "
            "reasoning_effort parameter, on top of the built-in match for gpt-5, "
            "o3, Claude, and Gemini models. Use this to opt in a model the "
            "built-in detection doesn't recognize yet."
        ),
    )
    structured_patch_models: list[str] = Field(
        default_factory=list,
        description=(
            "Model IDs (or glob prefixes ending in '*') that prefer the "
            "structured_patch edit tool over search_replace_block. Runtime "
            "override layer ON TOP of the controllable model→tool-variant map "
            "(mewbo_core/prompts/model_variants.yaml), which now holds the "
            "built-in defaults (GPT-5/o3/o4/Codex/GPT-4); only set this to "
            "override or extend without editing that file."
        ),
        json_schema_extra={"x-advanced": True},
    )

    @field_validator("reasoning_effort", mode="before")
    @classmethod
    def _normalize_reasoning_effort(cls, value: Any) -> str:
        if value is None:
            return ""
        normalized = str(value).strip().lower()
        if normalized in {"low", "medium", "high", "none"}:
            return normalized
        return ""

    @field_validator("reasoning_effort_models", mode="before")
    @classmethod
    def _normalize_reasoning_effort_models(cls, value: Any) -> list[str]:
        return [entry.lower() for entry in _coerce_list(value)]

    @field_validator("structured_patch_models", mode="before")
    @classmethod
    def _normalize_structured_patch_models(cls, value: Any) -> list[str]:
        return [entry.lower() for entry in _coerce_list(value)]

    @field_validator("proxy_model_prefix", mode="before")
    @classmethod
    def _normalize_proxy_model_prefix(cls, value: Any) -> str:
        normalized = str(value).strip().strip("/") if value is not None else ""
        return normalized or "openai"

    def _resolve_api_base(self) -> str | None:
        base = self.api_base.strip()
        return base or None

    def _models_endpoint(self) -> str:
        base = self._resolve_api_base()
        if not base:
            raise ValueError("llm.api_base is not set.")
        base = base.rstrip("/")
        if base.endswith("/v1"):
            return f"{base}/models"
        return f"{base}/v1/models"

    def list_models(self, *, timeout: float = 8.0) -> list[str]:
        api_key = self.api_key.strip()
        if not api_key:
            raise ValueError("llm.api_key is not set.")
        request = Request(
            self._models_endpoint(),
            headers={"Authorization": f"Bearer {api_key}"},
        )
        try:
            with urlopen(request, timeout=timeout) as response:
                payload = json.loads(response.read().decode("utf-8"))
        except HTTPError as exc:
            raise ValueError(f"Model listing failed: HTTP {exc.code}") from exc
        except URLError as exc:
            raise ValueError(f"Model listing failed: {exc.reason}") from exc
        data = payload.get("data", [])
        return sorted([item.get("id") for item in data if item.get("id")])

    def resolve_available_model(self, model: str, *, fallback: str, timeout: float = 4.0) -> str:
        """Return *model* if the proxy still advertises it, else *fallback*.

        Guards a persisted/stale model id (e.g. a wiki reindex replaying an old
        submission) against a model the proxy has since retired, which would
        otherwise fast-fail the whole run on an invalid-model 400. Best-effort:
        if the model list can't be fetched we trust *model* (the caller's
        retry/fallback ladder is the backstop). The provider prefix is ignored
        on both sides (``openai/x`` matches a bare ``x`` the proxy advertises).
        """
        if not model:
            return fallback
        try:
            available = self.list_models(timeout=timeout)
        except Exception:  # noqa: BLE001 — unreachable/unset proxy ⇒ trust model
            return model
        if not available:
            return model

        def _bare(name: str) -> str:
            return name.split("/", 1)[-1].strip().lower()

        target = _bare(model)
        if any(_bare(entry) == target for entry in available):
            return model
        return fallback

    def validate_models(self) -> ConfigCheck:
        if not self._resolve_api_base():
            return ConfigCheck(
                name="llm",
                enabled=True,
                ok=True,
                reason="api_base not set; using direct provider routing",
            )
        if not self.api_key.strip():
            return ConfigCheck(
                name="llm",
                enabled=True,
                ok=False,
                reason="llm.api_key is not set",
            )
        try:
            models = self.list_models()
        except ValueError as exc:
            return ConfigCheck(name="llm", enabled=True, ok=False, reason=str(exc))
        missing: list[str] = []
        compact_explicit = {m for m in self.compact_models if m and m != "default"}
        for model_name in {
            self.default_model,
            self.action_plan_model,
            self.tool_model,
            self.title_model,
            *compact_explicit,
        }:
            if model_name and model_name not in models:
                missing.append(model_name)
        if missing:
            return ConfigCheck(
                name="llm",
                enabled=True,
                ok=False,
                reason="Configured model not found in API",
                metadata={"missing_models": missing, "available_models": models},
            )
        return ConfigCheck(name="llm", enabled=True, ok=True, metadata={"available_models": models})

    def effective_fallback_models(self) -> list[str]:
        """Resolve the active fallback ladder from this instance.

        The precedence rule lives HERE, on the data that owns it, so the
        module-level accessor (which reads the process-wide config) and any
        validator holding a not-yet-installed ``AppConfig`` cannot disagree
        about how long the ladder is.
        """
        if self.fallback.enabled:
            return list(self.fallback.models) or list(self.fallback_models)
        return list(self.fallback_models)

effective_fallback_models() -> list[str]

Resolve the active fallback ladder from this instance.

The precedence rule lives HERE, on the data that owns it, so the module-level accessor (which reads the process-wide config) and any validator holding a not-yet-installed AppConfig cannot disagree about how long the ladder is.

Source code in packages/mewbo_core/src/mewbo_core/config.py
690
691
692
693
694
695
696
697
698
699
700
def effective_fallback_models(self) -> list[str]:
    """Resolve the active fallback ladder from this instance.

    The precedence rule lives HERE, on the data that owns it, so the
    module-level accessor (which reads the process-wide config) and any
    validator holding a not-yet-installed ``AppConfig`` cannot disagree
    about how long the ladder is.
    """
    if self.fallback.enabled:
        return list(self.fallback.models) or list(self.fallback_models)
    return list(self.fallback_models)

resolve_available_model(model: str, *, fallback: str, timeout: float = 4.0) -> str

Return model if the proxy still advertises it, else fallback.

Guards a persisted/stale model id (e.g. a wiki reindex replaying an old submission) against a model the proxy has since retired, which would otherwise fast-fail the whole run on an invalid-model 400. Best-effort: if the model list can't be fetched we trust model (the caller's retry/fallback ladder is the backstop). The provider prefix is ignored on both sides (openai/x matches a bare x the proxy advertises).

Source code in packages/mewbo_core/src/mewbo_core/config.py
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
def resolve_available_model(self, model: str, *, fallback: str, timeout: float = 4.0) -> str:
    """Return *model* if the proxy still advertises it, else *fallback*.

    Guards a persisted/stale model id (e.g. a wiki reindex replaying an old
    submission) against a model the proxy has since retired, which would
    otherwise fast-fail the whole run on an invalid-model 400. Best-effort:
    if the model list can't be fetched we trust *model* (the caller's
    retry/fallback ladder is the backstop). The provider prefix is ignored
    on both sides (``openai/x`` matches a bare ``x`` the proxy advertises).
    """
    if not model:
        return fallback
    try:
        available = self.list_models(timeout=timeout)
    except Exception:  # noqa: BLE001 — unreachable/unset proxy ⇒ trust model
        return model
    if not available:
        return model

    def _bare(name: str) -> str:
        return name.split("/", 1)[-1].strip().lower()

    target = _bare(model)
    if any(_bare(entry) == target for entry in available):
        return model
    return fallback

LSPConfig

Bases: BaseModel

Language Server Protocol integration settings.

Source code in packages/mewbo_core/src/mewbo_core/config.py
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
class LSPConfig(BaseModel):
    """Language Server Protocol integration settings."""

    model_config = ConfigDict(extra="forbid", json_schema_extra={"title": "LSP"})

    enabled: bool = Field(
        True,
        description=(
            "Master switch for the native LSP tool (hover/diagnostics/"
            "go-to-definition). When off, or when the pygls dependency isn't "
            "installed, the tool is never registered and the agent works from "
            "grep/read alone."
        ),
    )
    servers: dict[str, dict[str, Any]] = Field(
        default_factory=dict,
        description=(
            "Override or extend built-in server definitions. "
            'Set {"pyright": {"disabled": true}} to disable a built-in, '
            "or add custom servers with command/extensions/root_markers."
        ),
    )

LangfuseConfig

Bases: BaseModel

Langfuse LLM observability integration.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
class LangfuseConfig(BaseModel):
    """Langfuse LLM observability integration."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Langfuse", "x-group": "integrations", "x-order": 2},
    )

    enabled: bool = Field(False, description="Enable Langfuse tracing for all LLM calls.")
    host: str = Field(
        "",
        description="Langfuse server URL.",
        examples=["https://langfuse.server.local"],
    )
    project_id: str = Field(
        "",
        description="Langfuse project ID for constructing dashboard URLs.",
        examples=["clvh22gis002oru6ay1rm2eh0"],
    )
    public_key: str = Field(
        "",
        description="Langfuse project public key.",
        examples=["pk-lf-xxxxxxxxxxxxxxxx"],
        json_schema_extra={"x-secret": True},
    )
    secret_key: str = Field(
        "",
        description="Langfuse project secret key.",
        examples=["sk-lf-xxxxxxxxxxxxxxxx"],
        json_schema_extra={"x-secret": True},
    )

    @field_validator("enabled", mode="before")
    @classmethod
    def _normalize_enabled(cls, value: Any) -> bool:
        return _coerce_bool(value, default=False)

    def evaluate(self) -> tuple[bool, str | None, dict[str, Any]]:
        if not self.enabled:
            return False, "disabled via config", {}
        missing: list[str] = []
        if not self.public_key:
            missing.append("langfuse.public_key")
        if not self.secret_key:
            missing.append("langfuse.secret_key")
        if missing:
            return (
                False,
                "missing langfuse.public_key/langfuse.secret_key",
                {"required_config": missing},
            )
        try:
            from langfuse.langchain import CallbackHandler  # noqa: F401
        except ModuleNotFoundError as exc:
            message = str(exc).lower()
            if "langchain" in message:
                return False, "langchain not installed", {}
            return False, "langfuse not installed", {}
        return True, None, {}

MongoDBConfig

Bases: EnvOverridable

MongoDB connection settings.

Source code in packages/mewbo_core/src/mewbo_core/config.py
2910
2911
2912
2913
2914
2915
2916
2917
2918
2919
2920
2921
2922
2923
2924
2925
2926
2927
2928
2929
2930
2931
2932
2933
2934
2935
2936
2937
2938
2939
2940
2941
2942
2943
2944
class MongoDBConfig(EnvOverridable):
    """MongoDB connection settings."""

    # validate_default=True so the normalizers below run even when the field
    # falls back to its default (the common `model_validate({})` path).
    model_config = ConfigDict(validate_default=True, json_schema_extra={"title": "MongoDB"})

    uri: str = Field(
        "mongodb://localhost:27017",
        description=(
            "MongoDB connection URI (includes host, port, credentials). "
            "Overridden by the MEWBO_MONGODB_URI environment variable."
        ),
        examples=["mongodb://user:pass@localhost:27017/mewbo?authSource=admin"],
        json_schema_extra={"x-env-var": "MEWBO_MONGODB_URI"},
    )
    database: str = Field(
        "mewbo",
        description=(
            "MongoDB database name for session storage. "
            "Overridden by the MEWBO_MONGODB_DATABASE environment variable."
        ),
        examples=["mewbo"],
        json_schema_extra={"x-env-var": "MEWBO_MONGODB_DATABASE"},
    )

    @field_validator("uri", mode="before")
    @classmethod
    def _normalize_uri(cls, value: Any) -> str:
        return str(value).strip() if value else "mongodb://localhost:27017"

    @field_validator("database", mode="before")
    @classmethod
    def _normalize_database(cls, value: Any) -> str:
        return str(value).strip() if value else "mewbo"

PermissionsConfig

Bases: BaseModel

Tool execution permission policy.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
class PermissionsConfig(BaseModel):
    """Tool execution permission policy."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Permissions", "x-group": "agent", "x-order": 2},
    )

    policy_path: str = Field(
        "",
        description="Path to a JSON or TOML permission policy file. Empty uses built-in defaults.",
        examples=["./configs/policy.json"],
    )
    approval_mode: str = Field(
        "ask",
        description=(
            "Default approval mode: 'ask' prompts the user, 'allow' auto-approves, 'deny' blocks."
        ),
        examples=["ask"],
    )

    @field_validator("approval_mode", mode="before")
    @classmethod
    def _normalize_approval_mode(cls, value: Any) -> str:
        if value is None:
            return "ask"
        normalized = str(value).strip().lower()
        if normalized in {"allow", "auto", "approve", "yes"}:
            return "allow"
        if normalized in {"deny", "never", "no"}:
            return "deny"
        return "ask"

PluginsConfig

Bases: BaseModel

Plugin system configuration.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
class PluginsConfig(BaseModel):
    """Plugin system configuration."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Plugins", "x-group": "plugins", "x-order": 1},
    )

    enabled: bool = Field(
        True,
        description=(
            "Turn the whole plugin system on, including Mewbo's own built-in "
            "suites such as `widget_builder`.\n\n"
            "While this is off, no plugin contributes anything to the agent, "
            "whether it is built in or installed from a marketplace: no agent "
            "definitions, no skills, no hooks, no MCP tools. The "
            "`enabled_plugins` and `marketplaces` settings below are ignored "
            "entirely until you turn it back on."
        ),
    )
    enabled_plugins: list[str] = Field(
        default_factory=list,
        description=(
            "Plugin names to enable. Empty = all installed plugins. "
            "Format: 'plugin-name' or 'plugin-name@marketplace'."
        ),
    )
    marketplaces: list[str] = Field(
        default_factory=lambda: ["anthropics/claude-plugins-official"],
        description=(
            "Marketplace catalogs holding a marketplace.json plugin index, on any "
            "git host. Each entry is a full git URL "
            "(https/ssh/git, or scp-style git@host:owner/repo), a 'host/owner/repo' "
            "shorthand, or a bare 'owner/repo' (cloned from marketplace_default_host)."
        ),
    )
    marketplace_default_host: str = Field(
        "github.com",
        description=(
            "Default git host for bare 'owner/repo' marketplace entries. Full URLs "
            "and 'host/owner/repo' entries ignore this."
        ),
    )
    install_path: str = Field(
        "",
        description=(
            "Override install path for Mewbo-managed plugins. "
            "Defaults to $MEWBO_HOME/plugins/ (via resolve_mewbo_home)."
        ),
    )

    @field_validator("enabled", mode="before")
    @classmethod
    def _normalize_enabled(cls, value: Any) -> bool:
        return _coerce_bool(value, default=True)

    @field_validator("enabled_plugins", "marketplaces", mode="before")
    @classmethod
    def _normalize_string_lists(cls, value: Any) -> list[str]:
        return _coerce_list(value)

    @field_validator("install_path", mode="before")
    @classmethod
    def _normalize_install_path(cls, value: Any) -> str:
        raw = str(value).strip() if value else ""
        if raw:
            return str(Path(raw).expanduser().resolve())
        return ""

    @field_validator("marketplace_default_host", mode="before")
    @classmethod
    def _normalize_default_host(cls, value: Any) -> str:
        raw = str(value).strip() if value else ""
        return raw or "github.com"

    def resolve_install_dir(self) -> Path:
        """Uses install_path if set, otherwise resolve_mewbo_home() / 'plugins'."""
        if self.install_path:
            return Path(self.install_path)
        return resolve_mewbo_home() / "plugins"

    def resolve_registry_paths(self) -> list[Path]:
        """Paths to search for installed_plugins.json: CC cache + our own."""
        paths = [
            Path.home() / ".claude" / "plugins" / "installed_plugins.json",
            self.resolve_install_dir() / "installed_plugins.json",
        ]
        return [p for p in paths if p.parent.is_dir()]

    def resolve_marketplace_dirs(self, *, sync: bool = True) -> list[Path]:
        """Paths to search for marketplace.json caches.

        Scans both Claude Code's and our own marketplace directories.
        When *sync* is True and ``self.marketplaces`` lists repos that aren't
        yet cloned locally, ``sync_marketplaces`` clones them first.
        """
        dirs: list[Path] = []
        # 1. Check Claude Code's cache (read-only)
        cc_base = Path.home() / ".claude" / "plugins" / "marketplaces"
        if cc_base.is_dir():
            dirs.extend(sorted(d for d in cc_base.iterdir() if d.is_dir()))

        # 2. Check our own cache
        own_base = self.resolve_install_dir() / "marketplaces"
        if own_base.is_dir():
            dirs.extend(sorted(d for d in own_base.iterdir() if d.is_dir()))

        # 3. Ensure every configured catalog is cloned (skip ones already present).
        if sync and self.marketplaces:
            from mewbo_core.tooling.plugins import marketplace_dir_name, sync_marketplaces

            existing_names = {d.name for d in dirs}
            missing = []
            for entry in self.marketplaces:
                canonical = marketplace_dir_name(entry, default_host=self.marketplace_default_host)
                # A clone directory named by the bare repo leaf (Claude Code's
                # own cache layout) holds the same catalog — reuse it instead of
                # re-cloning under the canonical name.
                legacy_leaf = entry.rstrip("/").split("/")[-1]
                if canonical not in existing_names and legacy_leaf not in existing_names:
                    missing.append(entry)
            if missing:
                synced = sync_marketplaces(
                    missing,
                    self.resolve_install_dir(),
                    default_host=self.marketplace_default_host,
                )
                dirs.extend(synced)

        return dirs

resolve_install_dir() -> Path

Uses install_path if set, otherwise resolve_mewbo_home() / 'plugins'.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1729
1730
1731
1732
1733
def resolve_install_dir(self) -> Path:
    """Uses install_path if set, otherwise resolve_mewbo_home() / 'plugins'."""
    if self.install_path:
        return Path(self.install_path)
    return resolve_mewbo_home() / "plugins"

resolve_marketplace_dirs(*, sync: bool = True) -> list[Path]

Paths to search for marketplace.json caches.

Scans both Claude Code's and our own marketplace directories. When sync is True and self.marketplaces lists repos that aren't yet cloned locally, sync_marketplaces clones them first.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
def resolve_marketplace_dirs(self, *, sync: bool = True) -> list[Path]:
    """Paths to search for marketplace.json caches.

    Scans both Claude Code's and our own marketplace directories.
    When *sync* is True and ``self.marketplaces`` lists repos that aren't
    yet cloned locally, ``sync_marketplaces`` clones them first.
    """
    dirs: list[Path] = []
    # 1. Check Claude Code's cache (read-only)
    cc_base = Path.home() / ".claude" / "plugins" / "marketplaces"
    if cc_base.is_dir():
        dirs.extend(sorted(d for d in cc_base.iterdir() if d.is_dir()))

    # 2. Check our own cache
    own_base = self.resolve_install_dir() / "marketplaces"
    if own_base.is_dir():
        dirs.extend(sorted(d for d in own_base.iterdir() if d.is_dir()))

    # 3. Ensure every configured catalog is cloned (skip ones already present).
    if sync and self.marketplaces:
        from mewbo_core.tooling.plugins import marketplace_dir_name, sync_marketplaces

        existing_names = {d.name for d in dirs}
        missing = []
        for entry in self.marketplaces:
            canonical = marketplace_dir_name(entry, default_host=self.marketplace_default_host)
            # A clone directory named by the bare repo leaf (Claude Code's
            # own cache layout) holds the same catalog — reuse it instead of
            # re-cloning under the canonical name.
            legacy_leaf = entry.rstrip("/").split("/")[-1]
            if canonical not in existing_names and legacy_leaf not in existing_names:
                missing.append(entry)
        if missing:
            synced = sync_marketplaces(
                missing,
                self.resolve_install_dir(),
                default_host=self.marketplace_default_host,
            )
            dirs.extend(synced)

    return dirs

resolve_registry_paths() -> list[Path]

Paths to search for installed_plugins.json: CC cache + our own.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1735
1736
1737
1738
1739
1740
1741
def resolve_registry_paths(self) -> list[Path]:
    """Paths to search for installed_plugins.json: CC cache + our own."""
    paths = [
        Path.home() / ".claude" / "plugins" / "installed_plugins.json",
        self.resolve_install_dir() / "installed_plugins.json",
    ]
    return [p for p in paths if p.parent.is_dir()]

ProjectConfig

Bases: BaseModel

A project directory exposed to the REST API for session scoping.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
class ProjectConfig(BaseModel):
    """A project directory exposed to the REST API for session scoping."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Project"},
    )

    path: str = Field(
        "",
        description=(
            "Absolute path to the project's root directory on the API host's "
            "filesystem (tilde is expanded). The directory must already exist, "
            "because Mewbo does not create it, and a session request against "
            "this project fails if the path is missing."
        ),
    )
    description: str = Field(
        "",
        description=(
            "Short blurb shown next to this project's name in project pickers. "
            "Purely informational, with no effect on behavior."
        ),
    )
    allowed_paths: list[str] = Field(
        default_factory=list,
        description=(
            "Additional directories a session may reach while this project is "
            "its active one, on top of the project's own path — a sibling "
            "checkout or a shared data directory, say. Both halves of the "
            "boundary honour it: the entry is subtracted from "
            "agent.shell_sandbox's deny set for shell subprocesses, and added "
            "to the roots agent.path_scope_to_active_project allows a path "
            "argument to resolve under, so it re-permits a directory that "
            "would otherwise be out of scope as belonging to another project. "
            "Absolute or ~-relative; a path that does not exist is ignored."
        ),
    )

    @field_validator("path", mode="before")
    @classmethod
    def _normalize_path(cls, value: Any) -> str:
        raw = str(value).strip() if value else ""
        if raw:
            return str(Path(raw).expanduser().resolve())
        return ""

    @field_validator("allowed_paths", mode="before")
    @classmethod
    def _normalize_allowed_paths(cls, value: Any) -> list[str]:
        raw_list = value if isinstance(value, list) else [value] if value else []
        normalized: list[str] = []
        for raw in raw_list:
            text = str(raw).strip() if raw else ""
            if text:
                normalized.append(str(Path(text).expanduser().resolve()))
        return normalized

RetiredKey

Bases: BaseModel

One config key that no longer exists, and what to tell whoever still sets it.

Deleting a field does not delete the key from the OPERATOR'S app.json, and what that key does next is opposite in the two section kinds: under extra="ignore" it is swallowed in silence, so a knob that stopped working looks exactly like one that works; under extra="forbid" it is a ValidationError during startup, so the process will not boot.

A retired key is therefore pruned from the payload before validation — so a forbid section never sees it — and announced once per key per process, so the silence becomes an instruction to delete the line.

Cost: O(one record) — one walk per declared retired key over the config document, on the load path only.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3588
3589
3590
3591
3592
3593
3594
3595
3596
3597
3598
3599
3600
3601
3602
3603
3604
3605
3606
3607
3608
3609
3610
3611
3612
3613
3614
3615
3616
3617
3618
3619
3620
3621
3622
3623
3624
3625
3626
3627
3628
3629
3630
3631
3632
3633
3634
3635
3636
3637
3638
3639
3640
3641
3642
3643
3644
3645
3646
3647
3648
3649
class RetiredKey(BaseModel):
    """One config key that no longer exists, and what to tell whoever still sets it.

    Deleting a field does not delete the key from the OPERATOR'S ``app.json``,
    and what that key does next is opposite in the two section kinds: under
    ``extra="ignore"`` it is swallowed in silence, so a knob that stopped
    working looks exactly like one that works; under ``extra="forbid"`` it is a
    ``ValidationError`` during startup, so the process will not boot.

    A retired key is therefore pruned from the payload before validation — so a
    ``forbid`` section never sees it — and announced once per key per process,
    so the silence becomes an instruction to delete the line.

    Cost: ``O(one record)`` — one walk per declared retired key over the config
    document, on the load path only.
    """

    model_config = ConfigDict(extra="forbid")

    path: tuple[str, ...] = Field(
        ..., min_length=1, description="Dotted location of the key, split into segments."
    )
    guidance: str = Field(
        ..., description="What the operator should do instead, in one sentence."
    )

    @property
    def dotted(self) -> str:
        """The key as an operator would read it in ``app.json``."""
        return ".".join(self.path)

    def prune(self, payload: dict[str, Any]) -> tuple[dict[str, Any], bool]:
        """Return ``payload`` without this key, plus whether it was there at all.

        Copies only the dicts ALONG the key's own path and shares every other
        branch, so the caller's document is never mutated and nothing else in it
        is duplicated. A whole-document deep copy would be the obvious
        alternative and is not available: a payload reaching ``model_validate``
        may already hold constructed submodels, which do not survive a
        JSON round trip.
        """
        return self._without(payload, self.path)

    @classmethod
    def _without(cls, node: dict[str, Any], path: tuple[str, ...]) -> tuple[dict[str, Any], bool]:
        """Rebuild ``node`` minus ``path``, sharing every branch off that path."""
        head, rest = path[0], path[1:]
        if head not in node:
            return node, False
        if not rest:
            trimmed = dict(node)
            del trimmed[head]
            return trimmed, True
        child = node[head]
        if not isinstance(child, dict):
            return node, False
        rebuilt, removed = cls._without(child, rest)
        if not removed:
            return node, False
        trimmed = dict(node)
        trimmed[head] = rebuilt
        return trimmed, True

dotted: str property

The key as an operator would read it in app.json.

prune(payload: dict[str, Any]) -> tuple[dict[str, Any], bool]

Return payload without this key, plus whether it was there at all.

Copies only the dicts ALONG the key's own path and shares every other branch, so the caller's document is never mutated and nothing else in it is duplicated. A whole-document deep copy would be the obvious alternative and is not available: a payload reaching model_validate may already hold constructed submodels, which do not survive a JSON round trip.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3619
3620
3621
3622
3623
3624
3625
3626
3627
3628
3629
def prune(self, payload: dict[str, Any]) -> tuple[dict[str, Any], bool]:
    """Return ``payload`` without this key, plus whether it was there at all.

    Copies only the dicts ALONG the key's own path and shares every other
    branch, so the caller's document is never mutated and nothing else in it
    is duplicated. A whole-document deep copy would be the obvious
    alternative and is not available: a payload reaching ``model_validate``
    may already hold constructed submodels, which do not survive a
    JSON round trip.
    """
    return self._without(payload, self.path)

RetryConfig

Bases: BaseModel

Automatic LLM-call retry / fallback resilience knobs.

Same-model retry hardening (full-jitter backoff, circuit breaker, retry budget, wall-clock deadline, doom-loop halt) is always on; cross-model fallback is opt-in via llm.fallback. Defaults are calibrated from production agent loops, not the tighter vendor-SDK defaults.

Source code in packages/mewbo_core/src/mewbo_core/config.py
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
2204
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
2226
2227
2228
2229
2230
2231
2232
2233
2234
2235
2236
2237
2238
2239
2240
2241
2242
2243
2244
2245
2246
2247
2248
2249
2250
2251
2252
2253
2254
2255
2256
2257
2258
2259
2260
2261
2262
2263
2264
2265
2266
2267
2268
2269
2270
2271
class RetryConfig(BaseModel):
    """Automatic LLM-call retry / fallback resilience knobs.

    Same-model retry hardening (full-jitter backoff, circuit breaker, retry
    budget, wall-clock deadline, doom-loop halt) is always on; cross-model
    fallback is opt-in via ``llm.fallback``. Defaults are calibrated from
    production agent loops, not the tighter vendor-SDK defaults.
    """

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Retry"},
    )

    backoff_base: float = Field(
        DEFAULT_BACKOFF_BASE,
        description="Base seconds for full-jitter backoff: random(0, min(cap, base*2^(n-1))).",
    )
    backoff_cap: float = Field(DEFAULT_BACKOFF_CAP, description="Maximum backoff delay in seconds.")
    retry_after_cap: float = Field(
        DEFAULT_RETRY_AFTER_CAP,
        description=(
            "Upper bound in seconds applied to a server's Retry-After header "
            "before the loop sleeps on it; caps how long one misbehaving "
            "response can stall a run."
        ),
    )
    turn_deadline: float = Field(
        DEFAULT_TURN_DEADLINE,
        description=(
            "Wall-clock seconds budget for one logical LLM call across all "
            "retries and fallbacks. Checked before each attempt AND before "
            "advancing to the next model, so a spent budget stops the chain "
            "without claiming a model it never called. Must exceed "
            "llm_call_timeout x llm_call_retries plus one more attempt, or no "
            "fallback model is ever reachable. 0 disables."
        ),
    )
    fallback_retries: int = Field(
        DEFAULT_FALLBACK_RETRIES,
        description="Attempts per fallback model after the primary is exhausted.",
    )
    circuit_breaker_threshold: int = Field(
        DEFAULT_CB_THRESHOLD,
        description=(
            "Consecutive per-model failures before that model is cooled down "
            "and skipped (when an alternative exists). 0 disables this "
            "heuristic. It does NOT disable the cooldown a provider declares "
            "for itself: a model whose quota the provider reports as exhausted "
            "is still skipped until that quota resets, since that is a stated "
            "fact rather than an inference this threshold tunes."
        ),
    )
    circuit_breaker_cooldown: float = Field(
        DEFAULT_CB_COOLDOWN,
        description=(
            "Seconds a model is skipped after tripping the circuit breaker. A "
            "provider-declared reset time overrides this for that model."
        ),
    )
    budget_capacity: float = Field(
        DEFAULT_BUDGET_CAPACITY,
        description=(
            "Token-bucket retry budget for transient failures (timeouts, 5xx, "
            "connection errors). Retries stop once the bucket drops to half "
            "capacity, so a sustained outage fails fast instead of storming."
        ),
    )
    rate_limit_budget_capacity: float = Field(
        DEFAULT_RATE_LIMIT_BUDGET_CAPACITY,
        description=(
            "Separate, much smaller token-bucket budget for rate-limit (429) "
            "retries. Retrying a server error bets that the server recovers in "
            "seconds and usually pays; retrying a throttle bets against a rate "
            "window the provider controls, and each attempt re-sends the whole "
            "prompt. Once this bucket is spent the run moves to the next model "
            "in the ladder rather than failing, because that model draws on a "
            "different quota pool. Retries stop at half capacity, as above."
        ),
    )
    doom_loop_threshold: int = Field(
        DEFAULT_DOOM_LOOP_THRESHOLD,
        description=(
            "Halt cleanly when the model repeats the same tool + identical input "
            "this many times in a row (no progress). 0 disables."
        ),
    )

RuntimeConfig

Bases: BaseModel

Runtime environment settings.

Source code in packages/mewbo_core/src/mewbo_core/config.py
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
class RuntimeConfig(BaseModel):
    """Runtime environment settings."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={
            "title": "Runtime",
            "x-group": "server",
            "x-order": 3,
            "x-advanced": True,
        },
    )

    envmode: str = Field(
        "dev",
        description=(
            "Free-text label for this deployment (e.g. dev, staging, prod). It "
            "becomes the Langfuse tracing `environment`, so traces from one "
            "deployment can be filtered and compared without staging traffic "
            "polluting production aggregates. Lowercased and punctuation-"
            "stripped on the way out, since Langfuse rejects other shapes. The "
            "trace `release` is the running Mewbo version and is not set here."
        ),
        examples=["dev"],
    )
    log_level: str = Field(
        "DEBUG",
        description="Logging verbosity. One of DEBUG, INFO, WARNING, ERROR, CRITICAL.",
        examples=["INFO"],
    )
    log_style: str = Field(
        "",
        description=(
            "Override for the CLI's terminal log format; prefer cli_log_style. "
            "Only the literal value 'dark' has any effect (dims the log line "
            "style for a dark background); anything else uses the plain format. "
            "Takes priority over cli_log_style when set; leave empty to use "
            "that instead."
        ),
        examples=[""],
    )
    cli_log_style: str = Field(
        "dark",
        description=(
            "CLI terminal log format: only the literal value 'dark' has any "
            "effect (dims the log line style for a dark background); any other "
            "value uses the plain format. Overridden by runtime.log_style when "
            "that's set."
        ),
        examples=["dark"],
    )
    preflight_enabled: bool = Field(
        False,
        description=("Run connectivity checks for LLM, Langfuse, and Home Assistant on startup."),
    )
    developer_mode: bool = Field(
        False,
        description=(
            "Enable developer mode. Unlocks graph-only (no-LLM) repository "
            "indexing so contributors can inspect AST graph construction "
            "without documentation generation. Read by the API, console, and CLI."
        ),
    )
    cache_dir: str = Field(
        "",
        description="Directory for tool caches. Defaults to $MEWBO_HOME/cache.",
        examples=["~/.mewbo/cache"],
        json_schema_extra={"x-protected": True},
    )
    session_dir: str = Field(
        "",
        description="Directory for session transcripts. Defaults to $MEWBO_HOME/sessions.",
        examples=["~/.mewbo/sessions"],
        json_schema_extra={"x-protected": True},
    )
    config_dir: str = Field(
        "",
        description="Root configuration directory. Defaults to $MEWBO_HOME.",
        examples=["~/.mewbo"],
        json_schema_extra={"x-protected": True},
    )
    result_export_dir: str = Field(
        "",
        description="Directory for large tool result exports. Empty to disable.",
        examples=["/tmp/mewbo-results"],
    )
    projects_home: str = Field(
        "",
        description="Directory for virtual project folders. Defaults to $MEWBO_HOME/projects.",
        examples=["~/.mewbo/projects"],
        json_schema_extra={"x-protected": True},
    )

    @field_validator("log_level", mode="before")
    @classmethod
    def _normalize_log_level(cls, value: Any) -> str:
        if not value:
            return "DEBUG"
        return str(value).strip().upper()

    @field_validator("cache_dir", "session_dir", "config_dir", mode="before")
    @classmethod
    def _normalize_paths(cls, value: Any, info: ValidationInfo) -> str:
        raw = str(value).strip() if value is not None else ""
        if raw:
            return raw
        home = resolve_mewbo_home()
        defaults = {
            "cache_dir": str(home / "cache"),
            "session_dir": str(home / "sessions"),
            "config_dir": str(home),
        }
        return defaults.get(info.field_name or "", str(home))

    @field_validator("projects_home", mode="before")
    @classmethod
    def _normalize_projects_home(cls, value: Any) -> str:
        raw = str(value).strip() if value is not None else ""
        if raw:
            return str(Path(raw).expanduser())
        return str(resolve_mewbo_home() / "projects")

    @field_validator("preflight_enabled", mode="before")
    @classmethod
    def _normalize_preflight_enabled(cls, value: Any) -> bool:
        return _coerce_bool(value, default=False)

    @field_validator("developer_mode", mode="before")
    @classmethod
    def _normalize_developer_mode(cls, value: Any) -> bool:
        return _coerce_bool(value, default=False)

SafetyConfig

Bases: BaseModel

Master switch for the operator-owned tool-call gate and session observer.

OFF by default. While off, .mewbo/policy/ and .mewbo/monitor/ are never read, no rule is built and no event is emitted — a deployment that never sets this section pays nothing. There is deliberately no per-project override and no other knob here: discovery path, rule precedence and the built-in self-protection rule are fixed by the plane itself, not config-tunable, because a knob that could redirect discovery or reorder rules would be a knob that could weaken the guardrail it configures.

Source code in packages/mewbo_core/src/mewbo_core/config.py
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
class SafetyConfig(BaseModel):
    """Master switch for the operator-owned tool-call gate and session observer.

    OFF by default. While off, ``.mewbo/policy/`` and ``.mewbo/monitor/`` are
    never read, no rule is built and no event is emitted — a deployment that
    never sets this section pays nothing. There is deliberately no per-project
    override and no other knob here: discovery path, rule precedence and the
    built-in self-protection rule are fixed by the plane itself, not
    config-tunable, because a knob that could redirect discovery or reorder
    rules would be a knob that could weaken the guardrail it configures.
    """

    model_config = ConfigDict(
        extra="forbid",
        json_schema_extra={"title": "Safety Plane", "x-group": "agent", "x-order": 3},
    )

    enabled: bool = Field(
        False,
        description=(
            "Master switch for the policy gate and session-budget observer "
            "read from .mewbo/policy/ and .mewbo/monitor/. OFF by default: no "
            "evaluation, no tokens, no latency until an operator turns this on."
        ),
    )

ScgConfig

Bases: BaseModel

Operator-facing knobs for the Source Capability Graph (agentic search).

Source code in packages/mewbo_core/src/mewbo_core/config.py
3390
3391
3392
3393
3394
3395
3396
3397
3398
3399
3400
3401
3402
3403
3404
3405
3406
3407
3408
class ScgConfig(BaseModel):
    """Operator-facing knobs for the Source Capability Graph (agentic search)."""

    model_config = ConfigDict(
        extra="forbid",
        json_schema_extra={"title": "SCG", "x-group": "workspace", "x-order": 3},
    )

    enabled: bool = Field(
        False,
        description=(
            "Master switch for the SCG feature (source mapping and "
            "orchestrated agentic-search runs)."
        ),
    )
    traversal: ScgTraversalConfig = Field(
        default_factory=lambda: ScgTraversalConfig.model_validate({}),
        description="Traversal defaults (the per-run search tier).",
    )

ScgTierModelsConfig

Bases: BaseModel

Per-tier model mapping: the tier picks the brain, not just the budget.

A tier maps to the LLM that drives the whole run (orchestrator session AND its probe sub-agents, which inherit the session model). An empty string falls back to llm.default_model. An explicit per-request model override (where the endpoint offers one) always wins over the tier map.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3347
3348
3349
3350
3351
3352
3353
3354
3355
3356
3357
3358
3359
3360
3361
3362
3363
3364
3365
3366
3367
3368
3369
class ScgTierModelsConfig(BaseModel):
    """Per-tier model mapping: the tier picks the brain, not just the budget.

    A tier maps to the LLM that drives the whole run (orchestrator session AND
    its probe sub-agents, which inherit the session model). An empty string
    falls back to ``llm.default_model``. An explicit per-request ``model``
    override (where the endpoint offers one) always wins over the tier map.
    """

    model_config = ConfigDict(extra="forbid", json_schema_extra={"title": "Tier models"})

    fast: str = Field(
        "openai/gpt-oss-120b",
        description="Model for `fast` tier runs (the tier still sets the low-latency budget).",
    )
    auto: str = Field(
        "openai/gpt-oss-120b",
        description="Model for `auto` tier runs (the tier still sets the balanced budget).",
    )
    deep: str = Field(
        "openai/gpt-oss-120b",
        description="Model for `deep` tier runs (the tier still sets the exhaustive budget).",
    )

ScgTraversalConfig

Bases: BaseModel

Traversal defaults for SCG search (the per-run tier budget knob).

Source code in packages/mewbo_core/src/mewbo_core/config.py
3372
3373
3374
3375
3376
3377
3378
3379
3380
3381
3382
3383
3384
3385
3386
3387
class ScgTraversalConfig(BaseModel):
    """Traversal defaults for SCG search (the per-run tier budget knob)."""

    model_config = ConfigDict(extra="forbid", json_schema_extra={"title": "Traversal"})

    default_tier: Literal["fast", "auto", "deep"] = Field(
        "auto",
        description=(
            "Default search tier, one budget knob over decomposition depth "
            "and probe fan-out. Overridable per run."
        ),
    )
    tier_models: ScgTierModelsConfig = Field(
        default_factory=lambda: ScgTierModelsConfig.model_validate({}),
        description="Which LLM each search tier runs on (fast/auto/deep).",
    )

SpeechConfig

Bases: BaseModel

Which gateway models handle speech, and which gateway serves them.

The three connection fields are all optional and all empty by default, because the common deployment has one gateway: speech falls back to llm.api_base/llm.api_key whenever these are blank, so an install that never writes a speech block still works. They exist for the deployment that genuinely splits the two, which is a real shape — a self-hosted synthesis backend beside a hosted chat provider — and refusing to represent it would only push the operator into running one gateway they do not want.

Source code in packages/mewbo_core/src/mewbo_core/config.py
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
829
830
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
class SpeechConfig(BaseModel):
    """Which gateway models handle speech, and which gateway serves them.

    The three connection fields are all optional and all empty by default,
    because the common deployment has one gateway: speech falls back to
    ``llm.api_base``/``llm.api_key`` whenever these are blank, so an install
    that never writes a ``speech`` block still works. They exist for the
    deployment that genuinely splits the two, which is a real shape — a
    self-hosted synthesis backend beside a hosted chat provider — and refusing
    to represent it would only push the operator into running one gateway they
    do not want.
    """

    model_config = ConfigDict(
        extra="forbid",
        validate_default=True,
        json_schema_extra={"title": "Speech", "x-group": "models", "x-order": 5},
    )

    api_base: str = Field(
        "",
        description=(
            "Base URL of the gateway that serves the speech models. Leave it "
            "empty to use the same gateway as the language models.\n\n"
            "Set this only when speech is served somewhere other than the "
            "endpoint under Language Model, such as a synthesis service running "
            "beside a hosted chat provider."
        ),
        examples=["", "https://my-litellm-proxy.example.com/v1"],
    )
    api_key: str = Field(
        "",
        description=(
            "Key for the speech gateway. Leave it empty to reuse the language "
            "model key.\n\n"
            "Only needed alongside a separate speech endpoint above. Setting one "
            "here without the other is almost always a mistake, because the key "
            "is then sent to the language model gateway that already had one."
        ),
        examples=["sk-xxxxxxxx"],
        json_schema_extra={"x-secret": True},
    )
    #: Mirrors ``mewbo_speech.gateway.DEFAULT_SPEECH_TIMEOUT``, and the two are
    #: pinned equal by ``tests/test_config_speech.py``. Duplicated rather than
    #: imported because core must not reach UP into a capability library, and
    #: the value cannot simply be left to the package: once this typed field
    #: exists, its default is what the accessor returns, so the package's own
    #: ``default=`` argument never runs again. A silent divergence here would
    #: change the deployed timeout while both files still read as correct.
    timeout: float = Field(
        90.0,
        gt=0,
        description=(
            "Seconds to wait for the speech gateway before giving up.\n\n"
            "Synthesis is not instant and scales with the length of the text: a "
            "sentence takes about half a second and a paragraph about four, so a "
            "short timeout cuts off long answers. Raising it has a cost too, "
            "because a request that is going to fail holds a server slot for the "
            "whole wait."
        ),
    )
    tts: SpeechTtsConfig = Field(
        default_factory=lambda: SpeechTtsConfig.model_validate({}),
        description="Reading an answer aloud.",
    )
    stt: SpeechSttConfig = Field(
        default_factory=lambda: SpeechSttConfig.model_validate({}),
        description="Turning a recording into text.",
    )

    @field_validator("api_base", "api_key", mode="before")
    @classmethod
    def _normalize_connection(cls, value: Any) -> str:
        """Trim both, because only blank means "fall back to the llm section".

        A key or URL pasted with a trailing newline is not blank, so it wins the
        ``or`` the reader falls back through and is then sent verbatim. The
        gateway answers that with an auth failure or a bad URL, neither of which
        points at the whitespace. Nothing downstream trims, so this is the only
        place it can happen.
        """
        return str(value).strip() if value is not None else ""

SpeechSttConfig

Bases: BaseModel

Speech to text: which model turns a recording into words.

Source code in packages/mewbo_core/src/mewbo_core/config.py
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
class SpeechSttConfig(BaseModel):
    """Speech to text: which model turns a recording into words."""

    model_config = ConfigDict(
        extra="forbid",
        validate_default=True,
        json_schema_extra={"title": "Speech to text"},
    )

    model: str = Field(
        "nova-3",
        description=(
            "Model that transcribes a recording. Leave it empty to turn dictation off.\n\n"
            "This is a model id your LLM gateway advertises. Transcription models are "
            "separate from both chat models and the text-to-speech model above, so the "
            "id here will not appear in the model picker used for answers."
        ),
        examples=["nova-3"],
    )
    @field_validator("model", mode="before")
    @classmethod
    def _normalize_model(cls, value: Any) -> str:
        return str(value).strip() if value is not None else ""

SpeechTtsConfig

Bases: BaseModel

Text to speech: which model reads an answer aloud, and in whose voice.

Source code in packages/mewbo_core/src/mewbo_core/config.py
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
class SpeechTtsConfig(BaseModel):
    """Text to speech: which model reads an answer aloud, and in whose voice."""

    model_config = ConfigDict(
        extra="forbid",
        validate_default=True,
        json_schema_extra={"title": "Text to speech"},
    )

    model: str = Field(
        "supertonic-3",
        description=(
            "Model that turns text into audio. Leave it empty to turn read aloud off.\n\n"
            "This is a model id your LLM gateway advertises, and speech models are a "
            "separate family from the chat models the answer itself runs on. A chat "
            "model named here is refused by the gateway at the moment someone presses "
            "play, not when this page is saved."
        ),
        examples=["supertonic-3", "supertonic-3-hd"],
    )
    voice: str = Field(
        "nova",
        description=(
            "Voice the reader speaks in. Type the name your gateway knows it by.\n\n"
            "This was once a fixed list of eleven, which was wrong for a self-hosted "
            "gateway: a backend can carry its own trained voice style, and a name "
            "absent from that list was refused here before the gateway ever saw it. "
            "The gateway decides what a voice is. A name it does not know fails the "
            "request when someone presses play, the same way an unknown model does."
        ),
        examples=["nova", "alloy", "shimmer"],
    )
    response_format: Literal["wav", "flac"] = Field(
        "wav",
        description=(
            "Audio format the gateway returns. WAV is the safe default and FLAC is "
            "the same audio in a smaller file.\n\n"
            "Only these two are accepted. Asking for MP3, Opus, AAC or raw PCM fails "
            "the request, so they are not offered here."
        ),
    )

    @field_validator("model", mode="before")
    @classmethod
    def _normalize_model(cls, value: Any) -> str:
        return str(value).strip() if value is not None else ""

    DEFAULT_VOICE: ClassVar[str] = "nova"

    @field_validator("voice", mode="before")
    @classmethod
    def _normalize_voice(cls, value: Any) -> str:
        """Trim padding, and resolve an empty value to the default.

        Case is PRESERVED. Folding it presumed every voice name was OpenAI's;
        an operator-defined style may be capitalised, and lowercasing it sends
        the gateway a name it need not recognise.

        Blank resolves rather than travelling as empty, because the gateway
        answers an omitted voice with an opaque 500 — the one failure this
        field can still prevent locally, now that which names EXIST is the
        gateway's fact rather than ours.
        """
        if value is None:
            return cls.DEFAULT_VOICE
        return str(value).strip() or cls.DEFAULT_VOICE

StorageConfig

Bases: EnvOverridable

Session storage backend configuration.

Source code in packages/mewbo_core/src/mewbo_core/config.py
2947
2948
2949
2950
2951
2952
2953
2954
2955
2956
2957
2958
2959
2960
2961
2962
2963
2964
2965
2966
2967
2968
2969
2970
2971
2972
2973
2974
2975
2976
2977
2978
2979
2980
class StorageConfig(EnvOverridable):
    """Session storage backend configuration."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={
            "title": "Storage",
            "x-group": "server",
            "x-order": 2,
            "x-advanced": True,
        },
    )

    driver: str = Field(
        "json",
        description=(
            "Storage driver: 'json' (filesystem) or 'mongodb'. "
            "Overridden by the MEWBO_STORAGE_DRIVER environment variable."
        ),
        examples=["json", "mongodb"],
        json_schema_extra={"x-env-var": "MEWBO_STORAGE_DRIVER"},
    )
    mongodb: MongoDBConfig = Field(
        default_factory=lambda: MongoDBConfig.model_validate({}),
        description="MongoDB connection settings (used when driver is 'mongodb').",
    )

    @field_validator("driver", mode="before")
    @classmethod
    def _normalize_driver(cls, value: Any) -> str:
        raw = str(value).strip().lower() if value else "json"
        if raw not in {"json", "mongodb"}:
            raise ValueError(f"Unknown storage driver {raw!r}. Expected 'json' or 'mongodb'.")
        return raw

TokenBudgetConfig

Bases: BaseModel

Token budget and auto-compaction thresholds.

Source code in packages/mewbo_core/src/mewbo_core/config.py
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
class TokenBudgetConfig(BaseModel):
    """Token budget and auto-compaction thresholds."""

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Token Budget", "x-group": "models", "x-order": 2},
    )

    default_context_window: int = Field(
        128000,
        description=(
            "Default context window size in tokens used when the "
            "model is not listed in model_context_windows."
        ),
        examples=[128000],
    )
    auto_compact_threshold: float = Field(
        0.8,
        description=(
            "Fraction of the context window (0.0-1.0) that triggers "
            "automatic conversation compaction."
        ),
        examples=[0.8],
    )
    model_context_windows: dict[str, int] = Field(
        default_factory=dict,
        description=(
            "Override only: per-model context window in tokens. Keys are model "
            "names (with or without provider prefix). The authoritative source "
            "is LiteLLM's model catalogue; populate this only to cap below the "
            "model's real max, or for models LiteLLM doesn't know yet — which "
            "includes every model renamed by a proxy, since the catalogue is "
            "keyed on the real name. A model the catalogue cannot resolve falls "
            "back to default_context_window and logs a one-time warning naming "
            "it; that number then drives compaction and every utilisation "
            "reading, so an unlisted model silently budgets against a guess."
        ),
    )

    @field_validator("default_context_window", mode="before")
    @classmethod
    def _normalize_context_window(cls, value: Any) -> int:
        try:
            parsed = int(value)
        except (TypeError, ValueError):
            return 128000
        return max(parsed, 1)

    @field_validator("auto_compact_threshold", mode="before")
    @classmethod
    def _normalize_compact_threshold(cls, value: Any) -> float:
        try:
            parsed = float(value)
        except (TypeError, ValueError):
            return 0.8
        return min(max(parsed, 0.0), 1.0)

    @field_validator("model_context_windows", mode="before")
    @classmethod
    def _normalize_model_context_windows(cls, value: Any) -> dict[str, int]:
        if not isinstance(value, dict):
            return {}
        cleaned: dict[str, int] = {}
        for key, raw in value.items():
            try:
                cleaned[str(key)] = max(int(raw), 1)
            except (TypeError, ValueError):
                continue
        return cleaned

ToolSearchConfig

Bases: BaseModel

Deferred tool loading via on-demand schema fetching.

When mode='on', MCP tool schemas (and any spec with metadata.deferred=True) are stripped from the initial bind_tools call and surfaced to the model by name only via <available-deferred-tools>. The model fetches schemas it actually needs by calling the built-in tool_search tool. Mirrors Claude Code's ToolSearchTool mechanism, saving substantial context tokens on sessions with many MCP servers connected.

Source code in packages/mewbo_core/src/mewbo_core/config.py
2145
2146
2147
2148
2149
2150
2151
2152
2153
2154
2155
2156
2157
2158
2159
2160
2161
2162
2163
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
2178
2179
2180
2181
2182
class ToolSearchConfig(BaseModel):
    """Deferred tool loading via on-demand schema fetching.

    When ``mode='on'``, MCP tool schemas (and any spec with
    ``metadata.deferred=True``) are stripped from the initial ``bind_tools``
    call and surfaced to the model by name only via
    ``<available-deferred-tools>``. The model fetches schemas it actually
    needs by calling the built-in ``tool_search`` tool. Mirrors Claude
    Code's ``ToolSearchTool`` mechanism, saving substantial context tokens
    on sessions with many MCP servers connected.
    """

    model_config = ConfigDict(extra="forbid", json_schema_extra={"title": "Tool Search"})

    mode: Literal["off", "on", "auto"] = Field(
        "on",
        description=(
            "'on' (the default) always defers MCP tools and any spec with "
            "metadata.deferred=True; the model loads schemas on demand via "
            "tool_search, so no user MCP tool occupies the context window "
            "until it is actually needed. 'off' keeps every tool's schema in "
            "the initial bind. 'auto' defers only when the number of "
            "deferrable tools exceeds auto_threshold. Note that 'on' costs a "
            "zero-MCP session nothing: deferral only engages when the "
            "deferrable set is non-empty, so the two modes bind an identical "
            "list there. The range where they differ is 1..auto_threshold "
            "tools, where 'auto' spends ~240 tokens per tool every turn."
        ),
    )
    auto_threshold: int = Field(
        25,
        ge=0,
        description=(
            "In 'auto' mode, defer tool schemas only when more than this "
            "many deferrable tools (MCP + metadata.deferred specs) are "
            "registered. Ignored when mode is 'off' or 'on'."
        ),
    )

TriggersConfig

Bases: BaseModel

Reverse-invocation trigger subsystem.

The durable peer of the sub-agent hypervisor: a background watcher that fires time / cron / CI / forge-PR / webhook triggers and re-invokes the sessions that armed them. OFF by default (enabled=False), so the feature is un-enableable until an operator turns it on and a stock deployment pays nothing. The lower half of this section is the admission policy the schedule_trigger tool + the arm route enforce (mirrors mewbo_core.triggers.policy.TriggerPolicy field-for-field; to_policy builds one).

Source code in packages/mewbo_core/src/mewbo_core/config.py
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
class TriggersConfig(BaseModel):
    """Reverse-invocation trigger subsystem.

    The durable peer of the sub-agent hypervisor: a background watcher that
    fires time / cron / CI / forge-PR / webhook triggers and re-invokes the
    sessions that armed them. OFF by default (``enabled=False``), so the feature
    is un-enableable until an operator turns it on and a stock deployment pays
    nothing. The lower half of this section is the admission policy the
    ``schedule_trigger`` tool + the arm route enforce (mirrors
    ``mewbo_core.triggers.policy.TriggerPolicy`` field-for-field; ``to_policy``
    builds one).
    """

    model_config = ConfigDict(
        validate_default=True,
        json_schema_extra={"title": "Triggers", "x-group": "automation", "x-order": 1},
    )

    enabled: bool = Field(
        False,
        description=(
            "Turn the trigger watcher on. Nothing fires until you do.\n\n"
            "A trigger is how Mewbo starts a session later, on its own, with nobody "
            "watching: at a set time, on a repeating schedule, or when a CI run "
            "finishes, a pull request changes, or a webhook calls in. The watcher is "
            "the background loop that notices those moments and wakes the session "
            "that asked to be woken. While it is off, the trigger routes still work, "
            "so a session can arm a trigger and you can list, pause, or cancel it, "
            "but no trigger ever fires. Armed triggers simply wait until you turn "
            "the watcher on."
        ),
    )
    tick_interval_seconds: float = Field(
        5.0,
        ge=1.0,
        description=(
            "How often the watcher wakes up to look at the schedule, in seconds.\n\n"
            "On each pass it expires the triggers whose deadline has gone by and "
            "fires the time and cron triggers that have come due. A shorter interval "
            "wakes a session closer to the moment it asked for; a longer one costs "
            "the server less. This is also the cadence at which the forge poll below "
            "gets a chance to run."
        ),
    )
    poll_interval_seconds: float = Field(
        60.0,
        ge=5.0,
        description=(
            "How often the watcher asks the forge about CI runs and pull requests, "
            "in seconds.\n\n"
            "Time and cron triggers can be judged from the clock alone, but "
            "`ci.workflow` and `forge.pr` triggers cannot: the watcher has to call "
            "the forge's REST API to see what changed. Those calls are rate-limited "
            "and cost a round trip each, so they run on this deliberately coarser "
            "cadence rather than on every pass. Raise it if you are bumping into API "
            "limits; lower it if you want CI results picked up sooner."
        ),
    )
    max_consecutive_failures: int = Field(
        5,
        ge=1,
        description=(
            "How many errors in a row one trigger may hit before it is given up "
            "on.\n\n"
            "When a fire or a forge poll raises, the watcher records the error on "
            "the trigger and leaves it armed, so a passing outage never throws away "
            "a schedule. Once a trigger has failed this many times back to back "
            "without a single success in between, the watcher stops retrying it and "
            "moves it to `failed`. Any success resets the count to zero."
        ),
    )
    max_armed_per_session: int = Field(
        20,
        ge=1,
        description=(
            "The most triggers one session may have armed at the same time.\n\n"
            "Triggers are armed by the agent from inside a session, so this ceiling "
            "is what keeps a single session from filling the schedule with wakes. An "
            "attempt to arm one past the limit is refused, and the agent is told why. "
            "Cancelling a trigger, or letting one finish, frees the slot again."
        ),
    )
    max_fires_cap: int = Field(
        100,
        ge=1,
        description=(
            "The ceiling on how many times any single trigger may fire.\n\n"
            "A repeating trigger, a cron schedule for instance, can name its own "
            "`max_fires` limit when it is armed. This is the ceiling on that request: "
            "an attempt to arm a trigger asking for more is refused. A trigger that "
            "reaches its own limit completes and stops firing."
        ),
    )
    default_expiry_days: float = Field(
        7.0,
        gt=0.0,
        description=(
            "How long an armed trigger lives when it names no expiry of its own, in "
            "days.\n\n"
            "Every trigger expires eventually, so that a wake nobody remembers "
            "arming cannot linger forever. When the agent arms one without setting "
            "an expiry date, this many days from the moment of arming is stamped on "
            "it. Once that moment passes, the watcher expires the trigger instead of "
            "firing it."
        ),
    )
    cron_min_interval_seconds: int = Field(
        60,
        ge=1,
        description=(
            "The shortest gap allowed between two fires of a cron trigger, in "
            "seconds.\n\n"
            "A cron expression can be written to fire far more often than a session "
            "is worth waking, so this is the floor. When a cron trigger is armed, the "
            "gap between its first two fires is measured, and a schedule tighter than "
            "this is rejected there and then rather than being throttled later."
        ),
    )
    webhook_payload_max_bytes: int = Field(
        200_000,
        ge=1,
        description=(
            "How much of an incoming webhook body the woken session gets to see, in "
            "bytes.\n\n"
            "A webhook can carry a large payload, and all of it becomes context the "
            "session has to read. A body bigger than this is truncated rather than "
            "rejected: the call still fires the trigger, the session receives the "
            "first part of the body, and it is told the payload was cut short. When "
            "a signature is configured, it is checked against the whole body before "
            "any truncation happens."
        ),
    )

    def to_policy(self) -> Any:
        """Build the ``TriggerPolicy`` these fields describe.

        Lazy import keeps this config module free of any dependency on the
        triggers domain package (and sidesteps an import cycle, since the
        trigger store reads ``get_config_value`` from here).
        """
        from datetime import timedelta

        from mewbo_core.triggers.policy import TriggerPolicy

        return TriggerPolicy(
            max_armed_per_session=self.max_armed_per_session,
            max_fires_cap=self.max_fires_cap,
            default_expiry=timedelta(days=self.default_expiry_days),
            cron_min_interval_seconds=self.cron_min_interval_seconds,
            webhook_payload_max_bytes=self.webhook_payload_max_bytes,
        )

to_policy() -> Any

Build the TriggerPolicy these fields describe.

Lazy import keeps this config module free of any dependency on the triggers domain package (and sidesteps an import cycle, since the trigger store reads get_config_value from here).

Source code in packages/mewbo_core/src/mewbo_core/config.py
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
def to_policy(self) -> Any:
    """Build the ``TriggerPolicy`` these fields describe.

    Lazy import keeps this config module free of any dependency on the
    triggers domain package (and sidesteps an import cycle, since the
    trigger store reads ``get_config_value`` from here).
    """
    from datetime import timedelta

    from mewbo_core.triggers.policy import TriggerPolicy

    return TriggerPolicy(
        max_armed_per_session=self.max_armed_per_session,
        max_fires_cap=self.max_fires_cap,
        default_expiry=timedelta(days=self.default_expiry_days),
        cron_min_interval_seconds=self.cron_min_interval_seconds,
        webhook_payload_max_bytes=self.webhook_payload_max_bytes,
    )

UntrustedCwdRegistry

Directories whose own .mcp.json must never enter the merged config.

A working directory is normally a developer's own project, so its .mcp.json legitimately contributes MCP servers at the highest priority tier. That stops being true the moment the directory holds content this deployment did not author — a repository cloned for indexing being the worked example: its .mcp.json is attacker-supplied, and admitting it both OVERRIDES a same-named operator server and, because _deep_merge recurses, lets a file naming only env keep the operator's command and credential while adding a variable of its own. Merge order cannot fix that — the tier has to be excluded, not out-prioritised.

The caller that CREATED the directory is the only one that knows this, so it registers the root here once. Every resolution of the merged config then excludes it — the initial registry build and every re-resolution a tool runner performs at invocation time alike — without the knowledge having to be threaded through each of them. Registration is by explicit path, never by pattern-matching one: a heuristic on the directory name would be exactly the accidental control this exists to remove.

Registering a root covers that directory and everything beneath it, so one registration of a clone ROOT covers every job that clones into it.

Cost: O(registered roots) per lookup, on a path already doing file I/O.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4533
4534
4535
4536
4537
4538
4539
4540
4541
4542
4543
4544
4545
4546
4547
4548
4549
4550
4551
4552
4553
4554
4555
4556
4557
4558
4559
4560
4561
4562
4563
4564
4565
4566
4567
4568
4569
4570
4571
4572
4573
4574
4575
4576
4577
4578
4579
4580
4581
4582
4583
4584
4585
4586
4587
4588
4589
4590
4591
4592
4593
4594
4595
4596
4597
4598
4599
4600
4601
4602
4603
4604
4605
4606
4607
4608
4609
4610
4611
4612
4613
4614
class UntrustedCwdRegistry:
    """Directories whose own ``.mcp.json`` must never enter the merged config.

    A working directory is normally a developer's own project, so its
    ``.mcp.json`` legitimately contributes MCP servers at the highest priority
    tier. That stops being true the moment the directory holds content this
    deployment did not author — a repository cloned for indexing being the
    worked example: its ``.mcp.json`` is attacker-supplied, and admitting it
    both OVERRIDES a same-named operator server and, because ``_deep_merge``
    recurses, lets a file naming only ``env`` keep the operator's ``command``
    and credential while adding a variable of its own. Merge order cannot fix
    that — the tier has to be excluded, not out-prioritised.

    The caller that CREATED the directory is the only one that knows this, so
    it registers the root here once. Every resolution of the merged config then
    excludes it — the initial registry build and every re-resolution a tool
    runner performs at invocation time alike — without the knowledge having to
    be threaded through each of them. Registration is by explicit path, never
    by pattern-matching one: a heuristic on the directory name would be exactly
    the accidental control this exists to remove.

    Registering a root covers that directory and everything beneath it, so one
    registration of a clone ROOT covers every job that clones into it.

    Cost: ``O(registered roots)`` per lookup, on a path already doing file I/O.
    """

    def __init__(self) -> None:
        """Initialize an empty, lock-guarded registry."""
        self._lock = threading.Lock()
        self._roots: set[Path] = set()

    @staticmethod
    def _normalize(path: str | Path) -> Path | None:
        """Resolve *path* for comparison, or ``None`` when it cannot be read."""
        try:
            return Path(path).expanduser().resolve()
        except (OSError, RuntimeError, ValueError):
            return None

    def register(self, path: str | Path) -> None:
        """Mark *path* and everything under it as an untrusted working directory."""
        resolved = self._normalize(path)
        if resolved is None:
            _logger.warning("Could not mark %s as an untrusted cwd: unresolvable path", path)
            return
        with self._lock:
            self._roots.add(resolved)

    def unregister(self, path: str | Path) -> None:
        """Drop a previously registered root (no-op when absent)."""
        resolved = self._normalize(path)
        if resolved is None:
            return
        with self._lock:
            self._roots.discard(resolved)

    def clear(self) -> None:
        """Drop every registered root (tests)."""
        with self._lock:
            self._roots.clear()

    def roots(self) -> tuple[str, ...]:
        """Return the registered roots, sorted, for diagnostics."""
        with self._lock:
            return tuple(sorted(str(root) for root in self._roots))

    def contains(self, path: str | Path | None) -> bool:
        """True when *path* is at or below a registered untrusted root.

        A path that cannot be resolved is reported as untrusted: this gate is
        the one place where failing closed costs a feature and failing open
        spawns a process.
        """
        with self._lock:
            if not self._roots:
                return False
            roots = set(self._roots)
        resolved = self._normalize(path if path is not None else Path.cwd())
        if resolved is None:
            return True
        return any(resolved == root or root in resolved.parents for root in roots)

__init__() -> None

Initialize an empty, lock-guarded registry.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4560
4561
4562
4563
def __init__(self) -> None:
    """Initialize an empty, lock-guarded registry."""
    self._lock = threading.Lock()
    self._roots: set[Path] = set()

clear() -> None

Drop every registered root (tests).

Source code in packages/mewbo_core/src/mewbo_core/config.py
4590
4591
4592
4593
def clear(self) -> None:
    """Drop every registered root (tests)."""
    with self._lock:
        self._roots.clear()

contains(path: str | Path | None) -> bool

True when path is at or below a registered untrusted root.

A path that cannot be resolved is reported as untrusted: this gate is the one place where failing closed costs a feature and failing open spawns a process.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4600
4601
4602
4603
4604
4605
4606
4607
4608
4609
4610
4611
4612
4613
4614
def contains(self, path: str | Path | None) -> bool:
    """True when *path* is at or below a registered untrusted root.

    A path that cannot be resolved is reported as untrusted: this gate is
    the one place where failing closed costs a feature and failing open
    spawns a process.
    """
    with self._lock:
        if not self._roots:
            return False
        roots = set(self._roots)
    resolved = self._normalize(path if path is not None else Path.cwd())
    if resolved is None:
        return True
    return any(resolved == root or root in resolved.parents for root in roots)

register(path: str | Path) -> None

Mark path and everything under it as an untrusted working directory.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4573
4574
4575
4576
4577
4578
4579
4580
def register(self, path: str | Path) -> None:
    """Mark *path* and everything under it as an untrusted working directory."""
    resolved = self._normalize(path)
    if resolved is None:
        _logger.warning("Could not mark %s as an untrusted cwd: unresolvable path", path)
        return
    with self._lock:
        self._roots.add(resolved)

roots() -> tuple[str, ...]

Return the registered roots, sorted, for diagnostics.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4595
4596
4597
4598
def roots(self) -> tuple[str, ...]:
    """Return the registered roots, sorted, for diagnostics."""
    with self._lock:
        return tuple(sorted(str(root) for root in self._roots))

unregister(path: str | Path) -> None

Drop a previously registered root (no-op when absent).

Source code in packages/mewbo_core/src/mewbo_core/config.py
4582
4583
4584
4585
4586
4587
4588
def unregister(self, path: str | Path) -> None:
    """Drop a previously registered root (no-op when absent)."""
    resolved = self._normalize(path)
    if resolved is None:
        return
    with self._lock:
        self._roots.discard(resolved)

WebIdeConfig

Bases: EnvOverridable

Config for the per-session code-server "Open in Web IDE" feature.

Source code in packages/mewbo_core/src/mewbo_core/config.py
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
2028
2029
2030
2031
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
2100
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
2115
2116
2117
2118
class WebIdeConfig(EnvOverridable):
    """Config for the per-session code-server "Open in Web IDE" feature."""

    # ``validate_default=True`` so ``broker_url``'s normalizer still runs when
    # the key is absent from app.json; without it the validator only runs for a
    # key that is already present.
    model_config = ConfigDict(
        extra="forbid", validate_default=True, json_schema_extra={"title": "Web IDE"}
    )

    enabled: bool = Field(
        default=False,
        description=(
            "Turn on the 'Open in Web IDE' feature (per-session code-server "
            "containers via Docker). Also requires a MongoDB-backed session "
            "store; toggling this needs an API process restart to take effect "
            "since the /api/ide routes are registered at startup."
        ),
    )
    image: str = Field(
        default="codercom/code-server:latest",
        description="Docker image used to launch each session's code-server container.",
    )
    default_lifetime_hours: int = Field(
        default=1,
        ge=1,
        le=24,
        description=(
            "Hours a new Web IDE container stays up before it self-terminates, "
            "unless the session extends it first."
        ),
    )
    max_lifetime_hours: int = Field(
        default=8,
        ge=1,
        le=168,
        description=(
            "Hard ceiling on a session's total Web IDE lifetime across all "
            "extensions; a request to extend past this is rejected."
        ),
    )
    cpus: float = Field(
        default=1.0,
        ge=0.1,
        le=16.0,
        description=(
            "CPU core limit for each Web IDE container, e.g. 1.0 = one core "
            "(maps to Docker's --cpus / nano_cpus)."
        ),
    )
    memory: str = Field(
        default="1g",
        pattern=r"^\d+[mgMG]$",
        description=(
            "Memory limit for each Web IDE container, in Docker's --memory "
            "syntax: digits followed by m or g, e.g. '1g' or '512m'."
        ),
    )
    pids_limit: int = Field(
        default=512,
        ge=64,
        le=4096,
        description=(
            "Maximum number of processes/threads allowed inside a Web IDE "
            "container; bounds a runaway process from exhausting the host."
        ),
    )
    network: str = Field(
        default="mewbo-ide",
        pattern=r"^[a-zA-Z0-9_-]+$",
        description=(
            "Docker network each Web IDE container joins. Must be the same "
            "network the ide-proxy is attached to, or the proxy can't reach "
            "the container."
        ),
    )
    proxy_url: str = Field(
        default="http://127.0.0.1:5126",
        description=(
            "Base URL the API uses to reach the ide-proxy for readiness "
            "probes. The default suits a host-networked API, where the proxy "
            "is published on loopback 127.0.0.1:5126. A bridge-networked API "
            "(e.g. one joined to extra Docker networks via a compose override) "
            "cannot reach that loopback: attach it to the `network` above and "
            "point this at the proxy's in-network name, e.g. "
            "http://mewbo-ide-proxy:8080."
        ),
    )
    state_dir: str = Field(
        default="/tmp/mewbo-ide",
        description=(
            "Host directory where each Web IDE container's expiry-deadline file "
            "is written; the container's internal watchdog reads it to "
            "self-terminate on schedule."
        ),
    )
    broker_url: str = Field(
        default="",
        description=(
            "Base URL of the IDE broker service, e.g. http://127.0.0.1:5128. "
            "When this is set and the MEWBO_IDE_BROKER_TOKEN environment "
            "variable is present, the API delegates every container operation "
            "to the broker and needs no Docker access of its own — the broker "
            "holds the socket and builds each container spec from its own "
            "configuration. Leave empty to drive Docker directly from the API "
            "process, which requires giving that process the socket. The "
            "MEWBO_IDE_BROKER_URL environment variable overrides this value; "
            "the shared secret is read from the environment only, never from "
            "this file."
        ),
        json_schema_extra={"x-env-var": "MEWBO_IDE_BROKER_URL"},
    )

    @field_validator("broker_url", mode="before")
    @classmethod
    def _normalize_broker_url(cls, value: Any) -> str:
        return str(value).strip() if value else ""

WikiConfig

Bases: BaseModel

Operator-facing knobs for the wiki subsystem.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3263
3264
3265
3266
3267
3268
3269
3270
3271
3272
3273
3274
3275
3276
3277
3278
3279
3280
3281
3282
3283
3284
3285
3286
3287
3288
3289
3290
3291
3292
3293
3294
3295
3296
3297
3298
3299
3300
3301
3302
3303
3304
3305
3306
3307
3308
3309
3310
3311
3312
3313
3314
3315
3316
3317
3318
3319
3320
3321
3322
3323
3324
3325
3326
3327
3328
3329
3330
3331
3332
3333
3334
3335
3336
3337
3338
3339
3340
class WikiConfig(BaseModel):
    """Operator-facing knobs for the wiki subsystem."""

    model_config = ConfigDict(
        extra="forbid",
        json_schema_extra={"title": "Wiki", "x-group": "workspace", "x-order": 2},
    )

    default_model: str = Field(
        "",
        description=(
            "Model the wiki picker pre-selects for indexing (the wizard). "
            "Overrides ``llm.default_model`` for the wizard. Empty string "
            "means: fall back to ``llm.default_model``."
        ),
        examples=["openai/gpt-5.4-mini"],
    )
    default_qa_model: str = Field(
        "",
        description=(
            "Model the Q&A composer pre-selects. Typically smaller/faster "
            "than ``default_model`` because Q&A is a tight read-only loop "
            "where latency matters more than depth. Empty string means: "
            "fall back to ``default_model``, then to ``llm.default_model``."
        ),
        examples=["openai/gpt-5.4-nano"],
    )
    default_qa_fast_model: str = Field(
        "",
        description=(
            "Model the Q&A composer's fast mode pre-selects. Fast mode holds "
            "the retrieval surface itself with no probe fan-out, so a "
            "smaller/faster model than default_qa_model often suffices. Empty "
            "string means: fall back to default_qa_model, then default_model, "
            "then llm.default_model."
        ),
        examples=["openai/gpt-5.4-nano"],
    )
    qa_fast_step_budget: int = Field(
        15,
        ge=1,
        description=(
            "Tool-step ceiling for a fast-mode Q&A run — the root itself "
            "retrieves with no probe fan-out, so it needs far fewer steps than "
            "deep mode's session-wide budget. Budgets are config-tunable, "
            "never hardcoded."
        ),
    )
    default_depth: Literal["", "comprehensive", "concise"] = Field(
        "",
        description=(
            "Indexing depth the wizard pre-selects. Empty string means: use "
            "the wizard's own default (``comprehensive``)."
        ),
    )
    default_language: str = Field(
        "",
        description=(
            "Language code the wizard pre-selects (e.g. ``en``, ``es``). "
            "Empty string means: use the wizard's own default."
        ),
    )
    embedding: WikiEmbeddingConfig = Field(
        default_factory=lambda: WikiEmbeddingConfig.model_validate({}),
        description="Embedding settings for the wiki indexer.",
    )
    memory: WikiMemoryConfig = Field(
        default_factory=lambda: WikiMemoryConfig.model_validate({}),
        description="Multiplex memory-layer knobs (atomic insights over the graph).",
    )
    refresh: WikiRefreshConfig = Field(
        default_factory=lambda: WikiRefreshConfig.model_validate({}),
        description="On-demand incremental-refresh thresholds.",
    )
    phase_timeouts: WikiPhaseTimeoutsConfig = Field(
        default_factory=lambda: WikiPhaseTimeoutsConfig.model_validate({}),
        description="How long each long-running indexing phase may run before it is called wedged.",
    )

WikiEmbeddingConfig

Bases: BaseModel

Embedding settings for the wiki indexer.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3047
3048
3049
3050
3051
3052
3053
3054
3055
3056
3057
3058
3059
3060
3061
3062
3063
3064
3065
3066
3067
3068
3069
3070
3071
3072
3073
3074
3075
3076
3077
3078
3079
3080
3081
3082
3083
3084
3085
3086
3087
3088
3089
3090
3091
3092
3093
3094
3095
3096
3097
3098
3099
3100
3101
3102
3103
3104
3105
3106
3107
3108
3109
3110
3111
3112
3113
3114
3115
3116
3117
3118
3119
3120
3121
3122
3123
3124
3125
3126
3127
3128
3129
3130
3131
3132
3133
3134
3135
3136
3137
class WikiEmbeddingConfig(BaseModel):
    """Embedding settings for the wiki indexer."""

    model_config = ConfigDict(extra="forbid", json_schema_extra={"title": "Embedding"})

    enabled: bool = Field(
        True,
        description=(
            "When false, ``wiki_build_graph`` skips embedding generation and "
            "retrieval falls back to BM25 + graph traversal only."
        ),
        examples=[True, False],
    )
    model: str = Field(
        "openai/text-embedding-3-small",
        description=(
            "Embedding model ID routed through the LLM proxy. Must support "
            "the OpenAI ``/v1/embeddings`` shape (LiteLLM normalises Gemini "
            "and others to this shape). Pin a fast model here to speed up "
            "indexing, since embedding is per-node and runs synchronously."
        ),
        examples=["openai/gemini-embedding-001", "openai/text-embedding-3-large"],
    )
    # No ``dimensions`` knob — the Embedder reads ``len(vector)`` off the
    # actual response, which is always accurate. LangChain's
    # ``OpenAIEmbeddings`` only sends a ``dimensions`` parameter when the
    # caller wants truncation (OpenAI v3 family only); we don't expose
    # that here to keep the surface small.
    batch_size: int = Field(
        64,
        description=(
            "Number of graph nodes embedded per API call during indexing. "
            "Embedding throughput per connection is roughly flat regardless of "
            "batch size — a larger batch just takes proportionally longer per "
            "call — so use `concurrency` to speed up indexing, and use this "
            "knob only to stay under the provider's request payload limit."
        ),
        examples=[32, 64, 128],
        gt=0,
    )
    concurrency: int = Field(
        4,
        description=(
            "Maximum number of embedding requests issued to the provider "
            "concurrently during indexing. Embedding is I/O-bound — the "
            "indexer spends its time waiting on the network, not on CPU — so "
            "this, not `batch_size`, is what determines indexing speed. Keep "
            "it conservative: issuing too many requests at once trades "
            "throughput for HTTP 429 responses and retries. Raise it only "
            "after confirming headroom against your provider's rate limits."
        ),
        examples=[4, 8, 16],
        gt=0,
    )
    requests_per_minute: int | None = Field(
        None,
        description=(
            "Optional ceiling on embedding requests issued per minute. When "
            "set, the indexer paces itself against this budget — queuing "
            "work rather than firing it — instead of relying on `concurrency` "
            "alone. Leave unset to rely on `concurrency` plus automatic "
            "backoff on 429 responses."
        ),
        examples=[500, 3000],
        gt=0,
    )
    tokens_per_minute: int | None = Field(
        None,
        description=(
            "Optional ceiling on embedding tokens processed per minute, "
            "estimated from input text length. Paced the same way as "
            "`requests_per_minute`. Leave unset to rely on `concurrency` plus "
            "automatic backoff on 429 responses."
        ),
        examples=[1000000, 5000000],
        gt=0,
    )
    max_retries: int = Field(
        5,
        description=(
            "Number of times a rate-limited (HTTP 429) embedding request is "
            "retried before the indexing job fails. A retry honours the "
            "provider's `Retry-After` header when present, and falls back to "
            "exponential backoff with jitter otherwise; sustained "
            "rate-limiting also reduces `concurrency` for the remainder of "
            "the run so a rate-limited pass degrades to slower rather than "
            "failing outright."
        ),
        examples=[3, 5, 8],
        ge=0,
    )

WikiMemoryConfig

Bases: BaseModel

Knobs for the multiplex memory layer (atomic insights over the graph).

Source code in packages/mewbo_core/src/mewbo_core/config.py
3140
3141
3142
3143
3144
3145
3146
3147
3148
3149
3150
3151
3152
3153
3154
3155
3156
3157
3158
3159
3160
3161
3162
3163
3164
3165
3166
3167
3168
3169
3170
3171
3172
3173
3174
3175
3176
3177
3178
3179
3180
3181
3182
3183
3184
3185
3186
class WikiMemoryConfig(BaseModel):
    """Knobs for the multiplex memory layer (atomic insights over the graph)."""

    model_config = ConfigDict(extra="forbid", json_schema_extra={"title": "Memory"})

    enabled: bool = Field(
        True, description="Master switch for the memory layer (gates wiki_submit_insight)."
    )
    model: str = Field(
        "",
        description=(
            "Chat model for condense + LLM dedup on the human/REST/MCP path. "
            "Empty → falls back to default_qa_model, then default_model."
        ),
    )
    # The ceiling mirrors ``mewbo_graph.wiki.memory_types.MAX_INSIGHT_CHARS`` by
    # VALUE, not import: graph sits ABOVE core, so this module cannot reach it.
    # Kept in lockstep by ``TestMaxInsightCharsCeiling`` in
    # ``tests/wiki/test_memory_config_ceiling.py``, which asserts this bound
    # against the canonical constant directly — the same arrangement
    # ``_normalize_projects`` uses for the reserved project name.
    max_insight_chars: int = Field(
        200,
        gt=0,
        le=200,
        description=(
            "Hard cap on a memory note's length. Lowering this shortens notes; "
            "it cannot be raised above 200, which is the length the stored note "
            "model itself declares — a larger value would let a note past this "
            "check and then fail validation as it was written."
        ),
    )
    max_anchors: int = Field(8, description="Max code anchors per note.")
    dedup_k: int = Field(5, description="kNN candidate window for fuzzy + LLM dedup tiers.")
    dedup_cosine: float = Field(0.6, description="Cosine floor for the LLM dedup tier.")
    fuzzy_jaccard: float = Field(0.85, description="Jaccard floor for the fuzzy dedup tier.")
    fusion_w_ppr: float = Field(
        0.1,
        description=(
            "Weight applied to a code node's score when it's surfaced only by "
            "following a memory note's anchor rather than direct text/code "
            "search. Raise it to rank memory-anchored context higher relative "
            "to direct hits; lower it toward 0 to favor direct hits."
        ),
    )
    hub_degree: int = Field(50, description="Degree above which an anchor is hub-damped.")
    expansion_hops: int = Field(1, description="Structural hops to expand from an anchor.")

WikiPhaseTimeoutsConfig

Bases: BaseModel

How long each long-running indexing phase may run before it is called wedged.

These are WEDGE DETECTORS, not pacing knobs. Every phase below runs off the agent's event loop, so raising a value never makes indexing feel slower and lowering one never makes it faster — the only thing a timeout decides is how long a stuck phase is allowed to look like a working one. Each default is sized from that phase's own honest worst case, which depends on the repositories and the hardware a deployment actually has; that is why they are knobs rather than constants. When a phase does expire, its work is NOT stopped (a thread mid-parse or mid-clone reaches no cancellation point) — the tool reports the phase as wedged and the work continues in the background.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3206
3207
3208
3209
3210
3211
3212
3213
3214
3215
3216
3217
3218
3219
3220
3221
3222
3223
3224
3225
3226
3227
3228
3229
3230
3231
3232
3233
3234
3235
3236
3237
3238
3239
3240
3241
3242
3243
3244
3245
3246
3247
3248
3249
3250
3251
3252
3253
3254
3255
3256
3257
3258
3259
3260
class WikiPhaseTimeoutsConfig(BaseModel):
    """How long each long-running indexing phase may run before it is called wedged.

    These are WEDGE DETECTORS, not pacing knobs. Every phase below runs off the
    agent's event loop, so raising a value never makes indexing feel slower and
    lowering one never makes it faster — the only thing a timeout decides is how
    long a stuck phase is allowed to look like a working one. Each default is
    sized from that phase's own honest worst case, which depends on the
    repositories and the hardware a deployment actually has; that is why they are
    knobs rather than constants. When a phase does expire, its work is NOT
    stopped (a thread mid-parse or mid-clone reaches no cancellation point) — the
    tool reports the phase as wedged and the work continues in the background.
    """

    model_config = ConfigDict(
        extra="forbid", json_schema_extra={"title": "Phase timeouts", "x-advanced": True}
    )

    clone_s: float = Field(
        1800.0,
        gt=0,
        description=(
            "Seconds `wiki_clone_repo` waits for the checkout. Sized from the "
            "credential chain's own worst case rather than from a clone's "
            "duration: the chain tries up to five candidates and each git "
            "attempt is capped at 300s, so a repository whose every stored "
            "credential has been revoked legitimately spends 1500s before it "
            "reaches the anonymous attempt that succeeds. Raise it for very "
            "large repositories on slow links."
        ),
        examples=[1800.0, 3600.0],
    )
    scan_s: float = Field(
        900.0,
        gt=0,
        description=(
            "Seconds `wiki_scan_tree` waits for the tree walk. The scan reads "
            "and hashes every file it keeps, so its cost tracks total bytes on "
            "disk rather than file count, and it touches no network. Raise it "
            "for a very large monorepo or a slow filesystem."
        ),
        examples=[900.0, 1800.0],
    )
    graph_build_s: float = Field(
        3600.0,
        gt=0,
        description=(
            "Seconds `wiki_build_graph` waits for the tree-sitter parse, the "
            "graph write and the embedding pass. The longest of the phases and "
            "the one with the widest spread: it scales with parseable source "
            "size, and its embedding leg waits on the LLM proxy. Raise it for a "
            "large repository or a slow embedding model."
        ),
        examples=[3600.0, 7200.0],
    )

WikiRefreshConfig

Bases: BaseModel

Thresholds for the on-demand incremental refresh.

Source code in packages/mewbo_core/src/mewbo_core/config.py
3189
3190
3191
3192
3193
3194
3195
3196
3197
3198
3199
3200
3201
3202
3203
class WikiRefreshConfig(BaseModel):
    """Thresholds for the on-demand incremental refresh."""

    model_config = ConfigDict(extra="forbid", json_schema_extra={"title": "Refresh"})

    default_mode: Literal["auto", "full", "incremental"] = Field(
        "auto", description="Default re-index strategy when none is requested."
    )
    closure_max_depth: int = Field(4, description="Reverse-dependency closure depth cap.")
    drift_keep: float = Field(0.90, description="Cosine ≥ this keeps a memory anchor (no LLM).")
    drift_invalidate: float = Field(0.75, description="Cosine < this invalidates a memory anchor.")
    page_keep: float = Field(0.05, description="Doc staleness < this → keep.")
    page_edit: float = Field(0.35, description="Doc staleness < this → edit.")
    page_regen: float = Field(0.70, description="Doc staleness ≥ this → regenerate + review.")
    new_page_min: int = Field(5, description="Uncovered public symbols to propose a new page.")

effective_fallback_models() -> list[str]

Resolve the active fallback model chain honoring the opt-in policy.

Precedence: when llm.fallback.enabled is set, use llm.fallback.models (falling back to the flat llm.fallback_models if the typed list is empty). When fallback is disabled, a non-empty llm.fallback_models is still honored; otherwise there is no fallback.

A thin accessor over :meth:LLMConfig.effective_fallback_models, which owns the rule — the config-load-time budget check has to apply the SAME precedence to an AppConfig that is not the process-wide one yet.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4425
4426
4427
4428
4429
4430
4431
4432
4433
4434
4435
4436
4437
def effective_fallback_models() -> list[str]:
    """Resolve the active fallback model chain honoring the opt-in policy.

    Precedence: when ``llm.fallback.enabled`` is set, use ``llm.fallback.models``
    (falling back to the flat ``llm.fallback_models`` if the typed list is
    empty). When fallback is disabled, a non-empty ``llm.fallback_models`` is
    still honored; otherwise there is no fallback.

    A thin accessor over :meth:`LLMConfig.effective_fallback_models`, which owns
    the rule — the config-load-time budget check has to apply the SAME
    precedence to an ``AppConfig`` that is not the process-wide one yet.
    """
    return get_config().llm.effective_fallback_models()

ensure_app_config(path: str | Path) -> None

Write the default config file if missing.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4450
4451
4452
4453
4454
4455
def ensure_app_config(path: str | Path) -> None:
    """Write the default config file if missing."""
    target = Path(path)
    if target.exists():
        return
    AppConfig().write(target)

ensure_example_configs(app_path: str | Path | None = None, mcp_path: str | Path | None = None) -> tuple[Path, Path]

Write example config files if missing. Returns (app_path, mcp_path).

Source code in packages/mewbo_core/src/mewbo_core/config.py
4495
4496
4497
4498
4499
4500
4501
4502
4503
4504
4505
4506
4507
4508
4509
4510
4511
4512
4513
4514
4515
4516
4517
4518
4519
4520
4521
4522
4523
4524
4525
4526
4527
def ensure_example_configs(
    app_path: str | Path | None = None,
    mcp_path: str | Path | None = None,
) -> tuple[Path, Path]:
    """Write example config files if missing. Returns ``(app_path, mcp_path)``."""
    app_target = Path(app_path) if app_path else _default_example_path("app.example.json")
    if not app_target.exists():
        app_target.parent.mkdir(parents=True, exist_ok=True)
        app_target.write_text(
            json.dumps(_example_app_payload(), indent=2) + "\n",
            encoding="utf-8",
        )
    mcp_target = Path(mcp_path) if mcp_path else _default_example_path("mcp.example.json")
    if not mcp_target.exists():
        mcp_target.parent.mkdir(parents=True, exist_ok=True)
        mcp_target.write_text(
            json.dumps(
                {
                    "$schema": _MCP_SCHEMA_URL,
                    "servers": {
                        "codex_tools": {
                            "transport": "streamable_http",
                            "url": "http://127.0.0.1:6783/mcp/Codex-Tools-Personal",
                            "headers": {"Authorization": "Bearer YOUR_MCP_TOKEN"},
                        }
                    },
                },
                indent=2,
            )
            + "\n",
            encoding="utf-8",
        )
    return app_target, mcp_target

get_app_config_path() -> str

Return the configured app JSON path.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4373
4374
4375
4376
4377
def get_app_config_path() -> str:
    """Return the configured app JSON path."""
    if _APP_CONFIG_PATH_OVERRIDE:
        return str(_APP_CONFIG_PATH_OVERRIDE)
    return str(_resolve_config_path("app.json"))

get_config() -> AppConfig

Return cached AppConfig instance.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4389
4390
4391
4392
4393
4394
4395
4396
4397
4398
4399
4400
4401
4402
4403
4404
4405
4406
4407
def get_config() -> AppConfig:
    """Return cached AppConfig instance."""
    global _CONFIG_CACHE, _CONFIG_WARNED
    if _CONFIG_CACHE is not None:
        return _CONFIG_CACHE
    config_path = Path(get_app_config_path())
    if not config_path.exists() and not _CONFIG_WARNED:
        _logger.warning(
            "Config file not found at %s. Run /config init to scaffold examples.",
            config_path,
        )
        _CONFIG_WARNED = True
    base_payload = AppConfig().model_dump()
    file_payload = _load_json(get_app_config_path())
    merged = _deep_merge(base_payload, file_payload)
    if _APP_CONFIG_OVERRIDE:
        merged = _deep_merge(merged, _APP_CONFIG_OVERRIDE)
    _CONFIG_CACHE = AppConfig.model_validate(merged)
    return _CONFIG_CACHE

get_config_section(*keys: str) -> dict[str, Any]

Return a config section as a dictionary.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4440
4441
4442
4443
4444
4445
4446
4447
def get_config_section(*keys: str) -> dict[str, Any]:
    """Return a config section as a dictionary."""
    value = get_config_value(*keys, default={})
    if isinstance(value, BaseModel):
        return value.model_dump()
    if isinstance(value, dict):
        return value
    return {}

get_config_value(*keys: str, default: Any | None = None) -> Any

Return a nested config value or default.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4410
4411
4412
4413
4414
4415
4416
4417
4418
4419
4420
4421
4422
def get_config_value(*keys: str, default: Any | None = None) -> Any:
    """Return a nested config value or default."""
    current: Any = get_config()
    for key in keys:
        if isinstance(current, BaseModel):
            current = getattr(current, key, None)
        elif isinstance(current, dict):
            current = current.get(key)
        else:
            return default
        if current is None:
            return default
    return current

get_last_preflight() -> dict[str, dict[str, Any]] | None

Return the most recent preflight results if available.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4287
4288
4289
def get_last_preflight() -> dict[str, dict[str, Any]] | None:
    """Return the most recent preflight results if available."""
    return _LAST_PREFLIGHT

get_mcp_config_path() -> str

Return the configured MCP JSON path.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4380
4381
4382
4383
4384
4385
4386
def get_mcp_config_path() -> str:
    """Return the configured MCP JSON path."""
    if _MCP_CONFIG_DISABLED:
        return ""
    if _MCP_CONFIG_PATH_OVERRIDE:
        return str(_MCP_CONFIG_PATH_OVERRIDE)
    return str(_resolve_config_path("mcp.json"))

get_merged_mcp_config(cwd: str | None = None, *, extra_servers: dict[str, Any] | None = None, trust_cwd: bool = True) -> dict[str, Any]

Load and merge MCP configs: plugin extras + global + subtree + CWD .mcp.json.

Priority (lowest → highest): extra_servers < global < subtree (deep→shallow) < CWD. Returns the merged config dict with a servers key. When MCP is disabled (via set_mcp_config_path(None)), returns {}.

A server entry is a command this process will spawn, and the spawn happens during config resolution rather than at invocation — so no downstream tool allowlist can gate it. The two directory-derived tiers (cwd itself and the subtree walk beneath it) are therefore admitted only when the directory is trusted: pass trust_cwd=False, or register the directory via :data:register_untrusted_cwd, and BOTH are skipped entirely rather than merged at a lower priority — a partial merge would still let a repo-supplied env ride inside an operator's server. trust_cwd defaults to True as a COMPATIBILITY AFFORDANCE, not because trusting is the safe answer. The default exists so a developer's own project keeps contributing its .mcp.json, which is a real feature. It is not a judgement that an unspecified directory is safe, and it is why a caller who simply forgets this parameter gets the permissive behaviour.

Any cwd derived from request input MUST pass trust_cwd=False explicitly — a path off a query string, a repository checkout, a PR worktree, an indexing clone. Registering the directory via :data:register_untrusted_cwd covers directories this deployment creates, but it can never cover a path a CALLER names: such a path is by construction absent from that registry, and those are exactly the ones that need denying. There is no fallback behind this argument.

Cost: O(one directory tree) — the subtree walk is depth-capped, and is skipped altogether for an untrusted directory.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4694
4695
4696
4697
4698
4699
4700
4701
4702
4703
4704
4705
4706
4707
4708
4709
4710
4711
4712
4713
4714
4715
4716
4717
4718
4719
4720
4721
4722
4723
4724
4725
4726
4727
4728
4729
4730
4731
4732
4733
4734
4735
4736
4737
4738
4739
4740
4741
4742
4743
4744
4745
4746
4747
4748
4749
4750
4751
4752
4753
4754
4755
4756
4757
4758
4759
4760
4761
4762
4763
4764
4765
4766
4767
4768
4769
4770
4771
4772
4773
4774
def get_merged_mcp_config(
    cwd: str | None = None,
    *,
    extra_servers: dict[str, Any] | None = None,
    trust_cwd: bool = True,
) -> dict[str, Any]:
    """Load and merge MCP configs: plugin extras + global + subtree + CWD ``.mcp.json``.

    Priority (lowest → highest): extra_servers < global < subtree (deep→shallow) < CWD.
    Returns the merged config dict with a ``servers`` key.
    When MCP is disabled (via ``set_mcp_config_path(None)``), returns ``{}``.

    A server entry is a ``command`` this process will spawn, and the spawn
    happens during config resolution rather than at invocation — so no
    downstream tool allowlist can gate it. The two directory-derived tiers
    (``cwd`` itself and the subtree walk beneath it) are therefore admitted
    only when the directory is trusted: pass ``trust_cwd=False``, or register
    the directory via :data:`register_untrusted_cwd`, and BOTH are skipped
    entirely rather than merged at a lower priority — a partial merge would
    still let a repo-supplied ``env`` ride inside an operator's server.
    **``trust_cwd`` defaults to ``True`` as a COMPATIBILITY AFFORDANCE, not
    because trusting is the safe answer.** The default exists so a developer's
    own project keeps contributing its ``.mcp.json``, which is a real feature.
    It is not a judgement that an unspecified directory is safe, and it is why
    a caller who simply forgets this parameter gets the permissive behaviour.

    **Any ``cwd`` derived from request input MUST pass ``trust_cwd=False``
    explicitly** — a path off a query string, a repository checkout, a PR
    worktree, an indexing clone. Registering the directory via
    :data:`register_untrusted_cwd` covers directories this deployment creates,
    but it can never cover a path a CALLER names: such a path is by
    construction absent from that registry, and those are exactly the ones
    that need denying. There is no fallback behind this argument.

    Cost: ``O(one directory tree)`` — the subtree walk is depth-capped, and is
    skipped altogether for an untrusted directory.
    """
    if _MCP_CONFIG_DISABLED:
        return {}

    admit_directory_tiers = trust_cwd and not _UNTRUSTED_CWDS.contains(cwd)

    # 0. Plugin MCP servers (lowest priority — user/project configs override)
    merged: dict[str, Any] = {}
    if extra_servers:
        merged = {"servers": dict(extra_servers)}

    # 1. Load global config
    global_config: dict[str, Any] = {}
    global_path = get_mcp_config_path()
    if global_path and Path(global_path).is_file():
        try:
            with open(global_path, encoding="utf-8") as fh:
                global_config = json.load(fh)
        except (OSError, json.JSONDecodeError) as exc:
            _logger.warning("Failed to read global MCP config %s: %s", global_path, exc)

    if not isinstance(global_config, dict):
        global_config = {}

    # 2-3. Discover subtree (deepest first) then CWD .mcp.json — both skipped
    # wholesale for an untrusted directory, per the trust contract above.
    subtree_configs: list[dict[str, Any]] = []
    cwd_config: dict[str, Any] | None = None
    if admit_directory_tiers:
        subtree_configs = _discover_subtree_mcp_json(cwd)
        cwd_config = _discover_cwd_mcp_json(cwd)
    else:
        _logger.debug(
            "Untrusted working directory %s: its .mcp.json and any beneath it are excluded",
            cwd or Path.cwd(),
        )

    # 4. Merge: extra_servers ← global ← subtree (deep→shallow) ← CWD
    merged = _deep_merge(merged, global_config)
    for sub_cfg in subtree_configs:
        merged = _deep_merge(merged, sub_cfg)
    if cwd_config:
        merged = _deep_merge(merged, cwd_config)

    return merged

get_version() -> str

Return the package version from pyproject.toml (via importlib.metadata).

Tries mewbo-core first (always installed), then the workspace package (only available in local dev with uv sync).

Source code in packages/mewbo_core/src/mewbo_core/config.py
148
149
150
151
152
153
154
155
156
157
158
159
def get_version() -> str:
    """Return the package version from pyproject.toml (via importlib.metadata).

    Tries ``mewbo-core`` first (always installed), then the workspace
    package (only available in local dev with ``uv sync``).
    """
    for name in _PACKAGE_NAMES:
        try:
            return _pkg_version(name)
        except PackageNotFoundError:
            continue
    return "0.0.0"

reset_config() -> None

Clear cached configuration and overrides.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4351
4352
4353
4354
4355
4356
4357
4358
4359
4360
def reset_config() -> None:
    """Clear cached configuration and overrides."""
    global _CONFIG_CACHE, _APP_CONFIG_OVERRIDE, _APP_CONFIG_PATH_OVERRIDE, _MCP_CONFIG_PATH_OVERRIDE
    global _MCP_CONFIG_DISABLED, _CONFIG_WARNED
    _CONFIG_CACHE = None
    _APP_CONFIG_OVERRIDE = {}
    _APP_CONFIG_PATH_OVERRIDE = None
    _MCP_CONFIG_PATH_OVERRIDE = None
    _MCP_CONFIG_DISABLED = False
    _CONFIG_WARNED = False

resolve_mewbo_home() -> Path

Return the Mewbo home directory ($MEWBO_HOME or ~/.mewbo).

Source code in packages/mewbo_core/src/mewbo_core/config.py
79
80
81
82
83
84
def resolve_mewbo_home() -> Path:
    """Return the Mewbo home directory (``$MEWBO_HOME`` or ``~/.mewbo``)."""
    env = os.environ.get("MEWBO_HOME")
    if env:
        return Path(env).expanduser().resolve()
    return Path.home() / ".mewbo"

set_app_config_path(path: str | Path) -> None

Override the app config path (tests only).

Source code in packages/mewbo_core/src/mewbo_core/config.py
4333
4334
4335
4336
4337
def set_app_config_path(path: str | Path) -> None:
    """Override the app config path (tests only)."""
    global _APP_CONFIG_PATH_OVERRIDE, _CONFIG_CACHE
    _APP_CONFIG_PATH_OVERRIDE = Path(path)
    _CONFIG_CACHE = None

set_config_override(payload: dict[str, Any], *, replace: bool = False) -> None

Override config values in-memory (tests/CLI).

Source code in packages/mewbo_core/src/mewbo_core/config.py
4363
4364
4365
4366
4367
4368
4369
4370
def set_config_override(payload: dict[str, Any], *, replace: bool = False) -> None:
    """Override config values in-memory (tests/CLI)."""
    global _APP_CONFIG_OVERRIDE, _CONFIG_CACHE
    if replace:
        _APP_CONFIG_OVERRIDE = payload
    else:
        _APP_CONFIG_OVERRIDE = _deep_merge(_APP_CONFIG_OVERRIDE, payload)
    _CONFIG_CACHE = None

set_mcp_config_path(path: str | Path | None) -> None

Override the MCP config path (tests only).

Source code in packages/mewbo_core/src/mewbo_core/config.py
4340
4341
4342
4343
4344
4345
4346
4347
4348
def set_mcp_config_path(path: str | Path | None) -> None:
    """Override the MCP config path (tests only)."""
    global _MCP_CONFIG_PATH_OVERRIDE, _MCP_CONFIG_DISABLED
    if path is None or str(path).strip() == "":
        _MCP_CONFIG_PATH_OVERRIDE = None
        _MCP_CONFIG_DISABLED = True
        return
    _MCP_CONFIG_DISABLED = False
    _MCP_CONFIG_PATH_OVERRIDE = Path(path)

start_preflight(config: AppConfig | None = None, *, disable_on_failure: bool = True, on_complete: Callable[[dict[str, dict[str, Any]]], None] | None = None) -> threading.Thread

Run config preflight checks in a background thread.

Source code in packages/mewbo_core/src/mewbo_core/config.py
4258
4259
4260
4261
4262
4263
4264
4265
4266
4267
4268
4269
4270
4271
4272
4273
4274
4275
4276
4277
4278
4279
4280
4281
4282
4283
4284
def start_preflight(
    config: AppConfig | None = None,
    *,
    disable_on_failure: bool = True,
    on_complete: Callable[[dict[str, dict[str, Any]]], None] | None = None,
) -> threading.Thread:
    """Run config preflight checks in a background thread."""
    target = config or get_config()

    def _runner() -> None:
        global _LAST_PREFLIGHT
        results = asyncio.run(target.preflight(disable_on_failure=disable_on_failure))
        _LAST_PREFLIGHT = results
        failures = {
            name: info
            for name, info in results.items()
            if info.get("enabled") and not info.get("ok")
        }
        for name, info in failures.items():
            reason = info.get("reason") or "unknown failure"
            _logger.warning("Preflight check failed for %s: %s", name, reason)
        if on_complete is not None:
            on_complete(results)

    thread = threading.Thread(target=_runner, daemon=True)
    thread.start()
    return thread

mewbo_core.components

Helpers for optional components and observability integration.

ComponentStatus dataclass

Describe whether a component is enabled and why.

Source code in packages/mewbo_core/src/mewbo_core/components.py
68
69
70
71
72
73
74
75
@dataclass(frozen=True)
class ComponentStatus:
    """Describe whether a component is enabled and why."""

    name: str
    enabled: bool
    reason: str | None = None
    metadata: dict[str, JsonValue] = field(default_factory=dict)

The trace — and the observation inside it — a new span must attach to.

trace_context is the ONLY channel through which a span can name a parent observation, and naming one is not optional: given a trace_id with no parent_span_id the SDK mints a RANDOM 16-hex span id, wraps it in a NonRecordingSpan and parents the new span to it, so the exported parentObservationId points at an observation that is never sent. Every span opened that way is an orphan by construction, at every depth — which is why a span tree built from those spans cannot be walked root→child and per-agent token attribution from a trace alone is impossible.

The cure is to pass trace_context only where it buys something:

  • Inside one task, ambient OTel context already carries the enclosing span, so a nested span needs no trace_context at all — omitting it is what makes it a real child rather than a phantom-parented root.
  • Across an asyncio.create_task boundary, the spawning span is no longer the ambient one by the time the child opens its first span. There the link must be captured on the spawning side and handed over, which is the one case where an explicit parent_span_id is the right answer.
  • With nothing ambient at all (the first span of an invocation) the trace id still has to be pinned, so the phantom parent is unavoidable — it lands once, on a genuine trace root, instead of on every span.

Every read of live tracer state is best-effort: tracing is a side effect at the edge, so a link that cannot be captured degrades to None and the span falls back to the behaviour it had before rather than failing a run.

Source code in packages/mewbo_core/src/mewbo_core/components.py
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
@dataclass(frozen=True, slots=True)
class LangfuseTraceLink:
    """The trace — and the observation inside it — a new span must attach to.

    ``trace_context`` is the ONLY channel through which a span can name a
    parent observation, and naming one is not optional: given a ``trace_id``
    with no ``parent_span_id`` the SDK mints a RANDOM 16-hex span id, wraps it
    in a ``NonRecordingSpan`` and parents the new span to it, so the exported
    ``parentObservationId`` points at an observation that is never sent. Every
    span opened that way is an orphan by construction, at every depth — which
    is why a span tree built from those spans cannot be walked root→child and
    per-agent token attribution from a trace alone is impossible.

    The cure is to pass ``trace_context`` only where it buys something:

    - **Inside one task**, ambient OTel context already carries the enclosing
      span, so a nested span needs no ``trace_context`` at all — omitting it is
      what makes it a real child rather than a phantom-parented root.
    - **Across an ``asyncio.create_task`` boundary**, the spawning span is no
      longer the ambient one by the time the child opens its first span. There
      the link must be captured on the spawning side and handed over, which is
      the one case where an explicit ``parent_span_id`` is the right answer.
    - **With nothing ambient at all** (the first span of an invocation) the
      trace id still has to be pinned, so the phantom parent is unavoidable —
      it lands once, on a genuine trace root, instead of on every span.

    Every read of live tracer state is best-effort: tracing is a side effect at
    the edge, so a link that cannot be captured degrades to ``None`` and the
    span falls back to the behaviour it had before rather than failing a run.
    """

    trace_id: str
    parent_span_id: str | None = None

    @classmethod
    def from_trace_context(cls, trace_context: TraceContext | None) -> LangfuseTraceLink | None:
        """Read a link out of the SDK's ``TraceContext`` mapping, if it holds one."""
        if not trace_context:
            return None
        trace_id = trace_context.get("trace_id")
        if not trace_id:
            return None
        return cls(trace_id=trace_id, parent_span_id=trace_context.get("parent_span_id"))

    @classmethod
    def current(cls) -> LangfuseTraceLink | None:
        """The link bound to this context by :func:`langfuse_session_context`."""
        return cls.from_trace_context(_LANGFUSE_TRACE_CONTEXT.get())

    @classmethod
    def capture(cls) -> LangfuseTraceLink | None:
        """Capture the span a task started from HERE, before the task starts.

        Call this on the spawning side of an ``asyncio.create_task`` boundary,
        while the spawning span is still ambient. The live ambient observation
        wins; a context with no ambient span falls back to whatever link is
        already bound, so capturing twice down one spawn path is idempotent
        rather than parent-erasing.
        """
        bound = cls.current()
        trace_id, span_id = cls._ambient_ids()
        if trace_id and span_id:
            return cls(trace_id=trace_id, parent_span_id=span_id)
        return bound

    @staticmethod
    def _ambient_ids() -> tuple[str | None, str | None]:
        """``(trace_id, observation_id)`` of the live ambient span, best-effort."""
        try:
            from langfuse import get_client

            client = get_client()
            trace_id = client.get_current_trace_id()
            span_id = client.get_current_observation_id()
        except Exception:  # pragma: no cover - defensive
            return None, None
        if not isinstance(trace_id, str) or not isinstance(span_id, str):
            return None, None
        return trace_id, span_id

    def as_trace_context(self) -> TraceContext:
        """Render this link as the mapping the Langfuse SDK accepts."""
        ctx: dict[str, str] = {"trace_id": self.trace_id}
        if self.parent_span_id:
            ctx["parent_span_id"] = self.parent_span_id
        return cast(TraceContext, ctx)

    def trace_context_for_new_span(self) -> TraceContext | None:
        """What a span opening under this link should pass as ``trace_context``.

        ``None`` means "nest ambiently" — the enclosing span of this same trace
        is already current, so OTel parents the new span for free and passing a
        context would replace that real parent with a phantom one.
        """
        ambient_trace_id, ambient_span_id = self._ambient_ids()
        if ambient_span_id and ambient_trace_id == self.trace_id:
            return None
        return self.as_trace_context()

    @contextmanager
    def bind(self) -> Iterator[None]:
        """Bind this link for the duration of the block.

        ``asyncio.create_task`` copies the calling context, so a task created
        inside this block carries the link and opens its first span as a real
        child of the captured observation.
        """
        token = _LANGFUSE_TRACE_CONTEXT.set(self.as_trace_context())
        try:
            yield
        finally:
            _LANGFUSE_TRACE_CONTEXT.reset(token)

as_trace_context() -> TraceContext

Render this link as the mapping the Langfuse SDK accepts.

Source code in packages/mewbo_core/src/mewbo_core/components.py
302
303
304
305
306
307
def as_trace_context(self) -> TraceContext:
    """Render this link as the mapping the Langfuse SDK accepts."""
    ctx: dict[str, str] = {"trace_id": self.trace_id}
    if self.parent_span_id:
        ctx["parent_span_id"] = self.parent_span_id
    return cast(TraceContext, ctx)

bind() -> Iterator[None]

Bind this link for the duration of the block.

asyncio.create_task copies the calling context, so a task created inside this block carries the link and opens its first span as a real child of the captured observation.

Source code in packages/mewbo_core/src/mewbo_core/components.py
321
322
323
324
325
326
327
328
329
330
331
332
333
@contextmanager
def bind(self) -> Iterator[None]:
    """Bind this link for the duration of the block.

    ``asyncio.create_task`` copies the calling context, so a task created
    inside this block carries the link and opens its first span as a real
    child of the captured observation.
    """
    token = _LANGFUSE_TRACE_CONTEXT.set(self.as_trace_context())
    try:
        yield
    finally:
        _LANGFUSE_TRACE_CONTEXT.reset(token)

capture() -> LangfuseTraceLink | None classmethod

Capture the span a task started from HERE, before the task starts.

Call this on the spawning side of an asyncio.create_task boundary, while the spawning span is still ambient. The live ambient observation wins; a context with no ambient span falls back to whatever link is already bound, so capturing twice down one spawn path is idempotent rather than parent-erasing.

Source code in packages/mewbo_core/src/mewbo_core/components.py
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
@classmethod
def capture(cls) -> LangfuseTraceLink | None:
    """Capture the span a task started from HERE, before the task starts.

    Call this on the spawning side of an ``asyncio.create_task`` boundary,
    while the spawning span is still ambient. The live ambient observation
    wins; a context with no ambient span falls back to whatever link is
    already bound, so capturing twice down one spawn path is idempotent
    rather than parent-erasing.
    """
    bound = cls.current()
    trace_id, span_id = cls._ambient_ids()
    if trace_id and span_id:
        return cls(trace_id=trace_id, parent_span_id=span_id)
    return bound

current() -> LangfuseTraceLink | None classmethod

The link bound to this context by :func:langfuse_session_context.

Source code in packages/mewbo_core/src/mewbo_core/components.py
266
267
268
269
@classmethod
def current(cls) -> LangfuseTraceLink | None:
    """The link bound to this context by :func:`langfuse_session_context`."""
    return cls.from_trace_context(_LANGFUSE_TRACE_CONTEXT.get())

from_trace_context(trace_context: TraceContext | None) -> LangfuseTraceLink | None classmethod

Read a link out of the SDK's TraceContext mapping, if it holds one.

Source code in packages/mewbo_core/src/mewbo_core/components.py
256
257
258
259
260
261
262
263
264
@classmethod
def from_trace_context(cls, trace_context: TraceContext | None) -> LangfuseTraceLink | None:
    """Read a link out of the SDK's ``TraceContext`` mapping, if it holds one."""
    if not trace_context:
        return None
    trace_id = trace_context.get("trace_id")
    if not trace_id:
        return None
    return cls(trace_id=trace_id, parent_span_id=trace_context.get("parent_span_id"))

trace_context_for_new_span() -> TraceContext | None

What a span opening under this link should pass as trace_context.

None means "nest ambiently" — the enclosing span of this same trace is already current, so OTel parents the new span for free and passing a context would replace that real parent with a phantom one.

Source code in packages/mewbo_core/src/mewbo_core/components.py
309
310
311
312
313
314
315
316
317
318
319
def trace_context_for_new_span(self) -> TraceContext | None:
    """What a span opening under this link should pass as ``trace_context``.

    ``None`` means "nest ambiently" — the enclosing span of this same trace
    is already current, so OTel parents the new span for free and passing a
    context would replace that real parent with a phantom one.
    """
    ambient_trace_id, ambient_span_id = self._ambient_ids()
    if ambient_span_id and ambient_trace_id == self.trace_id:
        return None
    return self.as_trace_context()

build_langfuse_handler(*, user_id: str, session_id: str, trace_name: str, version: str, release: str, trace_context: TraceContext | None = None) -> LangfuseCallbackHandler | None

Create a Langfuse callback handler when configured.

Source code in packages/mewbo_core/src/mewbo_core/components.py
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
def build_langfuse_handler(
    *,
    user_id: str,
    session_id: str,
    trace_name: str,
    version: str,
    release: str,
    trace_context: TraceContext | None = None,
) -> LangfuseCallbackHandler | None:
    """Create a Langfuse callback handler when configured."""
    status = resolve_langfuse_status()
    if not status.enabled:
        logging.debug("Langfuse disabled: {}", status.reason)
        return None

    config = get_config().langfuse
    _ensure_langfuse_client(config)

    from langfuse.langchain import CallbackHandler

    trace_context = trace_context or _LANGFUSE_TRACE_CONTEXT.get()
    session_id_value = _LANGFUSE_SESSION_ID.get() or session_id
    user_id_value = _LANGFUSE_USER_ID.get() or user_id

    try:
        handler = CallbackHandler(public_key=config.public_key or None, trace_context=trace_context)
        _attach_langfuse_metadata(
            handler,
            user_id=user_id_value,
            session_id=session_id_value,
            trace_name=trace_name,
            version=version,
            release=release,
        )
        return handler
    except Exception as exc:  # pragma: no cover - defensive
        logging.warning("Langfuse initialization failed: {}", exc)
        return None

format_component_status(statuses: Iterable[ComponentStatus]) -> str

Format component statuses for inclusion in prompts.

Source code in packages/mewbo_core/src/mewbo_core/components.py
175
176
177
178
179
180
181
182
def format_component_status(statuses: Iterable[ComponentStatus]) -> str:
    """Format component statuses for inclusion in prompts."""
    lines: list[str] = []
    for status in statuses:
        state = "enabled" if status.enabled else "disabled"
        reason = f" ({status.reason})" if status.reason else ""
        lines.append(f"- {status.name}: {state}{reason}")
    return "\n".join(lines)

Hand a parent span to tasks created inside this block.

The one seam for the asyncio.create_task boundary. It takes both shapes the boundary comes in: pass a link captured earlier when the task is launched somewhere other than where it was spawned (a deferred unit), or omit it to capture the ambient span right here. Either way it degrades to a plain no-op when there is nothing to hand over — Langfuse disabled, or no trace bound to this context.

Source code in packages/mewbo_core/src/mewbo_core/components.py
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
@contextmanager
def langfuse_child_task_link(link: LangfuseTraceLink | None = None) -> Iterator[None]:
    """Hand a parent span to tasks created inside this block.

    The one seam for the ``asyncio.create_task`` boundary. It takes both shapes
    the boundary comes in: pass a *link* captured earlier when the task is
    launched somewhere other than where it was spawned (a deferred unit), or
    omit it to capture the ambient span right here. Either way it degrades to a
    plain no-op when there is nothing to hand over — Langfuse disabled, or no
    trace bound to this context.
    """
    resolved = link if link is not None else LangfuseTraceLink.capture()
    if resolved is None:
        yield
        return
    with resolved.bind():
        yield

langfuse_invoke_config(*, user_id: str, session_id: str, trace_name: str) -> dict[str, Any]

Build the LangChain invoke/astream config that exports a trace.

The "langfuse_metadata 3-line pattern" (see tool_use_loop), extracted so the no-loop synthesis primitives (StructuredSynthesizer / DraftStreamer) export to Langfuse too. langfuse_session_context only propagates attributes — it creates no observation, so a model call that doesn't ATTACH the CallbackHandler produces zero exported spans (a known defect: realtime synthesis traced nothing, ever). Attaching the handler is the seam that makes the generation land in the session-grouped trace.

Returns a config dict carrying callbacks (+ metadata when the handler exposes langfuse_metadata), or {} when Langfuse is disabled / unavailable — pass it straight to model.ainvoke(messages, config=cfg or None). Cheap to call (no network, no flush) so it is safe on a latency-critical path; the handler batches/exports asynchronously.

version / release are read from config here so callers stay 3 args.

Source code in packages/mewbo_core/src/mewbo_core/components.py
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
def langfuse_invoke_config(
    *,
    user_id: str,
    session_id: str,
    trace_name: str,
) -> dict[str, Any]:
    """Build the LangChain ``invoke``/``astream`` ``config`` that exports a trace.

    The "langfuse_metadata 3-line pattern" (see ``tool_use_loop``), extracted so
    the no-loop synthesis primitives (``StructuredSynthesizer`` / ``DraftStreamer``)
    export to Langfuse too. ``langfuse_session_context`` only *propagates
    attributes* — it creates no observation, so a model call that doesn't ATTACH
    the ``CallbackHandler`` produces zero exported spans (a known defect: realtime
    synthesis traced nothing, ever). Attaching the handler is the seam that makes
    the generation land in the session-grouped trace.

    Returns a ``config`` dict carrying ``callbacks`` (+ ``metadata`` when the
    handler exposes ``langfuse_metadata``), or ``{}`` when Langfuse is disabled /
    unavailable — pass it straight to ``model.ainvoke(messages, config=cfg or None)``.
    Cheap to call (no network, no flush) so it is safe on a latency-critical path;
    the handler batches/exports asynchronously.

    ``version`` / ``release`` are read from config here so callers stay 3 args.
    """
    handler = build_langfuse_handler(
        user_id=user_id,
        session_id=session_id,
        trace_name=trace_name,
        version=get_version(),
        release=get_config_value("runtime", "envmode", default="Not Specified"),
    )
    if handler is None:
        return {}
    config: dict[str, Any] = {"callbacks": [handler]}
    metadata = getattr(handler, "langfuse_metadata", None)
    if isinstance(metadata, dict) and metadata:
        config["metadata"] = metadata
    return config

langfuse_propagate(*, tags: list[str] | None = None, metadata: dict[str, str] | None = None, session_id: str | None = None, user_id: str | None = None, trace_name: str | None = None, version: str | None = None) -> Iterator[None]

Propagate Langfuse attributes to all child observations.

Thin wrapper around langfuse.propagate_attributes that gracefully degrades when Langfuse is disabled or unavailable.

trace_name is the ONE channel that names a trace from outside an observation: a trace otherwise inherits the name of whatever runnable opened its root span, which is why an untouched export is a wall of identically named traces. version rides along for the same reason — both are first-class trace fields, never tags.

Source code in packages/mewbo_core/src/mewbo_core/components.py
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
@contextmanager
def langfuse_propagate(
    *,
    tags: list[str] | None = None,
    metadata: dict[str, str] | None = None,
    session_id: str | None = None,
    user_id: str | None = None,
    trace_name: str | None = None,
    version: str | None = None,
) -> Iterator[None]:
    """Propagate Langfuse attributes to all child observations.

    Thin wrapper around ``langfuse.propagate_attributes`` that gracefully
    degrades when Langfuse is disabled or unavailable.

    *trace_name* is the ONE channel that names a trace from outside an
    observation: a trace otherwise inherits the name of whatever runnable opened
    its root span, which is why an untouched export is a wall of identically
    named traces. *version* rides along for the same reason — both are
    first-class trace fields, never tags.
    """
    status = resolve_langfuse_status()
    if not status.enabled:
        yield
        return
    try:
        from langfuse import propagate_attributes
    except Exception:  # pragma: no cover - defensive
        yield
        return
    kwargs: dict[str, Any] = {}
    if tags:
        kwargs["tags"] = tags
    if metadata:
        kwargs["metadata"] = metadata
    if session_id:
        kwargs["session_id"] = session_id
    if user_id:
        kwargs["user_id"] = user_id
    if trace_name:
        kwargs["trace_name"] = trace_name
    if version:
        kwargs["version"] = version
    if not kwargs:
        yield
        return
    try:
        with propagate_attributes(**kwargs):
            yield
    except Exception:  # pragma: no cover - defensive
        yield

langfuse_session_context(session_id: str, *, user_id: str | None = None, invocation_id: str | None = None, source_platform: str | None = None, trace_name: str | None = None, tags: list[str] | None = None, metadata: dict[str, str] | None = None) -> Iterator[None]

Bind a Langfuse trace context to the current invocation.

Each call gets a unique trace (via invocation_id) while Langfuse groups traces under the same session_id.

trace_name names that trace. Omitting it does not leave the trace unnamed: it leaves it named after whichever LangChain runnable happened to open the root span, which is a property of the client library rather than of the work.

tags / metadata carry pre-derived trace provenance (see session_provenance.TraceProvenance) and are merged into the propagated baseline so every child observation — including nested LangChain CallbackHandler generations and langfuse_trace_span spans — is filterable by product, workspace, project, repo, branch, surface, etc. This seam stays taxonomy-free: it propagates whatever it is handed. When provenance is supplied its surface:<…> tag supersedes the coarse channel:<source_platform> fallback (kept only for callers that pass source_platform alone).

Source code in packages/mewbo_core/src/mewbo_core/components.py
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
@contextmanager
def langfuse_session_context(
    session_id: str,
    *,
    user_id: str | None = None,
    invocation_id: str | None = None,
    source_platform: str | None = None,
    trace_name: str | None = None,
    tags: list[str] | None = None,
    metadata: dict[str, str] | None = None,
) -> Iterator[None]:
    """Bind a Langfuse trace context to the current invocation.

    Each call gets a **unique trace** (via *invocation_id*) while
    Langfuse groups traces under the same *session_id*.

    *trace_name* names that trace. Omitting it does not leave the trace unnamed:
    it leaves it named after whichever LangChain runnable happened to open the
    root span, which is a property of the client library rather than of the work.

    *tags* / *metadata* carry pre-derived trace provenance
    (see ``session_provenance.TraceProvenance``) and are merged into the
    propagated baseline so every child observation — including nested LangChain
    CallbackHandler generations and ``langfuse_trace_span`` spans — is filterable
    by product, workspace, project, repo, branch, surface, etc. This seam stays
    taxonomy-free: it propagates whatever it is handed. When provenance is
    supplied its ``surface:<…>`` tag supersedes the coarse
    ``channel:<source_platform>`` fallback (kept only for callers that pass
    *source_platform* alone).
    """
    trace_context = _build_langfuse_trace_context(session_id, invocation_id)
    token_ctx = _LANGFUSE_TRACE_CONTEXT.set(trace_context)
    token_session = _LANGFUSE_SESSION_ID.set(session_id)
    resolved_user = user_id or ANONYMOUS_USER_ID
    token_user = _LANGFUSE_USER_ID.set(resolved_user)

    # Use propagate_attributes so session_id, user_id, and baseline
    # tags automatically attach to every child observation (including
    # LangChain CallbackHandler generations).
    base_tags = ["mewbo"]
    base_metadata: dict[str, str] = {"sessionid": session_id[:12]}
    if tags or metadata:
        base_tags.extend(tags or [])
        base_metadata.update(metadata or {})
    elif source_platform:
        base_tags.append(f"channel:{source_platform}")
    propagate_cm = langfuse_propagate(
        session_id=session_id,
        user_id=resolved_user,
        trace_name=trace_name,
        version=get_version(),
        tags=base_tags,
        metadata=base_metadata,
    )
    propagate_cm.__enter__()
    try:
        yield
    finally:
        try:
            propagate_cm.__exit__(None, None, None)
        except Exception:  # pragma: no cover - defensive
            pass
        _LANGFUSE_TRACE_CONTEXT.reset(token_ctx)
        _LANGFUSE_SESSION_ID.reset(token_session)
        _LANGFUSE_USER_ID.reset(token_user)

langfuse_trace_span(name: str, *, as_type: ObservationKind = 'span', metadata: dict[str, str] | None = None, input_data: Any = None, level: str | None = None, attributes: dict[str, str] | None = None) -> Iterator[object | None]

Open a Langfuse observation bound to the current session trace context.

as_type selects how Langfuse renders the observation — an agent, a tool call and a plain span are the same OTel span with different types, and a trace built entirely of untyped spans loses the one axis the UI groups by. metadata is attached for filtering. input_data is set as the input. level sets the log level (e.g. "ERROR"). attributes are raw OTel span attributes, stamped as early as the SDK allows (see :func:_stamp_span_attributes).

Source code in packages/mewbo_core/src/mewbo_core/components.py
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
@contextmanager
def langfuse_trace_span(
    name: str,
    *,
    as_type: ObservationKind = "span",
    metadata: dict[str, str] | None = None,
    input_data: Any = None,
    level: str | None = None,
    attributes: dict[str, str] | None = None,
) -> Iterator[object | None]:
    """Open a Langfuse observation bound to the current session trace context.

    *as_type* selects how Langfuse renders the observation — an agent, a tool
    call and a plain span are the same OTel span with different types, and a
    trace built entirely of untyped spans loses the one axis the UI groups by.
    *metadata* is attached for filtering. *input_data* is set as the input.
    *level* sets the log level (e.g. ``"ERROR"``). *attributes* are raw OTel
    span attributes, stamped as early as the SDK allows (see
    :func:`_stamp_span_attributes`).
    """
    status = resolve_langfuse_status()
    if not status.enabled:
        yield None
        return
    link = LangfuseTraceLink.current()
    if link is None:
        yield None
        return
    try:
        from langfuse import get_client
    except Exception:  # pragma: no cover - defensive
        yield None
        return
    # Setup phase: if Langfuse fails, yield None (graceful degradation).
    # Body exceptions MUST propagate — never suppress them.
    span = None
    cm = None
    try:
        langfuse = get_client()
        # ``None`` here is the load-bearing case, not a degradation: it means an
        # enclosing span of this trace is ambient, so OTel parents this one for
        # real. Passing a context unconditionally is what parented every span to
        # a freshly-minted id that was never exported.
        cm = langfuse.start_as_current_observation(
            # ``as_type`` is overloaded per literal in the SDK's stubs, so a
            # variable of the union type has to be widened for the call to
            # resolve; the SDK validates the value itself.
            as_type=cast(Any, as_type),
            name=name,
            trace_context=link.trace_context_for_new_span(),
        )
        span = cm.__enter__()
        if span is not None:
            _stamp_span_attributes(span, attributes)
            update_kwargs: dict[str, Any] = {}
            if metadata:
                update_kwargs["metadata"] = metadata
            if input_data is not None:
                update_kwargs["input"] = input_data
            if level:
                update_kwargs["level"] = level
            if update_kwargs:
                span.update(**update_kwargs)
    except Exception:  # pragma: no cover - defensive
        logging.debug("Langfuse trace span setup failed.", exc_info=True)
    try:
        yield span
    except GeneratorExit:
        # The generator was closed rather than the traced body failing; there is
        # no operation outcome to record.
        raise
    except BaseException as exc:
        # A traced operation that raised must leave its cause ON the span.
        # Exiting the observation clean (which is all the ``finally`` below
        # does) is why a whole failure corpus carried ERROR-level spans with an
        # empty status_message and no exception event: the span recorded THAT
        # something failed and never WHAT. ``BaseException`` so a cancellation
        # or a deadline kill — the shapes a wedged call actually takes — is
        # recorded too, then re-raised untouched.
        record_span_exception(span, exc)
        raise
    finally:
        if cm is not None:
            try:
                cm.__exit__(None, None, None)
            except Exception:  # pragma: no cover - defensive
                pass

record_span_exception(span, exc=None, *, message=None, attributes=None)

Record an exception on a Langfuse span: ERROR status + an OTel event.

Two writes, because Langfuse reads them from different places: the span's own level/status_message drive the UI's error surface, while the exception tooling queries OTel exception events — span.update( level="ERROR") alone sets status with no cause attached. Setting the status here (rather than leaving it to each call site) is what makes the non-empty guarantee in :func:_span_status_message hold for every writer. Fully graceful: no span / no otel / disabled → no-op.

Source code in packages/mewbo_core/src/mewbo_core/components.py
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
def record_span_exception(span, exc=None, *, message=None, attributes=None):
    """Record an exception on a Langfuse span: ERROR status + an OTel event.

    Two writes, because Langfuse reads them from different places: the span's
    own ``level``/``status_message`` drive the UI's error surface, while the
    exception tooling queries OTel ``exception`` events — ``span.update(
    level="ERROR")`` alone sets status with no cause attached. Setting the
    status here (rather than leaving it to each call site) is what makes the
    non-empty guarantee in :func:`_span_status_message` hold for every writer.
    Fully graceful: no span / no otel / disabled → no-op.
    """
    if span is None:
        return
    detail = exc if exc is not None else (RuntimeError(message) if message else None)
    if detail is None:
        return
    try:
        span.update(level="ERROR", status_message=_span_status_message(detail))
    except Exception:  # pragma: no cover - never disrupt the run
        logging.debug("Langfuse span status update failed.", exc_info=True)
    otel = getattr(span, "_otel_span", None)
    try:
        if otel is not None and otel.is_recording():
            otel.record_exception(detail, attributes=attributes or None)
    except Exception:  # pragma: no cover - never disrupt the run
        logging.debug("Langfuse record_exception failed.", exc_info=True)

resolve_home_assistant_status() -> ComponentStatus

Determine whether the Home Assistant tool is configured.

Source code in packages/mewbo_core/src/mewbo_core/components.py
164
165
166
167
168
169
170
171
172
def resolve_home_assistant_status() -> ComponentStatus:
    """Determine whether the Home Assistant tool is configured."""
    enabled, reason, metadata = get_config().home_assistant.evaluate()
    return ComponentStatus(
        name="home_assistant_tool",
        enabled=enabled,
        reason=reason,
        metadata=metadata,
    )

resolve_langfuse_status() -> ComponentStatus

Determine whether Langfuse callbacks are available and configured.

Source code in packages/mewbo_core/src/mewbo_core/components.py
78
79
80
81
def resolve_langfuse_status() -> ComponentStatus:
    """Determine whether Langfuse callbacks are available and configured."""
    enabled, reason, metadata = get_config().langfuse.evaluate()
    return ComponentStatus(name="langfuse", enabled=enabled, reason=reason, metadata=metadata)

mewbo_core.permissions

Permission policies for tool execution.

PermissionDecision

Bases: str, Enum

Outcomes for a permission check.

Source code in packages/mewbo_core/src/mewbo_core/permissions.py
23
24
25
26
27
28
class PermissionDecision(str, Enum):
    """Outcomes for a permission check."""

    ALLOW = "allow"
    DENY = "deny"
    ASK = "ask"

PermissionPolicy

Evaluate permission rules for action steps.

Source code in packages/mewbo_core/src/mewbo_core/permissions.py
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
class PermissionPolicy:
    """Evaluate permission rules for action steps."""

    def __init__(
        self,
        rules: list[PermissionRule] | None = None,
        default_by_operation: dict[str, PermissionDecision] | None = None,
        default_decision: PermissionDecision = PermissionDecision.ASK,
    ) -> None:
        """Initialize the permission policy."""
        self._rules = rules or []
        self._default_by_operation = default_by_operation or {}
        self._default_decision = default_decision

    def decide(self, action_step: ActionStep) -> PermissionDecision:
        """Return the permission decision for an action step."""
        for rule in self._rules:
            if rule.matches(action_step):
                return rule.decision
        operation_decision = self._default_by_operation.get(action_step.operation)
        if operation_decision is not None:
            return operation_decision
        return self._default_decision

__init__(rules: list[PermissionRule] | None = None, default_by_operation: dict[str, PermissionDecision] | None = None, default_decision: PermissionDecision = PermissionDecision.ASK) -> None

Initialize the permission policy.

Source code in packages/mewbo_core/src/mewbo_core/permissions.py
49
50
51
52
53
54
55
56
57
58
def __init__(
    self,
    rules: list[PermissionRule] | None = None,
    default_by_operation: dict[str, PermissionDecision] | None = None,
    default_decision: PermissionDecision = PermissionDecision.ASK,
) -> None:
    """Initialize the permission policy."""
    self._rules = rules or []
    self._default_by_operation = default_by_operation or {}
    self._default_decision = default_decision

decide(action_step: ActionStep) -> PermissionDecision

Return the permission decision for an action step.

Source code in packages/mewbo_core/src/mewbo_core/permissions.py
60
61
62
63
64
65
66
67
68
def decide(self, action_step: ActionStep) -> PermissionDecision:
    """Return the permission decision for an action step."""
    for rule in self._rules:
        if rule.matches(action_step):
            return rule.decision
    operation_decision = self._default_by_operation.get(action_step.operation)
    if operation_decision is not None:
        return operation_decision
    return self._default_decision

PermissionRule dataclass

Rule describing a tool/action permission decision.

Source code in packages/mewbo_core/src/mewbo_core/permissions.py
31
32
33
34
35
36
37
38
39
40
41
42
43
@dataclass(frozen=True)
class PermissionRule:
    """Rule describing a tool/action permission decision."""

    tool_id: str = "*"
    operation: str = "*"
    decision: PermissionDecision = PermissionDecision.ASK

    def matches(self, action_step: ActionStep) -> bool:
        """Return True when the action step matches the rule pattern."""
        return fnmatch(action_step.tool_id, self.tool_id) and fnmatch(
            action_step.operation, self.operation
        )

matches(action_step: ActionStep) -> bool

Return True when the action step matches the rule pattern.

Source code in packages/mewbo_core/src/mewbo_core/permissions.py
39
40
41
42
43
def matches(self, action_step: ActionStep) -> bool:
    """Return True when the action step matches the rule pattern."""
    return fnmatch(action_step.tool_id, self.tool_id) and fnmatch(
        action_step.operation, self.operation
    )

approval_callback_from_config() -> Callable[[ActionStep], bool] | None

Return an approval callback based on config settings.

Source code in packages/mewbo_core/src/mewbo_core/permissions.py
149
150
151
152
153
154
155
156
157
def approval_callback_from_config() -> Callable[[ActionStep], bool] | None:
    """Return an approval callback based on config settings."""
    mode_raw = get_config_value("permissions", "approval_mode", default="")
    mode = str(mode_raw or "").strip().lower()
    if mode in {"allow", "auto", "approve", "yes"}:
        return lambda _: True
    if mode in {"deny", "never", "no"}:
        return lambda _: False
    return None

auto_approve(_: ActionStep) -> bool

Approval callback that always approves.

Source code in packages/mewbo_core/src/mewbo_core/permissions.py
160
161
162
def auto_approve(_: ActionStep) -> bool:
    """Approval callback that always approves."""
    return True

auto_deny(_: ActionStep) -> bool

Approval callback that always denies.

Source code in packages/mewbo_core/src/mewbo_core/permissions.py
165
166
167
def auto_deny(_: ActionStep) -> bool:
    """Approval callback that always denies."""
    return False

load_permission_policy(path: str | None = None) -> PermissionPolicy

Load permission policy configuration from disk or defaults.

Source code in packages/mewbo_core/src/mewbo_core/permissions.py
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
def load_permission_policy(path: str | None = None) -> PermissionPolicy:
    """Load permission policy configuration from disk or defaults."""
    if path is None:
        path = get_config_value("permissions", "policy_path")
    if not path:
        return _default_policy()
    if not os.path.exists(path):
        logging.warning("Permission policy file not found: {}", path)
        return _default_policy()
    try:
        payload = _load_policy_data(path)
    except (json.JSONDecodeError, OSError, tomllib.TOMLDecodeError) as exc:
        logging.warning("Failed to load permission policy: {}", exc)
        return _default_policy()

    rules: list[PermissionRule] = []
    for rule_data in payload.get("rules", []):
        decision = _parse_decision(rule_data.get("decision"))
        if decision is None:
            continue
        rules.append(
            PermissionRule(
                tool_id=str(rule_data.get("tool_id", "*")),
                operation=str(rule_data.get("operation", "*")),
                decision=decision,
            )
        )

    default_by_operation: dict[str, PermissionDecision] = {}
    for key, value in payload.get("default_by_operation", {}).items():
        parsed = _parse_decision(str(value))
        if parsed is not None:
            default_by_operation[str(key)] = parsed

    default_decision = _parse_decision(payload.get("default_decision"))
    if default_decision is None:
        default_decision = PermissionDecision.ASK

    return PermissionPolicy(
        rules=rules,
        default_by_operation=default_by_operation,
        default_decision=default_decision,
    )

mewbo_core.hooks

Hook manager for orchestration lifecycle events.

HookDispatch dataclass

The background threads behind every fire-and-forget hook, joinable.

A dispatch nobody can join is only observable by sleeping and hoping the thread won. That is a RACE, not a slow caller: on a loaded machine the observation is simply WRONG rather than late. Holding the handles here gives any caller that needs the outcome — a shutdown path, a test — a bounded wait to join on, while every hot-path caller keeps ignoring the return value and the fire-and-forget semantics are unchanged.

Process-wide by construction (:data:HOOK_DISPATCH) rather than per manager: the threads are the process's, a plugin-translated hook is dispatched with no manager in reach of the factory, and "has every dispatched hook landed" is the only question a joiner actually has.

Cost: O(1) per submit and O(1) per completion — a thread drops itself from the set when it ends, so nothing ever sweeps the set and the dispatch a hook pays for does not grow with how many are already in flight. Memory is bounded by what is genuinely in flight, for the same reason.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
@dataclass
class HookDispatch:
    """The background threads behind every fire-and-forget hook, joinable.

    A dispatch nobody can join is only observable by sleeping and hoping the
    thread won. That is a RACE, not a slow caller: on a loaded machine the
    observation is simply WRONG rather than late. Holding the handles here
    gives any caller that needs the outcome — a shutdown path, a test — a
    bounded ``wait`` to join on, while every hot-path caller keeps ignoring the
    return value and the fire-and-forget semantics are unchanged.

    Process-wide by construction (:data:`HOOK_DISPATCH`) rather than per
    manager: the threads are the process's, a plugin-translated hook is
    dispatched with no manager in reach of the factory, and "has every
    dispatched hook landed" is the only question a joiner actually has.

    Cost: ``O(1)`` per submit and ``O(1)`` per completion — a thread drops
    itself from the set when it ends, so nothing ever sweeps the set and the
    dispatch a hook pays for does not grow with how many are already in flight.
    Memory is bounded by what is genuinely in flight, for the same reason.
    """

    _threads: set[threading.Thread] = field(default_factory=set)
    _lock: threading.Lock = field(default_factory=threading.Lock)

    def submit(self, target: Callable[..., None], *args: Any) -> threading.Thread:
        """Run *target* on a daemon thread and return it, tracked until it ends."""
        thread = threading.Thread(target=self._run, args=(target, args), daemon=True)
        with self._lock:
            self._threads.add(thread)
        thread.start()
        return thread

    def _run(self, target: Callable[..., None], args: tuple[Any, ...]) -> None:
        """Run *target*, then drop this thread from the tracking set.

        The removal is the thread's OWN job because the alternative — sweeping
        the set for finished threads on the way into ``submit`` — is
        ``O(in-flight)`` work under a process-wide lock, so a burst of events
        makes every dispatch in the burst pay for the burst. Measured on one
        host, that sweep cost 49 µs at 200 live threads and 1.09 ms at 3000,
        against a lock-free ``Thread(...).start()`` before any of this existed.
        Discarding one member is ``O(1)``, and it happens off the hot path.

        The ``finally`` does not swallow anything: a raising hook still reaches
        ``threading.excepthook`` exactly as it did when the target ran as the
        thread's own callable.
        """
        try:
            target(*args)
        finally:
            with self._lock:
                self._threads.discard(threading.current_thread())

    def wait(self, timeout: float = 5.0) -> bool:
        """Join everything in flight within *timeout*; return whether all landed.

        Never raises and never re-raises a hook's own failure — a joiner is
        asking whether the work finished, not taking on responsibility for what
        it did. A timeout logs once and returns ``False``.

        Cost: ``O(in-flight dispatches)`` joins, bounded overall by *timeout*.
        No caller is on a hot path — production never joins; this is for a
        shutdown path or a test.
        """
        deadline = time.monotonic() + timeout
        with self._lock:
            pending = list(self._threads)
        for thread in pending:
            thread.join(max(0.0, deadline - time.monotonic()))
        still_running = [t for t in pending if t.is_alive()]
        if still_running:
            logger.warning(
                "{} hook dispatch(es) still in flight after {}s", len(still_running), timeout
            )
            return False
        return True

submit(target: Callable[..., None], *args: Any) -> threading.Thread

Run target on a daemon thread and return it, tracked until it ends.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
163
164
165
166
167
168
169
def submit(self, target: Callable[..., None], *args: Any) -> threading.Thread:
    """Run *target* on a daemon thread and return it, tracked until it ends."""
    thread = threading.Thread(target=self._run, args=(target, args), daemon=True)
    with self._lock:
        self._threads.add(thread)
    thread.start()
    return thread

wait(timeout: float = 5.0) -> bool

Join everything in flight within timeout; return whether all landed.

Never raises and never re-raises a hook's own failure — a joiner is asking whether the work finished, not taking on responsibility for what it did. A timeout logs once and returns False.

Cost: O(in-flight dispatches) joins, bounded overall by timeout. No caller is on a hot path — production never joins; this is for a shutdown path or a test.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
def wait(self, timeout: float = 5.0) -> bool:
    """Join everything in flight within *timeout*; return whether all landed.

    Never raises and never re-raises a hook's own failure — a joiner is
    asking whether the work finished, not taking on responsibility for what
    it did. A timeout logs once and returns ``False``.

    Cost: ``O(in-flight dispatches)`` joins, bounded overall by *timeout*.
    No caller is on a hot path — production never joins; this is for a
    shutdown path or a test.
    """
    deadline = time.monotonic() + timeout
    with self._lock:
        pending = list(self._threads)
    for thread in pending:
        thread.join(max(0.0, deadline - time.monotonic()))
    still_running = [t for t in pending if t.is_alive()]
    if still_running:
        logger.warning(
            "{} hook dispatch(es) still in flight after {}s", len(still_running), timeout
        )
        return False
    return True

HookManager dataclass

Container for hook callbacks used during orchestration.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
@dataclass
class HookManager:
    """Container for hook callbacks used during orchestration."""

    pre_tool_use: list[Callable[[ActionStep], ActionStep]] = field(default_factory=list)
    post_tool_use: list[Callable[[ActionStep, MockSpeaker], MockSpeaker]] = field(
        default_factory=list
    )
    permission_request: list[Callable[[ActionStep, PermissionDecision], PermissionDecision]] = (
        field(default_factory=list)
    )
    pre_compact: list[Callable[[list[EventRecord]], list[EventRecord]]] = field(
        default_factory=list
    )
    on_agent_start: list[Callable[[AgentHandle], None]] = field(default_factory=list)
    on_agent_stop: list[Callable[[AgentHandle], None]] = field(default_factory=list)
    on_session_start: list[Callable[[str], None]] = field(default_factory=list)
    # Returns an ``OutcomeAssertion`` to report an unmet purpose, or ``None``
    # to report nothing. Widened from ``None`` — every existing hook satisfies
    # the new signature unchanged, since returning nothing IS returning None.
    on_session_end: list[Callable[[str, str | None], OutcomeAssertion | None]] = field(
        default_factory=list
    )
    on_event: list[Callable[[str, EventRecord], None]] = field(default_factory=list)
    on_compact: list[Callable[..., None]] = field(default_factory=list)

    def run_pre_tool_use(self, action_step: ActionStep) -> ActionStep:
        """Apply pre-tool hooks to an action step.

        Args:
            action_step: Action step to process.

        Returns:
            Updated action step after hooks run.
        """
        for hook in self.pre_tool_use:
            try:
                action_step = hook(action_step)
            except Exception:
                logger.warning("Pre-tool hook failed", exc_info=True)
        return action_step

    def run_post_tool_use(self, action_step: ActionStep, result: MockSpeaker) -> MockSpeaker:
        """Apply post-tool hooks to a tool result.

        Args:
            action_step: Action step that was executed.
            result: Result returned by the tool.

        Returns:
            Updated result after hooks run.
        """
        for hook in self.post_tool_use:
            try:
                result = hook(action_step, result)
            except Exception:
                logger.warning("Post-tool hook failed", exc_info=True)
        return result

    def run_permission_request(
        self, action_step: ActionStep, decision: PermissionDecision
    ) -> PermissionDecision:
        """Apply permission hooks to a decision outcome.

        Args:
            action_step: Action step under review.
            decision: Current decision to modify.

        Returns:
            Updated permission decision after hooks run.
        """
        for hook in self.permission_request:
            try:
                decision = hook(action_step, decision)
            except Exception:
                logger.warning("Permission hook failed", exc_info=True)
        return decision

    def run_pre_compact(self, events: Iterable[EventRecord]) -> list[EventRecord]:
        """Apply compaction hooks to events prior to summarization.

        Args:
            events: Iterable of event records.

        Returns:
            List of event records after hooks run.
        """
        event_list: list[EventRecord] = list(events)
        for hook in self.pre_compact:
            try:
                event_list = hook(event_list)
            except Exception:
                logger.warning("Pre-compact hook failed", exc_info=True)
        return event_list

    def run_on_agent_start(self, handle: AgentHandle) -> None:
        """Notify hooks that an agent has started."""
        for hook in self.on_agent_start:
            try:
                hook(handle)
            except Exception:
                logger.warning("on_agent_start hook failed", exc_info=True)

    def run_on_agent_stop(self, handle: AgentHandle) -> None:
        """Notify hooks that an agent has stopped."""
        for hook in self.on_agent_stop:
            try:
                hook(handle)
            except Exception:
                logger.warning("on_agent_stop hook failed", exc_info=True)

    def run_on_session_start(self, session_id: str) -> None:
        """Notify hooks that a session has started."""
        for hook in self.on_session_start:
            try:
                hook(session_id)
            except Exception:
                logger.warning("on_session_start hook failed", exc_info=True)

    def run_on_session_end(
        self, session_id: str, error: str | None = None
    ) -> list[OutcomeAssertion]:
        """Notify hooks a session ended; collect any outcome assertions they report.

        A hook MAY return an :class:`OutcomeAssertion` to report that the
        session's purpose was not achieved. Returning ``None`` asserts nothing
        — see that class for why absence must stay distinguishable from success.

        The return value is COLLECTED, not discarded — discarding it is how a
        job that never reached its terminal state still presents as a clean
        completion, with the hook able to SEE the failure and nowhere to put
        it.

        Failure-isolated per hook like every other lifecycle dispatch, and
        deliberately strict about what it accepts: a hook returning something
        that is not an assertion is logged and IGNORED rather than coerced,
        because a truthy stray return would otherwise invent a failure.
        """
        assertions: list[OutcomeAssertion] = []
        for hook in self.on_session_end:
            try:
                reported = hook(session_id, error)
            except Exception:
                logger.warning("on_session_end hook failed", exc_info=True)
                continue
            if reported is None:
                continue
            if not isinstance(reported, OutcomeAssertion):
                logger.warning(
                    "on_session_end hook returned {}, not an OutcomeAssertion; ignoring",
                    type(reported).__name__,
                )
                continue
            if not reported.source:
                # Re-CONSTRUCTED rather than ``model_copy``d: the stamp is a
                # function name of unbounded length, and model_copy skips the
                # validator that bounds it.
                reported = OutcomeAssertion(
                    reason=reported.reason,
                    detail=reported.detail,
                    source=getattr(hook, "__name__", "") or type(hook).__name__,
                )
            assertions.append(reported)
        return assertions

    def run_on_event(self, session_id: str, event: EventRecord) -> None:
        """Notify hooks that an event was appended to a session transcript.

        Runs on the event-append hot path (via the ``SessionEventBus``
        observer), so every registered hook is itself fire-and-forget; this
        loop only dispatches and is failure-isolated per hook.
        """
        for hook in self.on_event:
            try:
                hook(session_id, event)
            except Exception:
                logger.warning("on_event hook failed", exc_info=True)

    def run_on_compact(
        self,
        session_id: str,
        *,
        summary: str = "",
        tokens_before: int = 0,
        tokens_saved: int = 0,
        events_summarized: int = 0,
    ) -> None:
        """Notify hooks that compaction occurred."""
        for hook in self.on_compact:
            try:
                hook(
                    session_id,
                    summary=summary,
                    tokens_before=tokens_before,
                    tokens_saved=tokens_saved,
                    events_summarized=events_summarized,
                )
            except Exception:
                logger.warning("on_compact hook failed", exc_info=True)

    @classmethod
    def load_from_config(cls, hooks_config: HooksConfig) -> HookManager:
        """Create a HookManager with hooks loaded from config."""
        manager = cls()
        pre_map = {"command": _make_command_hook, "http": _make_http_hook}
        post_map = {"command": _make_post_tool_hook, "http": _make_http_post_tool_hook}
        start_map = {"command": _make_session_hook, "http": _make_http_session_hook}
        end_map = {"command": _make_session_end_hook, "http": _make_http_session_end_hook}
        event_map = {"command": _make_event_command_hook, "http": _make_http_event_hook}
        for entry in hooks_config.pre_tool_use:
            manager.pre_tool_use.append(pre_map.get(entry.type, _make_command_hook)(entry))
        for entry in hooks_config.post_tool_use:
            manager.post_tool_use.append(post_map.get(entry.type, _make_post_tool_hook)(entry))
        for entry in hooks_config.on_session_start:
            manager.on_session_start.append(start_map.get(entry.type, _make_session_hook)(entry))
        for entry in hooks_config.on_session_end:
            manager.on_session_end.append(end_map.get(entry.type, _make_session_end_hook)(entry))
        for entry in hooks_config.on_event:
            manager.on_event.append(event_map.get(entry.type, _make_event_command_hook)(entry))
        return manager

load_from_config(hooks_config: HooksConfig) -> HookManager classmethod

Create a HookManager with hooks loaded from config.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
@classmethod
def load_from_config(cls, hooks_config: HooksConfig) -> HookManager:
    """Create a HookManager with hooks loaded from config."""
    manager = cls()
    pre_map = {"command": _make_command_hook, "http": _make_http_hook}
    post_map = {"command": _make_post_tool_hook, "http": _make_http_post_tool_hook}
    start_map = {"command": _make_session_hook, "http": _make_http_session_hook}
    end_map = {"command": _make_session_end_hook, "http": _make_http_session_end_hook}
    event_map = {"command": _make_event_command_hook, "http": _make_http_event_hook}
    for entry in hooks_config.pre_tool_use:
        manager.pre_tool_use.append(pre_map.get(entry.type, _make_command_hook)(entry))
    for entry in hooks_config.post_tool_use:
        manager.post_tool_use.append(post_map.get(entry.type, _make_post_tool_hook)(entry))
    for entry in hooks_config.on_session_start:
        manager.on_session_start.append(start_map.get(entry.type, _make_session_hook)(entry))
    for entry in hooks_config.on_session_end:
        manager.on_session_end.append(end_map.get(entry.type, _make_session_end_hook)(entry))
    for entry in hooks_config.on_event:
        manager.on_event.append(event_map.get(entry.type, _make_event_command_hook)(entry))
    return manager

run_on_agent_start(handle: AgentHandle) -> None

Notify hooks that an agent has started.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
317
318
319
320
321
322
323
def run_on_agent_start(self, handle: AgentHandle) -> None:
    """Notify hooks that an agent has started."""
    for hook in self.on_agent_start:
        try:
            hook(handle)
        except Exception:
            logger.warning("on_agent_start hook failed", exc_info=True)

run_on_agent_stop(handle: AgentHandle) -> None

Notify hooks that an agent has stopped.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
325
326
327
328
329
330
331
def run_on_agent_stop(self, handle: AgentHandle) -> None:
    """Notify hooks that an agent has stopped."""
    for hook in self.on_agent_stop:
        try:
            hook(handle)
        except Exception:
            logger.warning("on_agent_stop hook failed", exc_info=True)

run_on_compact(session_id: str, *, summary: str = '', tokens_before: int = 0, tokens_saved: int = 0, events_summarized: int = 0) -> None

Notify hooks that compaction occurred.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
def run_on_compact(
    self,
    session_id: str,
    *,
    summary: str = "",
    tokens_before: int = 0,
    tokens_saved: int = 0,
    events_summarized: int = 0,
) -> None:
    """Notify hooks that compaction occurred."""
    for hook in self.on_compact:
        try:
            hook(
                session_id,
                summary=summary,
                tokens_before=tokens_before,
                tokens_saved=tokens_saved,
                events_summarized=events_summarized,
            )
        except Exception:
            logger.warning("on_compact hook failed", exc_info=True)

run_on_event(session_id: str, event: EventRecord) -> None

Notify hooks that an event was appended to a session transcript.

Runs on the event-append hot path (via the SessionEventBus observer), so every registered hook is itself fire-and-forget; this loop only dispatches and is failure-isolated per hook.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
387
388
389
390
391
392
393
394
395
396
397
398
def run_on_event(self, session_id: str, event: EventRecord) -> None:
    """Notify hooks that an event was appended to a session transcript.

    Runs on the event-append hot path (via the ``SessionEventBus``
    observer), so every registered hook is itself fire-and-forget; this
    loop only dispatches and is failure-isolated per hook.
    """
    for hook in self.on_event:
        try:
            hook(session_id, event)
        except Exception:
            logger.warning("on_event hook failed", exc_info=True)

run_on_session_end(session_id: str, error: str | None = None) -> list[OutcomeAssertion]

Notify hooks a session ended; collect any outcome assertions they report.

A hook MAY return an :class:OutcomeAssertion to report that the session's purpose was not achieved. Returning None asserts nothing — see that class for why absence must stay distinguishable from success.

The return value is COLLECTED, not discarded — discarding it is how a job that never reached its terminal state still presents as a clean completion, with the hook able to SEE the failure and nowhere to put it.

Failure-isolated per hook like every other lifecycle dispatch, and deliberately strict about what it accepts: a hook returning something that is not an assertion is logged and IGNORED rather than coerced, because a truthy stray return would otherwise invent a failure.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
def run_on_session_end(
    self, session_id: str, error: str | None = None
) -> list[OutcomeAssertion]:
    """Notify hooks a session ended; collect any outcome assertions they report.

    A hook MAY return an :class:`OutcomeAssertion` to report that the
    session's purpose was not achieved. Returning ``None`` asserts nothing
    — see that class for why absence must stay distinguishable from success.

    The return value is COLLECTED, not discarded — discarding it is how a
    job that never reached its terminal state still presents as a clean
    completion, with the hook able to SEE the failure and nowhere to put
    it.

    Failure-isolated per hook like every other lifecycle dispatch, and
    deliberately strict about what it accepts: a hook returning something
    that is not an assertion is logged and IGNORED rather than coerced,
    because a truthy stray return would otherwise invent a failure.
    """
    assertions: list[OutcomeAssertion] = []
    for hook in self.on_session_end:
        try:
            reported = hook(session_id, error)
        except Exception:
            logger.warning("on_session_end hook failed", exc_info=True)
            continue
        if reported is None:
            continue
        if not isinstance(reported, OutcomeAssertion):
            logger.warning(
                "on_session_end hook returned {}, not an OutcomeAssertion; ignoring",
                type(reported).__name__,
            )
            continue
        if not reported.source:
            # Re-CONSTRUCTED rather than ``model_copy``d: the stamp is a
            # function name of unbounded length, and model_copy skips the
            # validator that bounds it.
            reported = OutcomeAssertion(
                reason=reported.reason,
                detail=reported.detail,
                source=getattr(hook, "__name__", "") or type(hook).__name__,
            )
        assertions.append(reported)
    return assertions

run_on_session_start(session_id: str) -> None

Notify hooks that a session has started.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
333
334
335
336
337
338
339
def run_on_session_start(self, session_id: str) -> None:
    """Notify hooks that a session has started."""
    for hook in self.on_session_start:
        try:
            hook(session_id)
        except Exception:
            logger.warning("on_session_start hook failed", exc_info=True)

run_permission_request(action_step: ActionStep, decision: PermissionDecision) -> PermissionDecision

Apply permission hooks to a decision outcome.

Parameters:

Name Type Description Default
action_step ActionStep

Action step under review.

required
decision PermissionDecision

Current decision to modify.

required

Returns:

Type Description
PermissionDecision

Updated permission decision after hooks run.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
def run_permission_request(
    self, action_step: ActionStep, decision: PermissionDecision
) -> PermissionDecision:
    """Apply permission hooks to a decision outcome.

    Args:
        action_step: Action step under review.
        decision: Current decision to modify.

    Returns:
        Updated permission decision after hooks run.
    """
    for hook in self.permission_request:
        try:
            decision = hook(action_step, decision)
        except Exception:
            logger.warning("Permission hook failed", exc_info=True)
    return decision

run_post_tool_use(action_step: ActionStep, result: MockSpeaker) -> MockSpeaker

Apply post-tool hooks to a tool result.

Parameters:

Name Type Description Default
action_step ActionStep

Action step that was executed.

required
result MockSpeaker

Result returned by the tool.

required

Returns:

Type Description
MockSpeaker

Updated result after hooks run.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
def run_post_tool_use(self, action_step: ActionStep, result: MockSpeaker) -> MockSpeaker:
    """Apply post-tool hooks to a tool result.

    Args:
        action_step: Action step that was executed.
        result: Result returned by the tool.

    Returns:
        Updated result after hooks run.
    """
    for hook in self.post_tool_use:
        try:
            result = hook(action_step, result)
        except Exception:
            logger.warning("Post-tool hook failed", exc_info=True)
    return result

run_pre_compact(events: Iterable[EventRecord]) -> list[EventRecord]

Apply compaction hooks to events prior to summarization.

Parameters:

Name Type Description Default
events Iterable[EventRecord]

Iterable of event records.

required

Returns:

Type Description
list[EventRecord]

List of event records after hooks run.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
def run_pre_compact(self, events: Iterable[EventRecord]) -> list[EventRecord]:
    """Apply compaction hooks to events prior to summarization.

    Args:
        events: Iterable of event records.

    Returns:
        List of event records after hooks run.
    """
    event_list: list[EventRecord] = list(events)
    for hook in self.pre_compact:
        try:
            event_list = hook(event_list)
        except Exception:
            logger.warning("Pre-compact hook failed", exc_info=True)
    return event_list

run_pre_tool_use(action_step: ActionStep) -> ActionStep

Apply pre-tool hooks to an action step.

Parameters:

Name Type Description Default
action_step ActionStep

Action step to process.

required

Returns:

Type Description
ActionStep

Updated action step after hooks run.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
def run_pre_tool_use(self, action_step: ActionStep) -> ActionStep:
    """Apply pre-tool hooks to an action step.

    Args:
        action_step: Action step to process.

    Returns:
        Updated action step after hooks run.
    """
    for hook in self.pre_tool_use:
        try:
            action_step = hook(action_step)
        except Exception:
            logger.warning("Pre-tool hook failed", exc_info=True)
    return action_step

OutcomeAssertion

Bases: BaseModel

A session-end hook's report that the session's PURPOSE was not achieved.

The return channel a session-end hook needs in order to CONTRADICT a clean terminal. A session can end with no exception, no halt and no blocked envelope — every signal the loop owns says success — while the job it existed to perform never reached its terminal state. Nothing the loop can see distinguishes that from a real completion. Only the hook holding the owning job can, and until this existed it had no way to say so: the return value of run_on_session_end was discarded, so the one component able to evaluate a session's real outcome could write a warning to its own job log and nothing more.

Absence is not an assertion. A hook that returns None — every command hook, every http hook, and any python hook that declines to judge — reports nothing, which is why the contract is an OPTIONAL RETURN and not a boolean: "did not report" and "reported success" must never collapse into each other, or the channel would invent an assertion for every hook that does not use it.

reason is a PRODUCT-OWNED token and deliberately not a Literal: core would otherwise have to learn every product's vocabulary before that product could tell the truth about itself. It never drives dispatch — the status this produces comes from the TYPE of this object, never from parsing the string — so it stays data, not a magic string.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
class OutcomeAssertion(BaseModel):
    """A session-end hook's report that the session's PURPOSE was not achieved.

    The return channel a session-end hook needs in order to CONTRADICT a clean
    terminal. A session can end with no exception, no halt and no blocked
    envelope — every signal the loop owns says success — while the job it
    existed to perform never reached its terminal state. Nothing the loop can
    see distinguishes that from a real completion. Only the hook holding the
    owning job can, and until this existed it had no way to say so: the return
    value of ``run_on_session_end`` was discarded, so the one component able to
    evaluate a session's real outcome could write a warning to its own job log
    and nothing more.

    **Absence is not an assertion.** A hook that returns ``None`` — every
    command hook, every http hook, and any python hook that declines to judge —
    reports nothing, which is why the contract is an OPTIONAL RETURN and not a
    boolean: "did not report" and "reported success" must never collapse into
    each other, or the channel would invent an assertion for every hook that
    does not use it.

    ``reason`` is a PRODUCT-OWNED token and deliberately not a ``Literal``:
    core would otherwise have to learn every product's vocabulary before that
    product could tell the truth about itself. It never drives dispatch — the
    status this produces comes from the TYPE of this object, never from parsing
    the string — so it stays data, not a magic string.
    """

    model_config = ConfigDict(extra="forbid")

    reason: str = Field(
        min_length=1,
        description="Short product-owned token naming the outcome that was not met.",
    )
    detail: str = Field(
        default="",
        description="One-line factual detail. Reaches clients, so it is bounded.",
    )
    source: str = Field(
        default="",
        description="Which hook asserted; stamped by the manager for forensics.",
    )

    @model_validator(mode="before")
    @classmethod
    def _bound_fields(cls, data: object) -> object:
        """Normalize and CLAMP at definition, so no call site can over-run a cap.

        Clamping rather than rejecting is the ``RunError`` precedent: these
        values are bounded because they are persisted and replayed to clients,
        but an assertion is an honesty signal and refusing an over-long one
        would discard the report entirely.
        """
        if not isinstance(data, dict):
            return data
        data = dict(data)
        for key, cap in (
            ("reason", _ASSERTION_TOKEN_MAX_CHARS),
            ("source", _ASSERTION_TOKEN_MAX_CHARS),
            ("detail", _ASSERTION_DETAIL_MAX_CHARS),
        ):
            raw = data.get(key)
            if not isinstance(raw, str):
                continue
            # A token is one bare word; a detail is one line.
            joiner = "_" if key != "detail" else " "
            data[key] = joiner.join(raw.split())[:cap]
        return data

default_hook_manager() -> HookManager

Create a hook manager with no custom hooks registered.

Returns:

Type Description
HookManager

Empty HookManager instance.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
726
727
728
729
730
731
732
def default_hook_manager() -> HookManager:
    """Create a hook manager with no custom hooks registered.

    Returns:
        Empty HookManager instance.
    """
    return HookManager()

merge_plugin_hooks(manager: HookManager, hooks_json: dict[str, Any], plugin_root: str) -> None

Translate Claude Code plugin hooks.json into HookManager callbacks.

Supports: PreToolUse, PostToolUse, SessionStart, SessionEnd. Substitutes ${CLAUDE_PLUGIN_ROOT} in command strings.

Source code in packages/mewbo_core/src/mewbo_core/hooks.py
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
def merge_plugin_hooks(
    manager: HookManager,
    hooks_json: dict[str, Any],
    plugin_root: str,
) -> None:
    """Translate Claude Code plugin hooks.json into HookManager callbacks.

    Supports: PreToolUse, PostToolUse, SessionStart, SessionEnd.
    Substitutes ${CLAUDE_PLUGIN_ROOT} in command strings.
    """
    from mewbo_core.config import HookEntry
    from mewbo_core.tooling.plugins import substitute_plugin_vars

    raw_hooks = hooks_json.get("hooks", {})
    for cc_event, entry_groups in raw_hooks.items():
        mapping = _PLUGIN_HOOK_MAP.get(cc_event)
        if mapping is None:
            continue
        slot_name, factory = mapping
        for group in entry_groups:
            matcher = group.get("matcher")
            for hook_def in group.get("hooks", []):
                if hook_def.get("type") != "command":
                    continue
                command = hook_def.get("command")
                if not command:
                    continue
                command = substitute_plugin_vars(command, plugin_root)
                entry = HookEntry(
                    type="command",
                    command=command,
                    matcher=matcher,
                    timeout=hook_def.get("timeout", 30),
                )
                getattr(manager, slot_name).append(factory(entry))

mewbo_core.common

Common helpers shared across the assistant runtime.

InstructionSource dataclass

A single source of project/user instructions.

Source code in packages/mewbo_core/src/mewbo_core/common.py
269
270
271
272
273
274
275
276
@dataclass
class InstructionSource:
    """A single source of project/user instructions."""

    content: str
    path: str
    level: str  # "user", "project", "rules", "local"
    priority: int  # Higher = takes precedence in composition

MockSpeaker

Bases: NamedTuple

Simple mock response container used across tools and tests.

Source code in packages/mewbo_core/src/mewbo_core/common.py
28
29
30
31
32
33
34
35
36
37
38
39
class MockSpeaker(NamedTuple):
    """Simple mock response container used across tools and tests."""

    content: str
    # Image content parts a multimodal tool returned, in LiteLLM's
    # ``{"type": "image_url", ...}`` shape. Defaulted and additive on purpose:
    # ``content`` stays a plain ``str`` for all ~100 construction sites and
    # every reader of it, so a tool that returns no image is unchanged in both
    # type and behaviour. The loop lifts these into the ``ToolMessage``
    # alongside the text, which is what makes them an image block inside the
    # provider's native ``tool_result``.
    images: tuple[dict[str, Any], ...] = ()

count_tokens(text: str, model: str = 'gpt-4') -> int

Estimate token count for text using tiktoken.

Falls back to a rough character-based estimate if encoding lookup fails.

Source code in packages/mewbo_core/src/mewbo_core/common.py
220
221
222
223
224
225
226
227
228
229
def count_tokens(text: str, model: str = "gpt-4") -> int:
    """Estimate token count for text using tiktoken.

    Falls back to a rough character-based estimate if encoding lookup fails.
    """
    try:
        enc = tiktoken.encoding_for_model(model)
        return len(enc.encode(text))
    except Exception:
        return len(text) // 4  # Rough fallback

discover_all_instructions(cwd: str | None = None) -> list[InstructionSource]

Discover instructions from all levels, ordered by priority (lowest first).

Levels (ascending priority): 1. User: ~/.claude/CLAUDE.md (priority 10) 2. Project: CLAUDE.md, .claude/CLAUDE.md walking up to git root (priority 20-29) 3. Rules: .claude/rules/*.md in CWD (priority 30) 4. Local: CLAUDE.local.md in CWD (priority 40)

Source code in packages/mewbo_core/src/mewbo_core/common.py
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
def discover_all_instructions(cwd: str | None = None) -> list[InstructionSource]:
    """Discover instructions from all levels, ordered by priority (lowest first).

    Levels (ascending priority):
    1. User:    ~/.claude/CLAUDE.md (priority 10)
    2. Project: CLAUDE.md, .claude/CLAUDE.md walking up to git root (priority 20-29)
    3. Rules:   .claude/rules/*.md in CWD (priority 30)
    4. Local:   CLAUDE.local.md in CWD (priority 40)
    """
    sources: list[InstructionSource] = []
    work_dir = Path(cwd) if cwd else Path.cwd()

    # 1. User level
    user_claude = Path.home() / ".claude" / "CLAUDE.md"
    if user_claude.is_file():
        content = user_claude.read_text(encoding="utf-8", errors="replace").strip()
        if content and not content.startswith(_NOLOAD_MARKER):
            sources.append(
                InstructionSource(content=content, path=str(user_claude), level="user", priority=10)
            )

    # 2. Project level — walk from CWD up to git root (or filesystem root)
    git_root = _find_git_root(work_dir)
    stop_at = git_root or Path(work_dir.anchor)
    current = work_dir
    depth = 0
    while current >= stop_at:
        for filename in ("CLAUDE.md", ".claude/CLAUDE.md"):
            candidate = current / filename
            if candidate.is_file():
                content = candidate.read_text(encoding="utf-8", errors="replace").strip()
                if content and not content.startswith(_NOLOAD_MARKER):
                    # Closer to CWD = higher priority within project level
                    prio = 20 + min(depth, 9)  # 20 (CWD) to 29 (root)
                    sources.append(
                        InstructionSource(
                            content=content, path=str(candidate), level="project", priority=prio
                        )
                    )
        parent = current.parent
        if parent == current:
            break
        current = parent
        depth += 1

    # 3. Rules level — .claude/rules/*.md in CWD
    rules_dir = work_dir / ".claude" / "rules"
    if rules_dir.is_dir():
        for md_file in sorted(rules_dir.glob("*.md")):
            if md_file.is_file():
                content = md_file.read_text(encoding="utf-8", errors="replace").strip()
                if content and not content.startswith(_NOLOAD_MARKER):
                    sources.append(
                        InstructionSource(
                            content=content, path=str(md_file), level="rules", priority=30
                        )
                    )

    # 4. Local level
    local_claude = work_dir / "CLAUDE.local.md"
    if local_claude.is_file():
        content = local_claude.read_text(encoding="utf-8", errors="replace").strip()
        if content and not content.startswith(_NOLOAD_MARKER):
            sources.append(
                InstructionSource(
                    content=content, path=str(local_claude), level="local", priority=40
                )
            )

    # Sort by priority (lowest first — will be composed in order, higher priority last)
    sources.sort(key=lambda s: s.priority)
    return sources

discover_project_instructions(cwd: str | None = None) -> str | None

Discover and load project instruction files. Uses hierarchical discovery.

Falls back to a root AGENTS.md when no hierarchical sources are found. Files containing <!-- mewbo:noload --> on the first line are skipped.

Additionally walks the subtree to build a lightweight index of nested instruction files so the model knows they exist and can read them on demand.

Returns the composed instruction text, or None if no files are found.

Source code in packages/mewbo_core/src/mewbo_core/common.py
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
def discover_project_instructions(cwd: str | None = None) -> str | None:
    """Discover and load project instruction files. Uses hierarchical discovery.

    Falls back to a root ``AGENTS.md`` when no hierarchical sources are found.
    Files containing ``<!-- mewbo:noload -->`` on the first line are skipped.

    Additionally walks the subtree to build a lightweight index of nested
    instruction files so the model knows they exist and can read them on demand.

    Returns the composed instruction text, or ``None`` if no files are found.
    """
    sources = discover_all_instructions(cwd)
    if not sources:
        work_dir = Path(cwd) if cwd else Path.cwd()
        agents_md = work_dir / "AGENTS.md"
        if agents_md.is_file():
            content = agents_md.read_text(encoding="utf-8", errors="replace").strip()
            if content and not content.startswith(_NOLOAD_MARKER):
                return content
        # Even with no direct sources, subtree files may exist — fall through

    # Compose direct sources with section headers
    parts: list[str] = []
    for src in sources:
        header = f"# Instructions ({src.level}: {Path(src.path).name})"
        parts.append(f"{header}\n\n{src.content}")

    # Discover subtree instruction files (index only — no content injection)
    subtree = discover_subtree_instructions(cwd)
    if subtree:
        work_dir = Path(cwd) if cwd else Path.cwd()
        lines = []
        for src in subtree:
            rel = Path(src.path).relative_to(work_dir)
            lines.append(f"- {rel}")
        from mewbo_core.llm.prompt_registry import get_prompt_registry

        heading = get_prompt_registry().render(
            "common.instruction_headings", has_root=bool(sources)
        )
        parts.append(heading + "\n\n" + "\n".join(lines))

    if not parts:
        return None
    return "\n\n---\n\n".join(parts)

discover_subtree_instructions(cwd: str | None = None, *, max_depth: int = _MAX_SUBTREE_DEPTH) -> list[InstructionSource]

Walk DOWN from CWD to find CLAUDE.md and AGENTS.md in subdirectories.

Returns lightweight InstructionSource entries with empty content. The model is made aware these files exist and can read them on demand. Respects the <!-- mewbo:noload --> marker (checked via first line).

Source code in packages/mewbo_core/src/mewbo_core/common.py
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
def discover_subtree_instructions(
    cwd: str | None = None,
    *,
    max_depth: int = _MAX_SUBTREE_DEPTH,
) -> list[InstructionSource]:
    """Walk DOWN from CWD to find CLAUDE.md and AGENTS.md in subdirectories.

    Returns lightweight ``InstructionSource`` entries with *empty* content.
    The model is made aware these files exist and can read them on demand.
    Respects the ``<!-- mewbo:noload -->`` marker (checked via first line).
    """
    work_dir = Path(cwd) if cwd else Path.cwd()
    found: list[InstructionSource] = []
    for dirpath, dirnames, _filenames in os.walk(work_dir):
        rel = Path(dirpath).relative_to(work_dir)
        depth = len(rel.parts)
        # Prune hidden dirs and common non-project dirs (must happen before any continue)
        dirnames[:] = [
            d
            for d in dirnames
            if not d.startswith(".") and d not in ("node_modules", "__pycache__", ".venv", "venv")
        ]
        if depth == 0:
            continue  # Skip CWD itself — already handled by discover_all_instructions
        if depth > max_depth:
            dirnames.clear()
            continue
        for filename in _INSTRUCTION_FILENAMES:
            candidate = Path(dirpath) / filename
            if candidate.is_file():
                try:
                    first_line = candidate.open(encoding="utf-8", errors="replace").readline()
                except OSError:
                    continue
                if first_line.strip().startswith(_NOLOAD_MARKER):
                    continue
                found.append(
                    InstructionSource(
                        content="",
                        path=str(candidate),
                        level="subtree",
                        priority=50,
                    )
                )
    found.sort(key=lambda s: s.path)
    return found

format_tool_input(tool_input: object) -> str

Format a tool input for logs and prompts.

Source code in packages/mewbo_core/src/mewbo_core/common.py
509
510
511
512
513
def format_tool_input(tool_input: object) -> str:
    """Format a tool input for logs and prompts."""
    if isinstance(tool_input, dict):
        return json.dumps(tool_input, ensure_ascii=True)
    return str(tool_input)

get_git_context(cwd: str | None = None, max_status_chars: int = 2000) -> str | None

Gather git context (branch, status, recent commits) for system prompt injection.

Returns formatted git context string, or None if not in a git repo.

Source code in packages/mewbo_core/src/mewbo_core/common.py
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
def get_git_context(cwd: str | None = None, max_status_chars: int = 2000) -> str | None:
    """Gather git context (branch, status, recent commits) for system prompt injection.

    Returns formatted git context string, or None if not in a git repo.
    """
    work_dir = cwd or str(Path.cwd())

    def _run_git(*args: str) -> str | None:
        try:
            result = subprocess.run(
                ["git", *args],
                cwd=work_dir,
                capture_output=True,
                text=True,
                timeout=5,
            )
            return result.stdout.strip() if result.returncode == 0 else None
        except (subprocess.TimeoutExpired, FileNotFoundError, OSError):
            return None

    branch = _run_git("rev-parse", "--abbrev-ref", "HEAD")
    if branch is None:
        return None  # Not a git repo

    default_branch = _run_git("rev-parse", "--abbrev-ref", "origin/HEAD")
    if default_branch:
        default_branch = default_branch.replace("origin/", "")

    status = _run_git("status", "--short")
    if status and len(status) > max_status_chars:
        status = status[:max_status_chars] + "\n[truncated]"

    recent_log = _run_git("log", "--oneline", "-n", "5")

    from mewbo_core.llm.prompt_registry import get_prompt_registry

    return get_prompt_registry().render(
        "common.git_context",
        branch=branch,
        default_branch=default_branch or "",
        status=status or "",
        recent_log=recent_log or "",
    )

get_logger(name: str | None = None)

Get the logger for the module.

Source code in packages/mewbo_core/src/mewbo_core/common.py
199
200
201
202
203
204
def get_logger(name: str | None = None):
    """Get the logger for the module."""
    _configure_logging()
    if not name:
        name = __name__
    return loguru_logger.bind(name=name)

get_mock_speaker() -> type[MockSpeaker]

Return a mock speaker for testing.

Source code in packages/mewbo_core/src/mewbo_core/common.py
42
43
44
def get_mock_speaker() -> type[MockSpeaker]:
    """Return a mock speaker for testing."""
    return MockSpeaker

get_system_prompt(name: str = 'action-planner') -> str

Get the system prompt for the task queue.

Routes through the central prompt registry when name has a file.* entry (the standalone, system.txt-sized prompts inventoried in prompts/registry/files.yaml), which the registry does not strip, so this shim applies the .strip() its callers expect. Names the registry does not inventory (tool prompts loaded by path) fall back to a raw file read.

Source code in packages/mewbo_core/src/mewbo_core/common.py
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
def get_system_prompt(name: str = "action-planner") -> str:
    """Get the system prompt for the task queue.

    Routes through the central prompt registry when *name* has a ``file.*``
    entry (the standalone, ``system.txt``-sized prompts inventoried in
    ``prompts/registry/files.yaml``), which the registry does not strip, so
    this shim applies the ``.strip()`` its callers expect. Names the registry
    does not inventory (tool prompts loaded by path) fall back to a raw file
    read.
    """
    from mewbo_core.llm.prompt_registry import get_prompt_registry

    registry = get_prompt_registry()
    prompt_id = f"file.{name}"
    if registry.has(prompt_id):
        return registry.render(prompt_id).strip()

    logging = get_logger(name="core.common.get_system_prompt")
    prompt_resource = resources.files("mewbo_core").joinpath("prompts").joinpath(f"{name}.txt")
    with resources.as_file(prompt_resource) as system_prompt_path:
        with open(system_prompt_path, encoding="utf-8") as system_prompt_file:
            system_prompt = system_prompt_file.read()
        logging.debug("Getting system prompt from `{}`", system_prompt_path)
    del logging
    return system_prompt.strip()

get_unique_timestamp() -> int

Get a unique timestamp for the task queue.

Source code in packages/mewbo_core/src/mewbo_core/common.py
232
233
234
235
236
def get_unique_timestamp() -> int:
    """Get a unique timestamp for the task queue."""
    current_timestamp = int(time.time())
    unique_timestamp = str(current_timestamp)
    return int("".join(str(x) for x in map(int, unique_timestamp)))

ha_render_system_prompt(all_entities: object | None = None, name: str = 'homeassistant-set-state') -> str

Render the Home Assistant Jinja2 system prompt.

Source code in packages/mewbo_core/src/mewbo_core/common.py
633
634
635
636
637
638
639
640
def ha_render_system_prompt(
    all_entities: object | None = None,
    name: str = "homeassistant-set-state",
) -> str:
    """Render the Home Assistant Jinja2 system prompt."""
    if all_entities is not None:
        all_entities = str(all_entities).strip()
    return render_jinja_prompt(name, ALL_ENTITIES=all_entities)

num_tokens_from_string(string: str, encoding_name: str = 'cl100k_base') -> int

Get the number of tokens in a string using a specific model.

Source code in packages/mewbo_core/src/mewbo_core/common.py
212
213
214
215
216
217
def num_tokens_from_string(string: str, encoding_name: str = "cl100k_base") -> int:
    """Get the number of tokens in a string using a specific model."""
    # TODO: Add support for dynamic model selection
    encoding = tiktoken.get_encoding(encoding_name)
    num_tokens = len(encoding.encode(string))
    return num_tokens

pydantic_to_openai_tool(model_cls: type, *, name: str, schema_generator: type | None = None) -> dict[str, object]

Build an OpenAI function-calling tool dict from a Pydantic model.

Uses the model's docstring as the tool description and its JSON schema as the parameters. Strips Pydantic's title fields that are irrelevant to function-calling. Output matches the shape used by the existing hand-written internal tool schemas (SPAWN_AGENT_SCHEMA, etc.) so callers can migrate piecewise.

Parameters:

Name Type Description Default
model_cls type

A Pydantic BaseModel subclass defining the tool args.

required
name str

The tool name (function name visible to the LLM).

required
schema_generator type | None

Optional GenerateJsonSchema subclass, passed straight to pydantic. This is the supported seam for changing how the schema is NAMED or shaped — a caller that post-processes the emitted dict instead has to find every $ref string AND the discriminator mapping, and one that finds only some of them leaves the schema self-inconsistent.

None

Returns:

Type Description
dict[str, object]

{"type": "function", "function": {"name", "description", "parameters"}}

Raises:

Type Description
TypeError

If model_cls is not a Pydantic BaseModel subclass.

Source code in packages/mewbo_core/src/mewbo_core/common.py
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
def pydantic_to_openai_tool(
    model_cls: type, *, name: str, schema_generator: type | None = None
) -> dict[str, object]:
    """Build an OpenAI function-calling tool dict from a Pydantic model.

    Uses the model's docstring as the tool description and its JSON schema
    as the parameters. Strips Pydantic's ``title`` fields that are irrelevant
    to function-calling. Output matches the shape used by the existing
    hand-written internal tool schemas (``SPAWN_AGENT_SCHEMA``, etc.) so
    callers can migrate piecewise.

    Args:
        model_cls: A Pydantic ``BaseModel`` subclass defining the tool args.
        name: The tool name (function name visible to the LLM).
        schema_generator: Optional ``GenerateJsonSchema`` subclass, passed
            straight to pydantic. This is the supported seam for changing how
            the schema is NAMED or shaped — a caller that post-processes the
            emitted dict instead has to find every ``$ref`` string AND the
            discriminator mapping, and one that finds only some of them leaves
            the schema self-inconsistent.

    Returns:
        ``{"type": "function", "function": {"name", "description", "parameters"}}``

    Raises:
        TypeError: If ``model_cls`` is not a Pydantic ``BaseModel`` subclass.
    """
    from pydantic import BaseModel as _BM

    if not (isinstance(model_cls, type) and issubclass(model_cls, _BM)):
        raise TypeError(
            "pydantic_to_openai_tool requires a Pydantic BaseModel subclass"
        )
    kwargs = {"schema_generator": schema_generator} if schema_generator else {}
    params = _strip_schema_titles(model_cls.model_json_schema(**kwargs))  # type: ignore[arg-type]
    return {
        "type": "function",
        "function": {
            "name": name,
            "description": inspect.cleandoc(model_cls.__doc__ or ""),
            "parameters": params,
        },
    }

render_jinja_prompt(name: str, **variables: object) -> str

Render a Jinja2 prompt template from mewbo_core/prompts/.

Looks up {name}.j2 first, falls back to {name}.txt for backward-compatibility with existing prompts that use Jinja2 syntax inside .txt files (e.g. homeassistant-*.txt).

Parameters:

Name Type Description Default
name str

Template stem without extension.

required
**variables object

Keyword variables bound to the template.

{}

Returns:

Type Description
str

Rendered prompt string.

Raises:

Type Description
RuntimeError

If no template is found (chained from the final jinja2.TemplateNotFound via __cause__). Other Jinja2 errors (TemplateSyntaxError, UndefinedError) propagate directly and are NOT swallowed.

Source code in packages/mewbo_core/src/mewbo_core/common.py
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
def render_jinja_prompt(name: str, **variables: object) -> str:
    """Render a Jinja2 prompt template from ``mewbo_core/prompts/``.

    Looks up ``{name}.j2`` first, falls back to ``{name}.txt`` for
    backward-compatibility with existing prompts that use Jinja2 syntax
    inside ``.txt`` files (e.g. ``homeassistant-*.txt``).

    Args:
        name: Template stem without extension.
        **variables: Keyword variables bound to the template.

    Returns:
        Rendered prompt string.

    Raises:
        RuntimeError: If no template is found (chained from the final
            ``jinja2.TemplateNotFound`` via ``__cause__``). Other Jinja2
            errors (``TemplateSyntaxError``, ``UndefinedError``) propagate
            directly and are NOT swallowed.
    """
    # These prompts are inventoried in the central registry as ``file.*``
    # entries (so they can grow per-model overrides later), but rendering stays
    # here on a tolerant default-``Undefined`` env: callers (HA) rely on a
    # missing variable rendering blank rather than raising, which the registry's
    # ``StrictUndefined`` deliberately does not allow. Making these strict is a
    # behaviour change for a later phase, not this verbatim extraction.
    #
    # autoescape stays off (the default): the rendered result is LLM prompt text,
    # never browser-facing HTML — HTML-entity-encoding it would corrupt the prompt.
    log = get_logger(name="core.common.render_jinja_prompt")
    template_env = Environment(loader=PackageLoader("mewbo_core", "prompts"))
    last_exc: TemplateNotFound | None = None
    for suffix in (".j2", ".txt"):
        try:
            template = template_env.get_template(f"{name}{suffix}")
        except TemplateNotFound as exc:
            last_exc = exc
            continue
        log.debug("Rendered prompt `{}{}`", name, suffix)
        return template.render(**variables)
    raise RuntimeError(f"No template found for prompt '{name}'") from last_exc

session_log_context(session_id: str, log_dir: str | None = None)

Context manager that logs all session output to a session log file.

Source code in packages/mewbo_core/src/mewbo_core/common.py
188
189
190
191
192
193
194
195
196
@contextmanager
def session_log_context(session_id: str, log_dir: str | None = None):
    """Context manager that logs all session output to a session log file."""
    _ensure_session_log_sink(session_id, log_dir=log_dir)
    try:
        with loguru_logger.contextualize(session_id=session_id):
            yield
    finally:
        _release_session_log_sink(session_id)

set_cli_log_file(log_file_path: str, *, overwrite: bool = False, quiet_console: bool = True) -> str

Stream all CLI logs to a file, keeping the terminal (TUI) output clean.

Adds an unfiltered loguru file sink at the active verbosity level. When quiet_console is set (the default), the stderr sink is removed so log lines cannot interleave with the Rich/Textual UI — they go only to the file. overwrite truncates the file at startup instead of appending.

Returns the absolute path actually written to.

Source code in packages/mewbo_core/src/mewbo_core/common.py
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
def set_cli_log_file(
    log_file_path: str, *, overwrite: bool = False, quiet_console: bool = True
) -> str:
    """Stream all CLI logs to a file, keeping the terminal (TUI) output clean.

    Adds an unfiltered loguru file sink at the active verbosity level. When
    ``quiet_console`` is set (the default), the stderr sink is removed so log
    lines cannot interleave with the Rich/Textual UI — they go only to the
    file. ``overwrite`` truncates the file at startup instead of appending.

    Returns the absolute path actually written to.
    """
    global _CLI_LOG_SINK_ID, _STDERR_SINK_ID
    _configure_logging()
    resolved = os.path.abspath(os.path.expanduser(log_file_path))
    parent = os.path.dirname(resolved)
    if parent:
        os.makedirs(parent, exist_ok=True)
    if _CLI_LOG_SINK_ID is not None:
        loguru_logger.remove(_CLI_LOG_SINK_ID)
        _CLI_LOG_SINK_ID = None
    _CLI_LOG_SINK_ID = loguru_logger.add(
        resolved,
        level=_resolve_log_level(),
        format=_session_log_format(),
        colorize=False,
        mode="w" if overwrite else "a",
        enqueue=True,
        diagnose=False,  # keep secret-bearing locals out of persisted tracebacks
    )
    if quiet_console and _STDERR_SINK_ID is not None:
        loguru_logger.remove(_STDERR_SINK_ID)
        _STDERR_SINK_ID = None
    return resolved

utc_now_iso() -> str

Return an ISO-8601 UTC timestamp string (the shared storage timestamp).

Source code in packages/mewbo_core/src/mewbo_core/common.py
207
208
209
def utc_now_iso() -> str:
    """Return an ISO-8601 UTC timestamp string (the shared storage timestamp)."""
    return datetime.now(timezone.utc).isoformat()

mewbo_core.contracts.errors

Core error types for tool/runtime coordination.

ToolInputError

Bases: Exception

Raised when a tool input is invalid but the tool remains healthy.

Source code in packages/mewbo_core/src/mewbo_core/contracts/errors.py
7
8
class ToolInputError(Exception):
    """Raised when a tool input is invalid but the tool remains healthy."""

mewbo_core.session.notifications

Lightweight notification storage for Mewbo.

NotificationRecord dataclass

Typed record for serialized notifications.

Source code in packages/mewbo_core/src/mewbo_core/session/notifications.py
21
22
23
24
25
26
27
28
29
30
31
32
33
34
@dataclass(frozen=True)
class NotificationRecord:
    """Typed record for serialized notifications."""

    id: str
    title: str
    message: str
    level: str
    created_at: str
    dismissed: bool
    session_id: str | None = None
    dismissed_at: str | None = None
    event_type: str | None = None
    metadata: dict[str, object] | None = None

NotificationStore

JSON-backed notification store for single-user UI.

Source code in packages/mewbo_core/src/mewbo_core/session/notifications.py
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
class NotificationStore:
    """JSON-backed notification store for single-user UI."""

    def __init__(self, root_dir: str | None = None, filename: str = "notifications.json") -> None:
        """Initialize the notification store location."""
        if root_dir is None:
            root_dir = get_config_value("runtime", "session_dir", default="./data/sessions")
        root_dir = os.path.abspath(root_dir)
        os.makedirs(root_dir, exist_ok=True)
        self._path = os.path.join(root_dir, filename)
        self._lock = threading.Lock()

    def _load(self) -> list[dict[str, object]]:
        """Load notification records from disk."""
        if not os.path.exists(self._path):
            return []
        with open(self._path, encoding="utf-8") as handle:
            try:
                data = json.load(handle)
            except json.JSONDecodeError:
                return []
        if isinstance(data, list):
            return data
        return []

    def _save(self, data: list[dict[str, object]]) -> None:
        """Persist notification records to disk."""
        with open(self._path, "w", encoding="utf-8") as handle:
            json.dump(data, handle, indent=2)

    def list(self, *, include_dismissed: bool = False) -> list[dict[str, object]]:
        """Return notifications, optionally including dismissed ones."""
        with self._lock:
            data = self._load()
        if not include_dismissed:
            data = [item for item in data if not item.get("dismissed")]
        return sorted(
            data,
            key=lambda item: str(item.get("created_at", "")),
            reverse=True,
        )

    def add(
        self,
        *,
        title: str,
        message: str,
        level: str = "info",
        session_id: str | None = None,
        event_type: str | None = None,
        metadata: dict[str, object] | None = None,
    ) -> dict[str, object]:
        """Add a new notification record and return it."""
        record = NotificationRecord(
            id=uuid.uuid4().hex,
            title=title,
            message=message,
            level=level,
            created_at=_utc_now(),
            dismissed=False,
            session_id=session_id,
            dismissed_at=None,
            event_type=event_type,
            metadata=metadata,
        )
        payload = record.__dict__
        with self._lock:
            data = self._load()
            data.append(payload)
            self._save(data)
        return payload

    def dismiss(self, ids: Sequence[str]) -> int:
        """Mark notifications as dismissed."""
        if not ids:
            return 0
        dismissed_at = _utc_now()
        updated = 0
        with self._lock:
            data = self._load()
            for item in data:
                if item.get("id") in ids and not item.get("dismissed"):
                    item["dismissed"] = True
                    item["dismissed_at"] = dismissed_at
                    updated += 1
            self._save(data)
        return updated

    def clear(self, *, dismissed_only: bool = True) -> int:
        """Clear dismissed notifications (or all when requested)."""
        with self._lock:
            data = self._load()
            if dismissed_only:
                remaining = [item for item in data if not item.get("dismissed")]
            else:
                remaining = []
            removed = len(data) - len(remaining)
            self._save(remaining)
        return removed

__init__(root_dir: str | None = None, filename: str = 'notifications.json') -> None

Initialize the notification store location.

Source code in packages/mewbo_core/src/mewbo_core/session/notifications.py
40
41
42
43
44
45
46
47
def __init__(self, root_dir: str | None = None, filename: str = "notifications.json") -> None:
    """Initialize the notification store location."""
    if root_dir is None:
        root_dir = get_config_value("runtime", "session_dir", default="./data/sessions")
    root_dir = os.path.abspath(root_dir)
    os.makedirs(root_dir, exist_ok=True)
    self._path = os.path.join(root_dir, filename)
    self._lock = threading.Lock()

add(*, title: str, message: str, level: str = 'info', session_id: str | None = None, event_type: str | None = None, metadata: dict[str, object] | None = None) -> dict[str, object]

Add a new notification record and return it.

Source code in packages/mewbo_core/src/mewbo_core/session/notifications.py
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
def add(
    self,
    *,
    title: str,
    message: str,
    level: str = "info",
    session_id: str | None = None,
    event_type: str | None = None,
    metadata: dict[str, object] | None = None,
) -> dict[str, object]:
    """Add a new notification record and return it."""
    record = NotificationRecord(
        id=uuid.uuid4().hex,
        title=title,
        message=message,
        level=level,
        created_at=_utc_now(),
        dismissed=False,
        session_id=session_id,
        dismissed_at=None,
        event_type=event_type,
        metadata=metadata,
    )
    payload = record.__dict__
    with self._lock:
        data = self._load()
        data.append(payload)
        self._save(data)
    return payload

clear(*, dismissed_only: bool = True) -> int

Clear dismissed notifications (or all when requested).

Source code in packages/mewbo_core/src/mewbo_core/session/notifications.py
125
126
127
128
129
130
131
132
133
134
135
def clear(self, *, dismissed_only: bool = True) -> int:
    """Clear dismissed notifications (or all when requested)."""
    with self._lock:
        data = self._load()
        if dismissed_only:
            remaining = [item for item in data if not item.get("dismissed")]
        else:
            remaining = []
        removed = len(data) - len(remaining)
        self._save(remaining)
    return removed

dismiss(ids: Sequence[str]) -> int

Mark notifications as dismissed.

Source code in packages/mewbo_core/src/mewbo_core/session/notifications.py
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
def dismiss(self, ids: Sequence[str]) -> int:
    """Mark notifications as dismissed."""
    if not ids:
        return 0
    dismissed_at = _utc_now()
    updated = 0
    with self._lock:
        data = self._load()
        for item in data:
            if item.get("id") in ids and not item.get("dismissed"):
                item["dismissed"] = True
                item["dismissed_at"] = dismissed_at
                updated += 1
        self._save(data)
    return updated

list(*, include_dismissed: bool = False) -> list[dict[str, object]]

Return notifications, optionally including dismissed ones.

Source code in packages/mewbo_core/src/mewbo_core/session/notifications.py
67
68
69
70
71
72
73
74
75
76
77
def list(self, *, include_dismissed: bool = False) -> list[dict[str, object]]:
    """Return notifications, optionally including dismissed ones."""
    with self._lock:
        data = self._load()
    if not include_dismissed:
        data = [item for item in data if not item.get("dismissed")]
    return sorted(
        data,
        key=lambda item: str(item.get("created_at", "")),
        reverse=True,
    )

mewbo_core.session.share_store

Session share token storage.

ShareStore

JSON-backed share token store for session exports.

Source code in packages/mewbo_core/src/mewbo_core/session/share_store.py
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
class ShareStore:
    """JSON-backed share token store for session exports."""

    def __init__(self, root_dir: str | None = None, filename: str = "shares.json") -> None:
        """Initialize the share token store location."""
        if root_dir is None:
            root_dir = get_config_value("runtime", "session_dir", default="./data/sessions")
        root_dir = os.path.abspath(root_dir)
        os.makedirs(root_dir, exist_ok=True)
        self._path = os.path.join(root_dir, filename)
        self._lock = threading.Lock()

    def _load(self) -> dict[str, dict[str, object]]:
        """Load share token records from disk."""
        if not os.path.exists(self._path):
            return {}
        with open(self._path, encoding="utf-8") as handle:
            try:
                data = json.load(handle)
            except json.JSONDecodeError:
                return {}
        if isinstance(data, dict):
            return data
        return {}

    def _save(self, data: dict[str, dict[str, object]]) -> None:
        """Persist share token records to disk."""
        with open(self._path, "w", encoding="utf-8") as handle:
            json.dump(data, handle, indent=2)

    def create(self, session_id: str) -> dict[str, object]:
        """Create and store a new share token."""
        token = uuid.uuid4().hex
        record: dict[str, object] = {"session_id": session_id, "created_at": _utc_now()}
        with self._lock:
            data = self._load()
            data[token] = record
            self._save(data)
        return {"token": token, **record}

    def resolve(self, token: str) -> dict[str, object] | None:
        """Resolve a share token to its record."""
        if not token:
            return None
        with self._lock:
            data = self._load()
            record = data.get(token)
        if not record:
            return None
        return {"token": token, **record}

    def revoke(self, token: str) -> bool:
        """Revoke a share token."""
        if not token:
            return False
        with self._lock:
            data = self._load()
            if token not in data:
                return False
            data.pop(token, None)
            self._save(data)
        return True

__init__(root_dir: str | None = None, filename: str = 'shares.json') -> None

Initialize the share token store location.

Source code in packages/mewbo_core/src/mewbo_core/session/share_store.py
22
23
24
25
26
27
28
29
def __init__(self, root_dir: str | None = None, filename: str = "shares.json") -> None:
    """Initialize the share token store location."""
    if root_dir is None:
        root_dir = get_config_value("runtime", "session_dir", default="./data/sessions")
    root_dir = os.path.abspath(root_dir)
    os.makedirs(root_dir, exist_ok=True)
    self._path = os.path.join(root_dir, filename)
    self._lock = threading.Lock()

create(session_id: str) -> dict[str, object]

Create and store a new share token.

Source code in packages/mewbo_core/src/mewbo_core/session/share_store.py
49
50
51
52
53
54
55
56
57
def create(self, session_id: str) -> dict[str, object]:
    """Create and store a new share token."""
    token = uuid.uuid4().hex
    record: dict[str, object] = {"session_id": session_id, "created_at": _utc_now()}
    with self._lock:
        data = self._load()
        data[token] = record
        self._save(data)
    return {"token": token, **record}

resolve(token: str) -> dict[str, object] | None

Resolve a share token to its record.

Source code in packages/mewbo_core/src/mewbo_core/session/share_store.py
59
60
61
62
63
64
65
66
67
68
def resolve(self, token: str) -> dict[str, object] | None:
    """Resolve a share token to its record."""
    if not token:
        return None
    with self._lock:
        data = self._load()
        record = data.get(token)
    if not record:
        return None
    return {"token": token, **record}

revoke(token: str) -> bool

Revoke a share token.

Source code in packages/mewbo_core/src/mewbo_core/session/share_store.py
70
71
72
73
74
75
76
77
78
79
80
def revoke(self, token: str) -> bool:
    """Revoke a share token."""
    if not token:
        return False
    with self._lock:
        data = self._load()
        if token not in data:
            return False
        data.pop(token, None)
        self._save(data)
    return True

mewbo_core.llm.llm

Model configuration helpers for ChatLiteLLM.

ChatModel

Bases: Protocol

Protocol for LangChain-compatible chat models.

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
39
40
41
42
43
44
45
46
47
48
49
50
class ChatModel(Protocol):
    """Protocol for LangChain-compatible chat models."""

    def invoke(
        self, input_data: object, config: object | None = None, **kwargs: object
    ) -> BaseMessage:
        """Invoke the model synchronously."""

    async def ainvoke(
        self, input_data: object, config: object | None = None, **kwargs: object
    ) -> BaseMessage:
        """Invoke the model asynchronously."""

ainvoke(input_data: object, config: object | None = None, **kwargs: object) -> BaseMessage async

Invoke the model asynchronously.

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
47
48
49
50
async def ainvoke(
    self, input_data: object, config: object | None = None, **kwargs: object
) -> BaseMessage:
    """Invoke the model asynchronously."""

invoke(input_data: object, config: object | None = None, **kwargs: object) -> BaseMessage

Invoke the model synchronously.

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
42
43
44
45
def invoke(
    self, input_data: object, config: object | None = None, **kwargs: object
) -> BaseMessage:
    """Invoke the model synchronously."""

build_chat_model(model_name: str, *, openai_api_base: str | None = None, api_key: str | None = None) -> ChatModel

Build a ChatLiteLLM model with reasoning-effort compatibility.

openai_api_base and api_key default to llm.api_base and llm.api_key from config when None. Pass them explicitly only to override the configured values (e.g. tests, multi-tenant routing).

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
def build_chat_model(
    model_name: str,
    *,
    openai_api_base: str | None = None,
    api_key: str | None = None,
) -> ChatModel:
    """Build a ChatLiteLLM model with reasoning-effort compatibility.

    ``openai_api_base`` and ``api_key`` default to ``llm.api_base`` and
    ``llm.api_key`` from config when ``None``. Pass them explicitly only
    to override the configured values (e.g. tests, multi-tenant routing).
    """
    try:
        ChatLiteLLM = _observability_chat_litellm_class()
    except ImportError as exc:  # pragma: no cover - dependency guard
        raise ImportError("langchain-litellm is required to build ChatLiteLLM") from exc

    if openai_api_base is None:
        openai_api_base = str(get_config_value("llm", "api_base", default="") or "")
    if api_key is None:
        api_key = str(get_config_value("llm", "api_key", default="") or "")
    proxy_prefix = (
        str(get_config_value("llm", "proxy_model_prefix", default="openai") or "openai")
        .strip()
        .strip("/")
        or "openai"
    )

    # Hydrate proxy capabilities once per process per distinct api_base, so the
    # cache-support gate below sees the proxy's advertised model_info instead
    # of just LiteLLM's bundled allowlist. No-op when api_base is unset.
    if openai_api_base:
        register_proxy_model_capabilities(openai_api_base, api_key)

    reasoning_effort = resolve_reasoning_effort(model_name)

    model_kwargs: dict[str, Any] = {
        "drop_params": True,  # Must be in model_kwargs to reach litellm.acompletion();
        # ChatLiteLLM has no drop_params field so top-level kwarg is silently ignored.
        #
        # The SAME routing trap, and it costs far more than a dropped param.
        # ``ChatLiteLLM.max_retries`` (set to 0 below) is consumed by the
        # LangChain-level tenacity decorator and is NEVER forwarded into
        # ``litellm.acompletion`` — it appears in neither ``_default_params`` nor
        # ``_client_params`` — so litellm builds the provider SDK client with its
        # own default of 2, i.e. three attempts. Each attempt is bounded by
        # ``request_timeout`` and re-sends the whole prompt, so the HTTP timeout
        # an operator reads as a 60s idle bound is really ~3x that before
        # anything surfaces, which is how a silent provider consumed the entire
        # per-call ceiling with no retry event and no error text. Retries belong
        # to the ToolUseLoop, which bounds and REPORTS each one.
        "max_retries": 0,
    }
    if reasoning_effort is not None:
        model_kwargs["reasoning_effort"] = reasoning_effort

    if model_supports_prompt_caching(model_name):
        # LiteLLM's hook attaches the right native marker for whichever provider
        # the call routes to (Anthropic content-block cache_control, Bedrock
        # with ttl-sanitisation).  We only declare *where* — at the system
        # message — and let LiteLLM own the per-provider syntax.  Auto-cache
        # providers (OpenAI) silently drop the kwarg and still surface savings
        # via usage_metadata.input_token_details.cache_read.
        model_kwargs["cache_control_injection_points"] = [
            {"location": "message", "role": "system", "control": {"type": "ephemeral"}}
        ]

    kwargs: dict[str, Any] = {
        "model": _resolve_litellm_model(model_name, openai_api_base, proxy_prefix),
        # Streaming is what makes every other bound mean what its docstring says.
        # ``ChatLiteLLM``'s constructor puts ``streaming`` into ``model_fields_set``
        # with ``False``, and langchain-core's ``_streaming_disabled`` treats that
        # as an opt-out that OVERRIDES an affirmative ``stream=True`` — so
        # ``astream()`` silently degraded to one buffered ``ainvoke``. Three things
        # were dead as a result: the transport bound became a ceiling on total
        # generation time (a buffered response sends nothing until it is complete,
        # so the wait for the first byte IS the whole generation), the first-token
        # and stream-idle arms of ``CallDeadline`` never armed because nothing
        # reported chunk progress, and ``agent_message_delta`` never fired, so the
        # CLI and console token streams both rendered nothing.
        "streaming": True,
        # Disable LiteLLM's built-in retries — the ToolUseLoop manages
        # retries itself with per-attempt timeouts and visible retry events.
        "request_timeout": float(
            get_config_value("llm", "request_timeout", default=DEFAULT_REQUEST_TIMEOUT)
            or DEFAULT_REQUEST_TIMEOUT
        ),
        "max_retries": 0,
    }
    if openai_api_base:
        kwargs["api_base"] = openai_api_base
    if api_key:
        kwargs["api_key"] = api_key
    if model_kwargs:
        kwargs["model_kwargs"] = model_kwargs

    chat = ChatLiteLLM(**kwargs)
    # Rescue token usage that litellm strands in ``_hidden_params`` so it
    # reaches ``usage_metadata`` (see ``_UsageNormalizingLiteLLM``). ChatLiteLLM
    # sets ``client`` to the ``litellm`` module at construction; wrap that. Skip
    # when absent (test doubles that stub ChatLiteLLM never set a client).
    inner_client = getattr(chat, "client", None)
    if inner_client is not None:
        chat.client = _UsageNormalizingLiteLLM(inner_client)
    return cast(ChatModel, chat)

model_prefers_structured_patch(model_name: str | None) -> bool

Return True if the model works better with the per-file structured_patch tool.

GPT-5-class, o3/o4, and Codex models use structured JSON tool calls that map naturally to file_edit_tool (structured_patch). Claude and Gemini are trained on diff/patch text formats and work better with aider_edit_block_tool (search_replace_block).

Precedence (the model→tool-variant map is now controllable data): 1. llm.structured_patch_models config allowlist (runtime override layer). 2. The operator-tunable prompts/model_variants.yaml map, loaded through ModelVariantRegistry — this is where the built-in defaults now live (gpt-5/o3/o4/codex/gpt-4), so they are editable without touching code. Its conservative defaults.edit_tool (search_replace_block) is the sane built-in fallback when no profile matches.

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
def model_prefers_structured_patch(model_name: str | None) -> bool:
    """Return True if the model works better with the per-file structured_patch tool.

    GPT-5-class, o3/o4, and Codex models use structured JSON tool calls that
    map naturally to ``file_edit_tool`` (structured_patch).  Claude and Gemini
    are trained on diff/patch text formats and work better with
    ``aider_edit_block_tool`` (search_replace_block).

    Precedence (the model→tool-variant map is now controllable data):
    1. ``llm.structured_patch_models`` config allowlist (runtime override layer).
    2. The operator-tunable ``prompts/model_variants.yaml`` map, loaded through
       ``ModelVariantRegistry`` — this is where the built-in defaults now live
       (gpt-5/o3/o4/codex/gpt-4), so they are editable without touching code.
       Its conservative ``defaults.edit_tool`` (``search_replace_block``) is the
       sane built-in fallback when no profile matches.
    """
    if not model_name:
        return False
    normalized = _strip_provider(model_name)
    raw = model_name.lower()
    allowlist = _normalize_model_list(
        get_config_value("llm", "structured_patch_models", default=[])
    )
    if _matches_model_list(raw, allowlist) or _matches_model_list(normalized, allowlist):
        return True
    # Controllable data file (migrated built-in defaults; operator-editable).
    from mewbo_core.llm.model_variants import get_model_variant_registry

    return get_model_variant_registry().edit_tool_for(model_name) == "structured_patch"

model_supports_prompt_caching(model_name: str | None) -> bool

Return True when LiteLLM reports the model supports prompt caching.

Single source of truth: litellm.utils.supports_prompt_caching, which reads the bundled model_cost.json (extensible at runtime via litellm.register_model). Returns False on unknown models or any lookup error so the caller can skip caching gracefully without crashing the agent loop.

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
def model_supports_prompt_caching(model_name: str | None) -> bool:
    """Return True when LiteLLM reports the model supports prompt caching.

    Single source of truth: ``litellm.utils.supports_prompt_caching``, which
    reads the bundled ``model_cost.json`` (extensible at runtime via
    ``litellm.register_model``).  Returns False on unknown models or any
    lookup error so the caller can skip caching gracefully without crashing
    the agent loop.
    """
    if not model_name or _litellm_supports_prompt_caching is None:
        return False
    try:
        return bool(_litellm_supports_prompt_caching(_strip_provider(model_name)))
    except Exception:
        return False

model_supports_reasoning_effort(model_name: str | None) -> bool

Return True if the model is known to support reasoning_effort.

LiteLLM translates reasoning_effort per-provider: - Claude → output_config.effort - Gemini → thinking budget_tokens or thinking_level - OpenAI (o3/gpt-5) → native reasoning_effort

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
def model_supports_reasoning_effort(model_name: str | None) -> bool:
    """Return True if the model is known to support reasoning_effort.

    LiteLLM translates reasoning_effort per-provider:
    - Claude → output_config.effort
    - Gemini → thinking budget_tokens or thinking_level
    - OpenAI (o3/gpt-5) → native reasoning_effort
    """
    if not model_name:
        return False
    raw = model_name.lower()
    normalized = _strip_provider(model_name)
    allowlist = _normalize_model_list(
        get_config_value("llm", "reasoning_effort_models", default=[])
    )
    if _matches_model_list(raw, allowlist) or _matches_model_list(normalized, allowlist):
        return True
    return (
        normalized.startswith("gpt-5")
        or normalized.startswith("o3")
        or "claude" in normalized
        or "gemini" in normalized
    )

register_proxy_model_capabilities(api_base: str | None, api_key: str | None, *, timeout: float = 5.0) -> int

Pull proxy /v1/model/info and register advertised models.

Hydrates LiteLLM's local model_cost map with the routes the proxy operator defined.

This is the bridge that lets litellm.utils.supports_prompt_caching (and every other supports_* helper) report accurately for proxy-fronted custom model names — the SDK only consults its bundled model_cost.json by default, which doesn't know about routes the proxy operator defined.

Idempotent per process: each distinct api_base is fetched at most once. Failures are logged and swallowed — Stage 1's per-model gate just stays conservative, never crashes.

Returns the number of models newly registered (0 if cached, no-op, or error).

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
def register_proxy_model_capabilities(
    api_base: str | None,
    api_key: str | None,
    *,
    timeout: float = 5.0,
) -> int:
    """Pull proxy ``/v1/model/info`` and register advertised models.

    Hydrates LiteLLM's local ``model_cost`` map with the routes the proxy
    operator defined.

    This is the bridge that lets ``litellm.utils.supports_prompt_caching`` (and
    every other ``supports_*`` helper) report accurately for proxy-fronted
    custom model names — the SDK only consults its bundled ``model_cost.json``
    by default, which doesn't know about routes the proxy operator defined.

    Idempotent per process: each distinct ``api_base`` is fetched at most once.
    Failures are logged and swallowed — Stage 1's per-model gate just stays
    conservative, never crashes.

    Returns the number of models newly registered (0 if cached, no-op, or
    error).
    """
    if not api_base:
        return 0
    base = api_base.rstrip("/")
    if base in _REGISTERED_PROXY_BASES:
        return 0
    _REGISTERED_PROXY_BASES.add(base)
    try:
        import httpx
        import litellm

        url = base + "/model/info"
        headers = {"Authorization": f"Bearer {api_key}"} if api_key else {}
        resp = httpx.get(url, headers=headers, timeout=timeout)
        resp.raise_for_status()
        entries = (resp.json() or {}).get("data") or []
    except Exception as exc:
        _logger.info(
            "Skipping proxy model_info registration for %s: %s",
            base,
            exc,
        )
        return 0

    registered = 0
    for entry in entries:
        if not isinstance(entry, dict):
            continue
        name = entry.get("model_name")
        info = entry.get("model_info") or {}
        if not isinstance(name, str) or not isinstance(info, dict) or not info:
            continue
        # FORCED, never defaulted. Every model here is reached through the proxy
        # as an OpenAI-compatible endpoint, which is why the key is ``openai/``;
        # the upstream's own provider label describes how the PROXY reaches the
        # model and is not ours to carry. Letting it through was a silent break:
        # the proxy reports ``anthropic`` for a Claude-backed route, so the entry
        # was stored under an ``openai/`` key while declaring a different
        # provider, and the lookup — which infers the provider from that very
        # prefix — rejected its own record as a mismatch.
        info["litellm_provider"] = "openai"
        info.setdefault("mode", "chat")
        try:
            litellm.register_model({f"openai/{name}": info})
            registered += 1
        except Exception as exc:  # pragma: no cover - litellm guard
            _logger.debug("register_model failed for %s: %s", name, exc)
    if registered:
        # The budget layer memoizes catalogue answers, INCLUDING misses, and the
        # API answers usage reads without ever constructing a client — so a poll
        # arriving before this hydration would pin the pre-hydration window for
        # the life of the worker. Imported at the call site: ``session`` sits
        # above ``llm`` in the package graph, and this is the only edge.
        from mewbo_core.session.token_budget import forget_cached_context_windows

        forget_cached_context_windows()
    return registered

resolve_reasoning_effort(model_name: str | None) -> str | None

Resolve the reasoning effort for a model.

Returns the configured value if set, otherwise None (let the provider decide). Only returns a value when the model supports the parameter and a value is explicitly configured. LiteLLM translates this per-provider: - Claude: output_config.effort - Gemini: thinking budget_tokens or thinking_level - OpenAI: native reasoning_effort

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
def resolve_reasoning_effort(model_name: str | None) -> str | None:
    """Resolve the reasoning effort for a model.

    Returns the configured value if set, otherwise ``None`` (let the
    provider decide).  Only returns a value when the model supports
    the parameter *and* a value is explicitly configured.
    LiteLLM translates this per-provider:
    - Claude: output_config.effort
    - Gemini: thinking budget_tokens or thinking_level
    - OpenAI: native reasoning_effort
    """
    if not model_supports_reasoning_effort(model_name):
        return None
    configured = get_config_value("llm", "reasoning_effort", default="")
    if isinstance(configured, str) and configured.strip():
        return configured.strip().lower()
    env = os.environ.get("MEWBO_REASONING_EFFORT", "").strip().lower()
    if env:
        return env
    return None

response_text(response: Any) -> str

Visible assistant text from an LLM response, whatever shape it arrives in.

A cross-model normalization, which is why it lives at this seam rather than at each caller (see this package's CLAUDE.md: format differences are fixed here, never detected upstream). Three shapes reach us:

  • a plain str — most models;
  • a list of {"type": "text", "text": ...} blocks — Anthropic-style;
  • a list mixing thinking/reasoning blocks with the answer as a BARE STRING element — what a reasoning model returns through the proxy.

Callers used to take the FIRST type == "text" dict, which yields "" for that third shape. The failure is silent — no exception, no log, just an empty answer — so a session title and a compaction summary each simply stopped being produced the moment a reasoning model became the default. Reasoning blocks are deliberately dropped: they are the model's scratchpad, not its answer.

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
def response_text(response: Any) -> str:
    """Visible assistant text from an LLM response, whatever shape it arrives in.

    A cross-model normalization, which is why it lives at this seam rather than
    at each caller (see this package's CLAUDE.md: format differences are fixed
    here, never detected upstream). Three shapes reach us:

    * a plain ``str`` — most models;
    * a list of ``{"type": "text", "text": ...}`` blocks — Anthropic-style;
    * a list mixing ``thinking``/``reasoning`` blocks with the answer as a BARE
      STRING element — what a reasoning model returns through the proxy.

    Callers used to take the FIRST ``type == "text"`` dict, which yields ``""``
    for that third shape. The failure is silent — no exception, no log, just an
    empty answer — so a session title and a compaction summary each simply
    stopped being produced the moment a reasoning model became the default.
    Reasoning blocks are deliberately dropped: they are the model's scratchpad,
    not its answer.
    """
    raw = response.content if hasattr(response, "content") else response
    if isinstance(raw, str):
        return raw
    if not isinstance(raw, list):
        return str(raw)
    parts: list[str] = []
    for block in raw:
        if isinstance(block, str):
            parts.append(block)
        elif isinstance(block, dict) and block.get("type") == "text":
            text = block.get("text")
            if isinstance(text, str):
                parts.append(text)
    return "".join(parts)

sanitize_tool_schema(schema: Any) -> Any

Recursively fix JSON Schema issues that strict LLM providers reject.

Runs at the specs_to_langchain_tools funnel so every tool schema — MCP, built-in, plugin — is covered. Fixes are valid JSON Schema, safe for all providers.

Current fixes: - array without items → add "items": {} (required by OpenAI). - drop maxLength / maxItems / maxProperties, which make a grammar-constrained backend fail to start.

WHY THE UPPER BOUNDS GO. A backend that constrains decoding with a grammar (llama.cpp/Ollama, and anything else compiling JSON Schema to GBNF) expands an upper bound into that many literal repetitions, and inlines every $ref while doing it. In a RECURSIVE schema the two multiply. Measured: present_ui — whose Card.children is a oneOf over eleven component types including Card itself — made Ollama answer 400 Failed to initialize samplers: failed to parse grammar for the whole 17-tool request. Dropping these three keywords fixed it; dropping minLength, pattern, const, default, discriminator, anyOf or additionalProperties did not.

Magnitude is what bites, not presence: the same schema compiled with maxLength forced to 8, and failed again with maxItems raised to 2000. A size threshold would still be unsound, because the blow-up scales with nesting depth as well as with the bound, so a limit that is safe at one depth is fatal one level down. Dropping unconditionally is the only rule that does not need to know the shape of the schema.

Dropping is strictly PERMISSIVE and cannot truncate or reject a valid argument — it only stops advertising a ceiling. Lower bounds stay: they are small in practice and carry real intent. Callers that need the ceiling enforced still get it, because tool arguments are validated against the original schema on our side (EmitStructuredResponseTool.handle).

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
def sanitize_tool_schema(schema: Any) -> Any:
    """Recursively fix JSON Schema issues that strict LLM providers reject.

    Runs at the ``specs_to_langchain_tools`` funnel so every tool schema —
    MCP, built-in, plugin — is covered.  Fixes are valid JSON Schema, safe
    for all providers.

    Current fixes:
    - ``array`` without ``items`` → add ``"items": {}`` (required by OpenAI).
    - drop ``maxLength`` / ``maxItems`` / ``maxProperties``, which make a
      grammar-constrained backend fail to start.

    WHY THE UPPER BOUNDS GO. A backend that constrains decoding with a grammar
    (llama.cpp/Ollama, and anything else compiling JSON Schema to GBNF) expands
    an upper bound into that many literal repetitions, and inlines every
    ``$ref`` while doing it. In a RECURSIVE schema the two multiply. Measured:
    ``present_ui`` — whose ``Card.children`` is a ``oneOf`` over
    eleven component types including ``Card`` itself — made Ollama answer
    ``400 Failed to initialize samplers: failed to parse grammar`` for the whole
    17-tool request. Dropping these three keywords fixed it; dropping
    ``minLength``, ``pattern``, ``const``, ``default``, ``discriminator``,
    ``anyOf`` or ``additionalProperties`` did not.

    Magnitude is what bites, not presence: the same schema compiled with
    ``maxLength`` forced to 8, and failed again with ``maxItems`` raised to
    2000. **A size threshold would still be unsound**, because the blow-up
    scales with nesting depth as well as with the bound, so a limit that is
    safe at one depth is fatal one level down. Dropping unconditionally is the
    only rule that does not need to know the shape of the schema.

    Dropping is strictly PERMISSIVE and cannot truncate or reject a valid
    argument — it only stops advertising a ceiling. Lower bounds stay: they are
    small in practice and carry real intent. Callers that need the ceiling
    enforced still get it, because tool arguments are validated against the
    original schema on our side (``EmitStructuredResponseTool.handle``).
    """
    if not isinstance(schema, dict):
        return schema

    result: dict[str, Any] = {}
    for key, value in schema.items():
        if key in _UNBOUNDED_KEYWORDS:
            continue
        if key in ("properties", "$defs", "definitions") and isinstance(value, dict):
            result[key] = {k: sanitize_tool_schema(v) for k, v in value.items()}
        elif key in ("additionalProperties", "items") and isinstance(value, dict):
            result[key] = sanitize_tool_schema(value)
        elif key in ("anyOf", "oneOf", "allOf", "prefixItems", "items") and isinstance(value, list):
            result[key] = [sanitize_tool_schema(v) for v in value]
        else:
            result[key] = value

    # Array without items → add permissive default.
    schema_type = result.get("type")
    is_array = schema_type == "array" or (isinstance(schema_type, list) and "array" in schema_type)
    if is_array and "items" not in result:
        result["items"] = {}

    return result

specs_to_langchain_tools(specs: list[object]) -> list[dict[str, Any]]

Convert ToolSpecs to LangChain bind_tools() format.

Each spec must have tool_id, description, and metadata["schema"]. Specs without a schema are silently skipped.

Delegates to LangChain's :func:convert_to_openai_tool (Anthropic-format input) so that schema normalisation is handled by the library rather than hand-rolled here.

Source code in packages/mewbo_core/src/mewbo_core/llm/llm.py
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
def specs_to_langchain_tools(specs: list[object]) -> list[dict[str, Any]]:
    """Convert ToolSpecs to LangChain bind_tools() format.

    Each spec must have ``tool_id``, ``description``, and ``metadata["schema"]``.
    Specs without a schema are silently skipped.

    Delegates to LangChain's :func:`convert_to_openai_tool` (Anthropic-format
    input) so that schema normalisation is handled by the library rather than
    hand-rolled here.
    """
    from langchain_core.utils.function_calling import convert_to_openai_tool

    tools: list[dict[str, Any]] = []
    for spec in specs:
        if not getattr(spec, "enabled", True):
            continue
        metadata = getattr(spec, "metadata", None) or {}
        schema = metadata.get("schema")
        if not isinstance(schema, dict):
            continue
        # Anthropic-format dict: LangChain maps input_schema → parameters.
        tools.append(
            convert_to_openai_tool(
                {
                    "name": getattr(spec, "tool_id", ""),
                    "description": getattr(spec, "description", ""),
                    "input_schema": sanitize_tool_schema(schema),
                }
            )
        )
    return tools

mewbo_core.tooling.plugins

Plugin discovery, manifest parsing, marketplace reading, and install/uninstall.

All path resolution is done by the CALLER via PluginsConfig.resolve_*() methods. This module accepts resolved paths as parameters — no hardcoded ~/.mewbo/ or ~/.claude/ paths.

The only place ${CLAUDE_PLUGIN_ROOT} is resolved is substitute_plugin_vars. It is applied once per plugin at discovery time via _deep_substitute.

GitSubdirPluginSource

Bases: _GitPluginSourceBase

A plugin vendored from a subdirectory of a larger repo.

{"source": "git-subdir", "url": ..., "path": ...}. path is REQUIRED here (unlike the shared optional field on the base class) — it is the entire reason this discriminator exists as distinct from url. url resolves through the SAME shared :func:_resolve_git_url the github variant uses, not verbatim — a bare host/owner/repo shorthand must work here exactly as it does for repo, since this is the one variant most likely to be hand-written.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
class GitSubdirPluginSource(_GitPluginSourceBase):
    """A plugin vendored from a subdirectory of a larger repo.

    ``{"source": "git-subdir", "url": ..., "path": ...}``. ``path`` is
    REQUIRED here (unlike the shared optional field on the base class) — it
    is the entire reason this discriminator exists as distinct from ``url``.
    ``url`` resolves through the SAME shared :func:`_resolve_git_url` the
    ``github`` variant uses, not verbatim — a bare ``host/owner/repo``
    shorthand must work here exactly as it does for ``repo``, since this is
    the one variant most likely to be hand-written.
    """

    source: Literal["git-subdir"] = "git-subdir"
    url: str
    path: str

    def resolve_git_url(self, *, default_host: str = "github.com") -> str:
        """Resolve ``url`` through the shared host-agnostic resolver."""
        return _resolve_git_url(self.url, default_host=default_host)

resolve_git_url(*, default_host: str = 'github.com') -> str

Resolve url through the shared host-agnostic resolver.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
778
779
780
def resolve_git_url(self, *, default_host: str = "github.com") -> str:
    """Resolve ``url`` through the shared host-agnostic resolver."""
    return _resolve_git_url(self.url, default_host=default_host)

GithubPluginSource

Bases: _GitPluginSourceBase

{"source": "github", "repo": ...} — a host-agnostic repo shorthand.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
733
734
735
736
737
738
739
740
741
class GithubPluginSource(_GitPluginSourceBase):
    """``{"source": "github", "repo": ...}`` — a host-agnostic repo shorthand."""

    source: Literal["github"] = "github"
    repo: str

    def resolve_git_url(self, *, default_host: str = "github.com") -> str:
        """Resolve ``repo`` through the shared host-agnostic resolver."""
        return _resolve_git_url(self.repo, default_host=default_host)

resolve_git_url(*, default_host: str = 'github.com') -> str

Resolve repo through the shared host-agnostic resolver.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
739
740
741
def resolve_git_url(self, *, default_host: str = "github.com") -> str:
    """Resolve ``repo`` through the shared host-agnostic resolver."""
    return _resolve_git_url(self.repo, default_host=default_host)

PluginComponents dataclass

Fan-out of a single plugin's contributions to existing registries.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
195
196
197
198
199
200
201
202
203
204
205
206
207
208
@dataclass(frozen=True)
class PluginComponents:
    """Fan-out of a single plugin's contributions to existing registries."""

    manifest: PluginManifest | None
    skill_dirs: list[str] = field(default_factory=list)  # paths to skills/<name>/ dirs
    command_files: list[str] = field(default_factory=list)  # paths to commands/*.md files
    agent_files: list[str] = field(default_factory=list)  # paths to agents/*.md files
    mcp_config: dict | None = None  # parsed .mcp.json content (vars substituted)
    hooks_config: dict | None = None  # parsed hooks/hooks.json content
    # Plugin-contributed session tools (Mewbo extension to the Claude Code
    # plugin format). Each entry is a ``{tool_id, module, class}`` record used
    # by :class:`SessionToolRegistry` to instantiate per-session handlers.
    session_tool_entries: list[dict] = field(default_factory=list)

PluginFanOut dataclass

Aggregated components from all enabled plugins, ready for registry injection.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
@dataclass
class PluginFanOut:
    """Aggregated components from all enabled plugins, ready for registry injection."""

    components: list[PluginComponents]
    skill_dirs: list[str]
    command_files: list[str]
    agent_files: list[str]
    mcp_servers: dict[str, dict]
    hooks_configs: list[tuple[dict, str]]  # (hooks_json, plugin_root)
    session_tool_entries: list[dict] = field(default_factory=list)

PluginManifest dataclass

Parsed .claude-plugin/plugin.json manifest.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
@dataclass(frozen=True)
class PluginManifest:
    """Parsed .claude-plugin/plugin.json manifest."""

    name: str
    display_name: str | None = None
    description: str = ""
    version: str = ""
    author: str = ""
    install_path: str = ""  # absolute path to plugin root dir
    marketplace: str = ""  # which marketplace it came from
    scope: str = "user"  # "user" | "project" | "local"
    # Plugin-level capability gate. When non-empty, the values fan out onto
    # every agent / skill / command this plugin contributes — authors can
    # also repeat (or tighten) per file in the contribution's frontmatter.
    requires_capabilities: tuple[str, ...] = field(default_factory=tuple)

PluginSource

Bases: RootModel[PluginSourceUnion]

Parse seam for a plugin manifest's dict-shaped source field.

Wraps the discriminated union in a RootModel so a mode="before" validator can run BEFORE Pydantic reads the source discriminator: a bare {"repo": ...} entry with no source key has no tag for the union to dispatch on until this stamps one. Installed manifests carry that shape and must keep resolving as github.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
class PluginSource(RootModel[PluginSourceUnion]):
    """Parse seam for a plugin manifest's dict-shaped ``source`` field.

    Wraps the discriminated union in a ``RootModel`` so a ``mode="before"``
    validator can run BEFORE Pydantic reads the ``source`` discriminator: a
    bare ``{"repo": ...}`` entry with no ``source`` key has no tag for the union
    to dispatch on until this stamps one. Installed manifests carry that shape
    and must keep resolving as ``github``.
    """

    @model_validator(mode="before")
    @classmethod
    def _stamp_legacy_default(cls, data: object) -> object:
        if isinstance(data, Mapping) and "source" not in data:
            if "repo" in data:
                return {**data, "source": "github"}
            if "url" in data:
                return {**data, "source": "url"}
        return data

    @classmethod
    def parse(
        cls, data: Mapping[str, object]
    ) -> GithubPluginSource | UrlPluginSource | GitSubdirPluginSource:
        """Parse a raw ``source`` dict into its concrete variant."""
        return cls.model_validate(data).root

parse(data: Mapping[str, object]) -> GithubPluginSource | UrlPluginSource | GitSubdirPluginSource classmethod

Parse a raw source dict into its concrete variant.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
809
810
811
812
813
814
@classmethod
def parse(
    cls, data: Mapping[str, object]
) -> GithubPluginSource | UrlPluginSource | GitSubdirPluginSource:
    """Parse a raw ``source`` dict into its concrete variant."""
    return cls.model_validate(data).root

UrlPluginSource

Bases: _GitPluginSourceBase

{"source": "url", "url": ...} — used verbatim, never resolved.

Trap: the url type is already expected to be a full, clonable reference — routing it through :func:_resolve_git_url (as :class:GitSubdirPluginSource correctly does) would be harmless for a real URL but silently wrong for anything else callers pass here as-is (e.g. a local path in a test fixture).

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
class UrlPluginSource(_GitPluginSourceBase):
    """``{"source": "url", "url": ...}`` — used verbatim, never resolved.

    Trap: the ``url`` type is already expected to be a full, clonable
    reference — routing it through :func:`_resolve_git_url` (as
    :class:`GitSubdirPluginSource` correctly does) would be harmless for a
    real URL but silently wrong for anything else callers pass here as-is
    (e.g. a local path in a test fixture).
    """

    source: Literal["url"] = "url"
    url: str

    def resolve_git_url(self, *, default_host: str = "github.com") -> str:
        """Return ``url`` unchanged — never resolved, unlike ``repo``."""
        return self.url

resolve_git_url(*, default_host: str = 'github.com') -> str

Return url unchanged — never resolved, unlike repo.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
757
758
759
def resolve_git_url(self, *, default_host: str = "github.com") -> str:
    """Return ``url`` unchanged — never resolved, unlike ``repo``."""
    return self.url

discover_builtin_plugins(root: Path | str) -> list[PluginComponents]

Discover first-party plugins shipped inside the core package.

Walks immediate subdirectories of root (no registry indirection) and returns :class:PluginComponents for each directory that contains a .claude-plugin/plugin.json. Built-in plugins bypass installed_plugins.json because they ship with the core wheel — their presence is a property of the installation, not user action.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
def discover_builtin_plugins(root: Path | str) -> list[PluginComponents]:
    """Discover first-party plugins shipped inside the core package.

    Walks immediate subdirectories of *root* (no registry indirection)
    and returns :class:`PluginComponents` for each directory that
    contains a ``.claude-plugin/plugin.json``. Built-in plugins bypass
    ``installed_plugins.json`` because they ship with the core wheel —
    their presence is a property of the installation, not user action.
    """
    root = Path(root)
    if not root.is_dir():
        return []
    result: list[PluginComponents] = []
    for child in sorted(root.iterdir()):
        if not child.is_dir():
            continue
        manifest_path = child / ".claude-plugin" / "plugin.json"
        if not manifest_path.is_file():
            continue
        components = discover_plugin_components(child)
        # Mark scope = "built-in" so the origin is visible in /plugins listings.
        if components.manifest is not None:
            patched = replace(
                components.manifest, scope="built-in", marketplace="built-in"
            )
            components = replace(components, manifest=patched)
        result.append(components)
    return result

discover_installed_plugins(registry_paths: list[Path | str], *, enabled: list[str] | None = None) -> list[PluginComponents]

Read installed plugin registries and return discovered components.

Registry format::

{
    "version": 2,
    "plugins": {
        "name@marketplace": [
            {"scope": "user", "installPath": "/abs/path", "version": "1.0.0"}
        ]
    }
}
  • Takes the FIRST entry in the list for each key (highest-priority scope).
  • Filters by the name part (before @) when enabled is non-empty.
  • Deduplicates by name — first registry_path wins.
Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
def discover_installed_plugins(
    registry_paths: list[Path | str],
    *,
    enabled: list[str] | None = None,
) -> list[PluginComponents]:
    """Read installed plugin registries and return discovered components.

    Registry format::

        {
            "version": 2,
            "plugins": {
                "name@marketplace": [
                    {"scope": "user", "installPath": "/abs/path", "version": "1.0.0"}
                ]
            }
        }

    - Takes the FIRST entry in the list for each key (highest-priority scope).
    - Filters by the name part (before ``@``) when *enabled* is non-empty.
    - Deduplicates by name — first registry_path wins.
    """
    seen: set[str] = set()
    result: list[PluginComponents] = []

    for raw_path in registry_paths:
        registry_path = Path(raw_path)
        if not registry_path.is_file():
            continue
        try:
            registry: dict = json.loads(registry_path.read_text(encoding="utf-8"))
        except (OSError, json.JSONDecodeError) as exc:
            logging.warning("Failed to read registry {}: {}", registry_path, exc)
            continue

        plugins_map: dict[str, list[dict]] = registry.get("plugins", {})
        for key, entries in plugins_map.items():
            if not entries:
                continue

            # Split "name@marketplace"
            if "@" in key:
                plugin_name, marketplace = key.split("@", 1)
            else:
                plugin_name, marketplace = key, ""

            # Dedup — first seen wins
            if plugin_name in seen:
                continue

            if enabled:
                if plugin_name not in enabled:
                    continue

            # First entry = highest priority scope
            entry = entries[0]
            install_path = entry.get("installPath", "")
            scope = entry.get("scope", "user")

            if not install_path or not Path(install_path).is_dir():
                logging.debug(
                    "Plugin {} installPath '{}' does not exist — skipping",
                    plugin_name,
                    install_path,
                )
                continue

            # Skip entries with no manifest file — these are incomplete/stale
            # cache entries (e.g. repos that lack .claude-plugin/plugin.json).
            manifest_file = Path(install_path) / ".claude-plugin" / "plugin.json"
            if not manifest_file.is_file():
                stale_key = (plugin_name, str(manifest_file))
                if stale_key not in _LOGGED_STALE_CACHE:
                    _LOGGED_STALE_CACHE.add(stale_key)
                    logging.trace(
                        "Plugin {} has no manifest at {} — skipping stale cache entry",
                        plugin_name,
                        manifest_file,
                    )
                continue

            seen.add(plugin_name)
            components = discover_plugin_components(Path(install_path))

            # Overlay marketplace and scope onto the manifest
            if components.manifest is not None:
                patched_manifest = replace(
                    components.manifest, marketplace=marketplace, scope=scope
                )
                components = replace(components, manifest=patched_manifest)

            result.append(components)

    return result

discover_marketplace_plugins(marketplace_dirs: list[Path | str]) -> list[dict]

Read marketplace.json from each directory and return a flat list of plugin dicts.

Each dict contains: name, description, category, marketplace, installed (always False — caller can enrich).

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
def discover_marketplace_plugins(marketplace_dirs: list[Path | str]) -> list[dict]:
    """Read marketplace.json from each directory and return a flat list of plugin dicts.

    Each dict contains: ``name``, ``description``, ``category``, ``marketplace``,
    ``installed`` (always ``False`` — caller can enrich).
    """
    result: list[dict] = []
    for raw_dir in marketplace_dirs:
        mp_dir = Path(raw_dir)
        marketplace_json = mp_dir / ".claude-plugin" / "marketplace.json"
        if not marketplace_json.is_file():
            logging.debug("No marketplace.json found in {}", mp_dir)
            continue
        try:
            data: dict = json.loads(marketplace_json.read_text(encoding="utf-8"))
        except (OSError, json.JSONDecodeError) as exc:
            logging.warning("Failed to parse marketplace.json in {}: {}", mp_dir, exc)
            continue

        marketplace_name = data.get("name", str(mp_dir.name))
        for plugin in data.get("plugins", []):
            result.append(
                {
                    "name": plugin.get("name", ""),
                    "description": plugin.get("description", ""),
                    "category": plugin.get("category", ""),
                    "marketplace": marketplace_name,
                    "installed": False,
                }
            )

    return result

discover_plugin_components(plugin_dir: Path | str) -> PluginComponents

Scan plugin_dir for all plugin contributions.

Returns a :class:PluginComponents instance. If plugin.json is absent the manifest will be None but other components are still discovered.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
def discover_plugin_components(plugin_dir: Path | str) -> PluginComponents:
    """Scan *plugin_dir* for all plugin contributions.

    Returns a :class:`PluginComponents` instance.  If ``plugin.json`` is absent
    the manifest will be ``None`` but other components are still discovered.
    """
    plugin_dir = Path(plugin_dir)
    plugin_root = str(plugin_dir)

    # Parse plugin.json ONCE and reuse the raw dict for both the manifest
    # and the session_tools entries — no double file read.
    manifest_data = _read_manifest_data(plugin_dir)
    manifest = (
        _manifest_from_data(manifest_data, plugin_dir)
        if manifest_data is not None
        else None
    )

    # --- skills/
    # Return the parent skills/ directory so consumers can call load_extra_dir() directly.
    skill_dirs: list[str] = []
    skills_dir = plugin_dir / "skills"
    if skills_dir.is_dir():
        has_valid_skill = any(
            candidate.is_dir() and (candidate / "SKILL.md").is_file()
            for candidate in skills_dir.iterdir()
        )
        if has_valid_skill:
            skill_dirs.append(str(skills_dir))

    # --- commands/
    command_files: list[str] = []
    commands_dir = plugin_dir / "commands"
    if commands_dir.is_dir():
        command_files = sorted(str(p) for p in commands_dir.glob("*.md") if p.is_file())

    # --- agents/
    agent_files: list[str] = []
    agents_dir = plugin_dir / "agents"
    if agents_dir.is_dir():
        agent_files = sorted(str(p) for p in agents_dir.glob("*.md") if p.is_file())

    # --- .mcp.json
    mcp_config: dict | None = None
    mcp_path = plugin_dir / ".mcp.json"
    if mcp_path.is_file():
        try:
            raw: dict = json.loads(mcp_path.read_text(encoding="utf-8"))
            # Handle both Mewbo native {"servers": {...}} and Claude Code style
            # {"server-name": {...}} — if no top-level "servers" key, whole dict is servers
            substituted: dict = _deep_substitute(raw, plugin_root)
            mcp_config = substituted
        except (OSError, json.JSONDecodeError) as exc:
            logging.warning("Failed to parse .mcp.json in {}: {}", plugin_dir, exc)

    # --- hooks/hooks.json (no variable substitution — merger does it later)
    hooks_config: dict | None = None
    hooks_path = plugin_dir / "hooks" / "hooks.json"
    if hooks_path.is_file():
        try:
            hooks_config = json.loads(hooks_path.read_text(encoding="utf-8"))
        except (OSError, json.JSONDecodeError) as exc:
            logging.warning("Failed to parse hooks/hooks.json in {}: {}", plugin_dir, exc)

    # --- session_tools (Mewbo extension to the Claude Code plugin format).
    # Each entry is a ``{tool_id, module, class}`` record the core imports and
    # turns into a factory via :class:`SessionToolRegistry`. Read from the
    # same manifest_data parsed above — do not reopen plugin.json.
    session_tool_entries: list[dict] = []
    if manifest_data is not None:
        raw_entries = manifest_data.get("session_tools", [])
        if isinstance(raw_entries, list):
            session_tool_entries = [e for e in raw_entries if isinstance(e, dict)]

    return PluginComponents(
        manifest=manifest,
        skill_dirs=skill_dirs,
        command_files=command_files,
        agent_files=agent_files,
        mcp_config=mcp_config,
        hooks_config=hooks_config,
        session_tool_entries=session_tool_entries,
    )

install_plugin(name: str, marketplace: str, *, marketplace_dirs: list[Path], install_base: Path) -> PluginManifest

Install a plugin from a marketplace into install_base.

  • Finds the plugin entry in marketplace.json by searching marketplace_dirs.
  • For local sources (string starting with ./): copies the directory.
  • For git sources: clones the repo.
  • Updates install_base/installed_plugins.json.
  • Returns the parsed :class:PluginManifest.
Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
873
874
875
876
877
878
879
880
881
882
883
884
885
886
887
888
889
890
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
939
940
941
942
943
944
945
946
947
948
949
950
951
952
953
954
955
956
957
958
959
def install_plugin(
    name: str,
    marketplace: str,
    *,
    marketplace_dirs: list[Path],
    install_base: Path,
) -> PluginManifest:
    """Install a plugin from a marketplace into *install_base*.

    - Finds the plugin entry in marketplace.json by searching *marketplace_dirs*.
    - For local sources (string starting with ``./``): copies the directory.
    - For git sources: clones the repo.
    - Updates ``install_base/installed_plugins.json``.
    - Returns the parsed :class:`PluginManifest`.
    """
    plugin_entry: dict | None = None
    source_mp_dir: Path | None = None

    for mp_dir in marketplace_dirs:
        marketplace_json = mp_dir / ".claude-plugin" / "marketplace.json"
        if not marketplace_json.is_file():
            continue
        try:
            data: dict = json.loads(marketplace_json.read_text(encoding="utf-8"))
        except (OSError, json.JSONDecodeError):
            continue
        mp_name = data.get("name", "")
        if mp_name != marketplace:
            continue
        for entry in data.get("plugins", []):
            if entry.get("name") == name:
                plugin_entry = entry
                source_mp_dir = mp_dir
                break
        if plugin_entry is not None:
            break

    if plugin_entry is None:
        raise ValueError(f"Plugin '{name}' not found in marketplace '{marketplace}'")

    source = plugin_entry.get("source", "")
    version = plugin_entry.get("version", "latest")
    cache_dir = (
        install_base
        / "cache"
        / _sanitize_path_component(marketplace)
        / _sanitize_path_component(name)
        / _sanitize_path_component(version)
    )
    cache_dir.mkdir(parents=True, exist_ok=True)

    if isinstance(source, str) and source.startswith("./"):
        # Local path relative to the marketplace directory
        if source_mp_dir is None:
            raise ValueError(f"Plugin '{name}' not found in any marketplace")
        src_path = (source_mp_dir / source).resolve()
        if not src_path.is_relative_to(source_mp_dir.resolve()):
            raise ValueError(f"Plugin source path escapes marketplace directory: {source}")
        if cache_dir.exists():
            shutil.rmtree(cache_dir)
        shutil.copytree(src_path, cache_dir)
    elif isinstance(source, dict):
        try:
            plugin_source = PluginSource.parse(source)
        except ValidationError as exc:
            is_unsupported, kind = _unsupported_source_kind(exc)
            if not is_unsupported:
                # A recognized variant that is merely incomplete (e.g. a
                # git-subdir missing 'path') — let Pydantic's own error
                # through; it names the missing field.
                raise
            raise ValueError(
                f"Plugin '{name}': unsupported source type {kind!r}. "
                f"Mewbo supports: local ('./path'), github, url, git-subdir. "
                f"'npm' is a Claude Code source type Mewbo does not implement. "
                f"Source: {source!r}"
            ) from exc

        git_url = plugin_source.resolve_git_url()
        subdir = (plugin_source.path or "").strip("/")
        pinned_sha = (plugin_source.sha or "").strip()
        ref = (plugin_source.ref or "").strip()
        if (cache_dir / ".git").exists():
            logging.info("Plugin {} already cloned, skipping", name)
        elif subdir:
            _clone_git_subdir(git_url, cache_dir, subdir=subdir, ref=ref, sha=pinned_sha, name=name)
        else:
            subprocess.run(
                ["git", "clone", git_url, str(cache_dir)],
                check=True,
                capture_output=True,
            )
            checkout_ref = pinned_sha or ref  # sha wins when both are set
            if checkout_ref:
                _git_checkout(cache_dir, checkout_ref, name=name)
    else:
        raise ValueError(f"Unsupported plugin source for '{name}': {source!r}")

    # Update registry
    registry_path = install_base / "installed_plugins.json"
    registry: dict = {"version": 2, "plugins": {}}
    if registry_path.is_file():
        try:
            registry = json.loads(registry_path.read_text(encoding="utf-8"))
        except (OSError, json.JSONDecodeError):
            pass

    registry.setdefault("plugins", {})
    key = f"{name}@{marketplace}"
    registry["plugins"][key] = [
        {
            "scope": "user",
            "installPath": str(cache_dir),
            "version": version,
        }
    ]
    registry_path.write_text(json.dumps(registry, indent=2), encoding="utf-8")

    manifest = parse_plugin_manifest(cache_dir)
    if manifest is None:
        raise RuntimeError(f"Installed plugin '{name}' is missing a valid plugin.json")
    return manifest

load_all_plugin_components() -> PluginFanOut

Discover all enabled plugins and aggregate their components.

Uses PluginsConfig from the live config for all path resolution. Returns a :class:PluginFanOut that callers can inject into their registries. This is the single point of truth for "what do plugins contribute?" — used by both Orchestrator.__init__ and the API endpoints so they stay in sync.

Results are cached and only recomputed when the installed-plugins registry file changes (mtime comparison).

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
def load_all_plugin_components() -> PluginFanOut:
    """Discover all enabled plugins and aggregate their components.

    Uses ``PluginsConfig`` from the live config for all path resolution.
    Returns a :class:`PluginFanOut` that callers can inject into their
    registries.  This is the single point of truth for "what do plugins
    contribute?" — used by both ``Orchestrator.__init__`` and the API
    endpoints so they stay in sync.

    Results are cached and only recomputed when the installed-plugins
    registry file changes (mtime comparison).
    """
    global _fanout_cache, _fanout_cache_mtime

    from mewbo_core.config import get_config

    cfg = get_config().plugins
    if not cfg.enabled:
        return PluginFanOut([], [], [], [], {}, [], [])

    # Hydrate CLAUDE_PLUGIN_ROOT from config so shell commands in agent
    # templates can rely on it. setdefault lets .env / the host
    # environment override it without the config silently clobbering the value.
    os.environ.setdefault("CLAUDE_PLUGIN_ROOT", str(cfg.resolve_install_dir()))

    # Check cache freshness against registry file mtime
    registry_paths = cfg.resolve_registry_paths()
    max_mtime = 0.0
    for rp in registry_paths:
        try:
            max_mtime = max(max_mtime, Path(rp).stat().st_mtime)
        except OSError:
            pass

    # Include each built-in plugin root's mtime so the cache refreshes when a
    # developer edits a built-in plugin during local dev. Lookups are cheap
    # (one stat per root) and keep editable installs responsive.
    builtin_roots = _all_builtin_roots()
    for root in builtin_roots:
        try:
            max_mtime = max(max_mtime, root.stat().st_mtime)
        except OSError:
            pass

    if _fanout_cache is not None and max_mtime <= _fanout_cache_mtime:
        return _fanout_cache

    # Built-in components come first so they win the dedup race inside the
    # fan-out loop below (same-name plugin from a marketplace loses). The core
    # wheel's root leads; optional capability libraries (mewbo_graph) follow.
    builtin_components = [
        pc for root in builtin_roots for pc in discover_builtin_plugins(root)
    ]
    all_components = [
        *builtin_components,
        *discover_installed_plugins(
            registry_paths=cfg.resolve_registry_paths(),
            enabled=cfg.enabled_plugins or None,
        ),
    ]

    skill_dirs: list[str] = []
    command_files: list[str] = []
    agent_files: list[str] = []
    mcp_servers: dict[str, dict] = {}
    hooks_configs: list[tuple[dict, str]] = []
    session_tool_entries: list[dict] = []

    for pc in all_components:
        if pc.manifest is None:
            continue
        skill_dirs.extend(pc.skill_dirs)
        command_files.extend(pc.command_files)
        agent_files.extend(pc.agent_files)
        session_tool_entries.extend(pc.session_tool_entries)
        if pc.mcp_config:
            # Normalize Claude Code "mcpServers" → "servers"
            raw = pc.mcp_config
            if "mcpServers" in raw and "servers" not in raw:
                raw = {**raw, "servers": raw["mcpServers"]}
            servers = raw.get("servers", raw)
            if isinstance(servers, dict):
                for srv_name, srv_cfg in servers.items():
                    if srv_name in ("servers", "mcpServers"):
                        continue
                    if isinstance(srv_cfg, dict):
                        mcp_servers.setdefault(srv_name, srv_cfg)
        if pc.hooks_config:
            hooks_configs.append((pc.hooks_config, pc.manifest.install_path))

    result = PluginFanOut(
        components=all_components,
        skill_dirs=skill_dirs,
        command_files=command_files,
        agent_files=agent_files,
        mcp_servers=mcp_servers,
        hooks_configs=hooks_configs,
        session_tool_entries=session_tool_entries,
    )
    _fanout_cache = result
    _fanout_cache_mtime = max_mtime
    return result

marketplace_dir_name(entry: str, *, default_host: str = 'github.com') -> str

Stable, collision-free local cache-dir name for a marketplace entry.

Derived from the resolved git URL's host + path (scheme, user@, and a trailing .git stripped) so entries that share a leaf name but differ in host or owner never collide. For example anthropics/claude-plugins-official → github.com-anthropics-claude-plugins-official.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
156
157
158
159
160
161
162
163
164
165
166
167
168
169
def marketplace_dir_name(entry: str, *, default_host: str = "github.com") -> str:
    """Stable, collision-free local cache-dir name for a marketplace entry.

    Derived from the *resolved* git URL's host + path (scheme, ``user@``, and a
    trailing ``.git`` stripped) so entries that share a leaf name but differ in
    host or owner never collide. For example
    ``anthropics/claude-plugins-official`` →
    ``github.com-anthropics-claude-plugins-official``.
    """
    url = _resolve_git_url(entry, default_host=default_host)
    name = _URL_SCHEME_RE.sub("", url)  # drop scheme
    name = name.split("@", 1)[-1]  # drop scp/ssh user@
    name = name.replace(":", "/")  # scp host:path and host:port → path separator
    return _sanitize_path_component(name.removesuffix(".git"))

parse_plugin_manifest(plugin_dir: Path | str) -> PluginManifest | None

Parse .claude-plugin/plugin.json from plugin_dir.

Returns None on any error (missing file, bad JSON, missing name).

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
286
287
288
289
290
291
292
293
294
295
def parse_plugin_manifest(plugin_dir: Path | str) -> PluginManifest | None:
    """Parse ``.claude-plugin/plugin.json`` from *plugin_dir*.

    Returns ``None`` on any error (missing file, bad JSON, missing name).
    """
    plugin_dir = Path(plugin_dir)
    data = _read_manifest_data(plugin_dir)
    if data is None:
        return None
    return _manifest_from_data(data, plugin_dir)

register_builtin_root(root: Path | str) -> None

Register an extra built-in plugin root (idempotent; down-only push).

A capability library above core in the DAG calls this on import so its bundled plugin suites are discovered by :func:load_all_plugin_components alongside the core wheel's own builtin_plugins/ — without core ever importing the library. Invalidates the fan-out cache so a freshly registered root is picked up on the next load.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
507
508
509
510
511
512
513
514
515
516
517
518
519
520
def register_builtin_root(root: Path | str) -> None:
    """Register an extra built-in plugin root (idempotent; down-only push).

    A capability library above core in the DAG calls this on import so its
    bundled plugin suites are discovered by :func:`load_all_plugin_components`
    alongside the core wheel's own ``builtin_plugins/`` — without core ever
    importing the library. Invalidates the fan-out cache so a freshly
    registered root is picked up on the next load.
    """
    global _fanout_cache
    path = Path(root)
    if path not in _EXTRA_BUILTIN_ROOTS:
        _EXTRA_BUILTIN_ROOTS.append(path)
        _fanout_cache = None  # force re-discovery to include the new root

substitute_plugin_vars(text: str, plugin_root: str) -> str

Replace ${CLAUDE_PLUGIN_ROOT} with plugin_root in text.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
216
217
218
def substitute_plugin_vars(text: str, plugin_root: str) -> str:
    """Replace ``${CLAUDE_PLUGIN_ROOT}`` with *plugin_root* in *text*."""
    return text.replace("${CLAUDE_PLUGIN_ROOT}", plugin_root)

sync_marketplaces(marketplace_repos: list[str], install_base: Path, *, default_host: str = 'github.com') -> list[Path]

Ensure marketplace catalogs are cloned locally, return their directory paths.

Each entry in marketplace_repos is resolved host-agnostically via :func:_resolve_git_url — a full git URL, an host/owner/repo shorthand, or a bare owner/repo (cloned from default_host). The catalog is cloned into install_base/marketplaces/<marketplace_dir_name(entry)>/ if not already present. Cloning uses plain git clone, so it inherits the ambient git credential helpers, SSH agent, and TLS configuration — private and self-hosted catalogs work without a GitHub-specific path. Returns the list of marketplace directories (same contract as PluginsConfig.resolve_marketplace_dirs()).

This is the bridge between the plugins.marketplaces config list and the filesystem-based marketplace discovery. Without it, standalone deployments (e.g. Docker without ~/.claude) would have zero marketplaces.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
def sync_marketplaces(
    marketplace_repos: list[str],
    install_base: Path,
    *,
    default_host: str = "github.com",
) -> list[Path]:
    """Ensure marketplace catalogs are cloned locally, return their directory paths.

    Each entry in *marketplace_repos* is resolved host-agnostically via
    :func:`_resolve_git_url` — a full git URL, an ``host/owner/repo`` shorthand,
    or a bare ``owner/repo`` (cloned from *default_host*). The catalog is cloned
    into ``install_base/marketplaces/<marketplace_dir_name(entry)>/`` if not
    already present. Cloning uses plain ``git clone``, so it inherits the
    ambient git credential helpers, SSH agent, and TLS configuration — private
    and self-hosted catalogs work without a GitHub-specific path. Returns the
    list of marketplace directories (same contract as
    ``PluginsConfig.resolve_marketplace_dirs()``).

    This is the bridge between the ``plugins.marketplaces`` config list and the
    filesystem-based marketplace discovery.  Without it, standalone deployments
    (e.g. Docker without ``~/.claude``) would have zero marketplaces.
    """
    marketplaces_base = install_base / "marketplaces"
    marketplaces_base.mkdir(parents=True, exist_ok=True)

    dirs: list[Path] = []
    for entry in marketplace_repos:
        mp_dir = marketplaces_base / marketplace_dir_name(entry, default_host=default_host)

        if mp_dir.is_dir():
            # Already cloned — check for marketplace.json
            marker = mp_dir / ".claude-plugin" / "marketplace.json"
            if marker.is_file():
                dirs.append(mp_dir)
                continue
            # Dir exists but looks broken — try updating
            if (mp_dir / ".git").is_dir():
                try:
                    subprocess.run(
                        ["git", "-C", str(mp_dir), "pull", "--ff-only"],
                        check=True,
                        capture_output=True,
                        timeout=60,
                    )
                except (subprocess.CalledProcessError, subprocess.TimeoutExpired, OSError) as exc:
                    logging.warning("Failed to update marketplace {}: {}", entry, exc)
                if marker.is_file():
                    dirs.append(mp_dir)
                continue

        # Not yet cloned — shallow clone (only need marketplace.json + plugin dirs)
        git_url = _resolve_git_url(entry, default_host=default_host)
        logging.info("Cloning marketplace {} from {} into {}", entry, git_url, mp_dir)
        try:
            subprocess.run(
                ["git", "clone", "--depth=1", git_url, str(mp_dir)],
                check=True,
                capture_output=True,
                timeout=120,
            )
            dirs.append(mp_dir)
        except (subprocess.CalledProcessError, subprocess.TimeoutExpired, OSError) as exc:
            logging.warning("Failed to clone marketplace {}: {}", entry, exc)

    return dirs

uninstall_plugin(name: str, *, install_base: Path) -> bool

Remove name from the installed plugins registry and delete its cache directory.

Returns True if the plugin was found and removed, False otherwise.

Source code in packages/mewbo_core/src/mewbo_core/tooling/plugins.py
962
963
964
965
966
967
968
969
970
971
972
973
974
975
976
977
978
979
980
981
982
983
984
985
986
987
988
989
990
991
992
993
994
def uninstall_plugin(name: str, *, install_base: Path) -> bool:
    """Remove *name* from the installed plugins registry and delete its cache directory.

    Returns ``True`` if the plugin was found and removed, ``False`` otherwise.
    """
    registry_path = install_base / "installed_plugins.json"
    if not registry_path.is_file():
        return False

    try:
        registry: dict = json.loads(registry_path.read_text(encoding="utf-8"))
    except (OSError, json.JSONDecodeError) as exc:
        logging.warning("Failed to read registry for uninstall: {}", exc)
        return False

    plugins: dict = registry.get("plugins", {})
    keys_to_remove = [k for k in plugins if k.split("@")[0] == name]
    if not keys_to_remove:
        return False

    for key in keys_to_remove:
        entries = plugins.pop(key, [])
        for entry in entries:
            install_path = Path(entry.get("installPath", ""))
            if install_path.is_dir():
                try:
                    shutil.rmtree(install_path)
                    logging.info("Removed plugin cache: {}", install_path)
                except OSError as exc:
                    logging.warning("Failed to remove plugin cache {}: {}", install_path, exc)

    registry_path.write_text(json.dumps(registry, indent=2), encoding="utf-8")
    return True

mewbo_core.agents.agent_registry

Agent definition registry for the Mewbo assistant.

An agent definition is an agents/*.md file (YAML frontmatter + markdown body) loaded from a Claude Code plugin or personal/project directory. Agent definitions tell Mewbo what sub-agents are available, what tools they may use, and what their system prompt should be.

This module mirrors the structure of skills.py — same frontmatter regex, same frozen dataclass pattern, same registry pattern with no-override semantics.

AgentDef dataclass

An agent definition loaded from an agents/*.md file.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_registry.py
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
@dataclass(frozen=True)
class AgentDef:
    """An agent definition loaded from an agents/*.md file."""

    name: str
    description: str
    source_path: str  # absolute path to the .md file
    source: str  # "plugin:<plugin-name>" or "project" or "personal"
    body: str  # markdown body (becomes agent's system prompt)
    # Three-state: None = unrestricted, [] = grants nothing, non-empty = exactly
    # those. Every consumer must test it with ``is None``, never truthiness.
    allowed_tools: list[str] | None = None
    denied_tools: list[str] | None = None
    model: str | None = None  # "inherit" becomes None
    plugin_root: str = ""  # absolute path to the plugin that contributed this agent
    requires_capabilities: tuple[str, ...] = ()

AgentRegistry

Registry of agent definitions.

First-registered agent wins — later registrations with the same name are silently ignored (same semantics as the subtree skill discovery).

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_registry.py
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
class AgentRegistry:
    """Registry of agent definitions.

    First-registered agent wins — later registrations with the same name are
    silently ignored (same semantics as the subtree skill discovery).
    """

    def __init__(self) -> None:  # noqa: D107
        self._agents: dict[str, AgentDef] = {}

    def register(
        self,
        agent_def: AgentDef,
        *,
        capabilities: Iterable[str] = (),
        plugin_root: str = "",
    ) -> None:
        """Register an agent.  Does NOT override existing entries.

        When *capabilities* is non-empty, they are unioned into the
        agent's ``requires_capabilities`` before registration — the
        standard way a plugin fans its bundle-level requirements out
        over every contributed agent.

        When *plugin_root* is provided and the agent does not already
        have one, it is stamped on so downstream consumers (e.g. the
        ``${CLAUDE_PLUGIN_ROOT}`` substitution in ``spawn_agent``) can
        locate the plugin's on-disk assets.
        """
        if agent_def.name in self._agents:
            return
        agent_def = overlay_capabilities(agent_def, capabilities)
        if plugin_root and not agent_def.plugin_root:
            agent_def = replace(agent_def, plugin_root=plugin_root)
        self._agents[agent_def.name] = agent_def

    def get(
        self,
        name: str,
        session_capabilities: Iterable[str] = (),
    ) -> AgentDef | None:
        """Return the agent definition with the given name, or ``None``.

        Agents gated by ``requires_capabilities`` that the session hasn't
        advertised are treated as if they don't exist.
        """
        agent = self._agents.get(name)
        if agent is None:
            return None
        visible = filter_by_capabilities([agent], session_capabilities)
        return visible[0] if visible else None

    def list_all(self) -> list[AgentDef]:
        """Return all registered agent definitions."""
        return list(self._agents.values())

    def visible_for(self, session_capabilities: Iterable[str]) -> list[AgentDef]:
        """Return agents visible given the session's advertised capabilities."""
        return filter_by_capabilities(self._agents.values(), session_capabilities)

    def render_catalog(self, session_capabilities: Iterable[str] = ()) -> str:
        """Render a compact agent catalog for system prompt injection.

        Applies capability filtering before rendering so capability-gated
        agents stay invisible to sessions that don't advertise them.
        """
        visible = self.visible_for(session_capabilities)
        if not visible:
            return ""
        agent_lines = "\n".join(f"- {agent.name}: {agent.description}" for agent in visible)
        return get_prompt_registry().render("catalog.agent_types", agent_lines=agent_lines)

get(name: str, session_capabilities: Iterable[str] = ()) -> AgentDef | None

Return the agent definition with the given name, or None.

Agents gated by requires_capabilities that the session hasn't advertised are treated as if they don't exist.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_registry.py
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
def get(
    self,
    name: str,
    session_capabilities: Iterable[str] = (),
) -> AgentDef | None:
    """Return the agent definition with the given name, or ``None``.

    Agents gated by ``requires_capabilities`` that the session hasn't
    advertised are treated as if they don't exist.
    """
    agent = self._agents.get(name)
    if agent is None:
        return None
    visible = filter_by_capabilities([agent], session_capabilities)
    return visible[0] if visible else None

list_all() -> list[AgentDef]

Return all registered agent definitions.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_registry.py
285
286
287
def list_all(self) -> list[AgentDef]:
    """Return all registered agent definitions."""
    return list(self._agents.values())

register(agent_def: AgentDef, *, capabilities: Iterable[str] = (), plugin_root: str = '') -> None

Register an agent. Does NOT override existing entries.

When capabilities is non-empty, they are unioned into the agent's requires_capabilities before registration — the standard way a plugin fans its bundle-level requirements out over every contributed agent.

When plugin_root is provided and the agent does not already have one, it is stamped on so downstream consumers (e.g. the ${CLAUDE_PLUGIN_ROOT} substitution in spawn_agent) can locate the plugin's on-disk assets.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_registry.py
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
def register(
    self,
    agent_def: AgentDef,
    *,
    capabilities: Iterable[str] = (),
    plugin_root: str = "",
) -> None:
    """Register an agent.  Does NOT override existing entries.

    When *capabilities* is non-empty, they are unioned into the
    agent's ``requires_capabilities`` before registration — the
    standard way a plugin fans its bundle-level requirements out
    over every contributed agent.

    When *plugin_root* is provided and the agent does not already
    have one, it is stamped on so downstream consumers (e.g. the
    ``${CLAUDE_PLUGIN_ROOT}`` substitution in ``spawn_agent``) can
    locate the plugin's on-disk assets.
    """
    if agent_def.name in self._agents:
        return
    agent_def = overlay_capabilities(agent_def, capabilities)
    if plugin_root and not agent_def.plugin_root:
        agent_def = replace(agent_def, plugin_root=plugin_root)
    self._agents[agent_def.name] = agent_def

render_catalog(session_capabilities: Iterable[str] = ()) -> str

Render a compact agent catalog for system prompt injection.

Applies capability filtering before rendering so capability-gated agents stay invisible to sessions that don't advertise them.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_registry.py
293
294
295
296
297
298
299
300
301
302
303
def render_catalog(self, session_capabilities: Iterable[str] = ()) -> str:
    """Render a compact agent catalog for system prompt injection.

    Applies capability filtering before rendering so capability-gated
    agents stay invisible to sessions that don't advertise them.
    """
    visible = self.visible_for(session_capabilities)
    if not visible:
        return ""
    agent_lines = "\n".join(f"- {agent.name}: {agent.description}" for agent in visible)
    return get_prompt_registry().render("catalog.agent_types", agent_lines=agent_lines)

visible_for(session_capabilities: Iterable[str]) -> list[AgentDef]

Return agents visible given the session's advertised capabilities.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_registry.py
289
290
291
def visible_for(self, session_capabilities: Iterable[str]) -> list[AgentDef]:
    """Return agents visible given the session's advertised capabilities."""
    return filter_by_capabilities(self._agents.values(), session_capabilities)

map_cc_tool_names(cc_names: list[str]) -> list[str]

Map a list of CC tool names to Mewbo tool IDs.

Unknown names pass through unchanged. Duplicates are removed while preserving first-occurrence order (multiple CC names may map to the same Mewbo tool ID).

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_registry.py
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
def map_cc_tool_names(cc_names: list[str]) -> list[str]:
    """Map a list of CC tool names to Mewbo tool IDs.

    Unknown names pass through unchanged.  Duplicates are removed while
    preserving first-occurrence order (multiple CC names may map to the same
    Mewbo tool ID).
    """
    seen: set[str] = set()
    result: list[str] = []
    for name in cc_names:
        mapped = CC_TOOL_MAP.get(name, name)
        if mapped not in seen:
            seen.add(mapped)
            result.append(mapped)
    return result

parse_agent_file(path: Path, source: str) -> AgentDef | None

Parse an agents/*.md file into an :class:AgentDef, or None on failure.

Source code in packages/mewbo_core/src/mewbo_core/agents/agent_registry.py
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
def parse_agent_file(path: Path, source: str) -> AgentDef | None:
    """Parse an ``agents/*.md`` file into an :class:`AgentDef`, or ``None`` on failure."""
    try:
        raw = path.read_text(encoding="utf-8")
    except OSError as exc:
        logging.warning("Failed to read {}: {}", path, exc)
        return None

    match = _FRONTMATTER_RE.match(raw)
    if match is None:
        # No frontmatter — infer a minimal agent def from the markdown body.
        # Third-party plugins (e.g. slack-block-kit-builder) sometimes ship
        # agent files as plain markdown without YAML frontmatter.
        path_key = str(path)
        if path_key not in _LOGGED_INFERRED_FRONTMATTER:
            _LOGGED_INFERRED_FRONTMATTER.add(path_key)
            logging.debug("No YAML frontmatter in {} — inferring from content", path)
        name = path.stem
        # Use first non-blank line (stripped of # prefix) as description
        description = ""
        for line in raw.splitlines():
            stripped = line.strip().lstrip("#").strip()
            if stripped:
                description = stripped[:200]
                break
        return AgentDef(
            name=name,
            description=description,
            source_path=str(path),
            source=source,
            body=raw,
        )

    try:
        meta = yaml.safe_load(match.group(1))
    except yaml.YAMLError as exc:
        logging.warning("Invalid YAML in {}: {}", path, exc)
        return None

    if not isinstance(meta, dict):
        logging.warning("Frontmatter is not a mapping in {}", path)
        return None

    # Name: required, fall back to filename stem
    name_raw = meta.get("name")
    name = name_raw if isinstance(name_raw, str) and name_raw.strip() else path.stem
    name = name.strip()

    # Description: "description" or "when-to-use" alias
    description = meta.get("description") or meta.get("when-to-use") or ""
    if isinstance(description, str):
        description = description.strip()
    else:
        description = ""

    # Allowed tools: space-delimited string or YAML list → mapped via CC_TOOL_MAP
    allowed_tools: list[str] | None = None
    tools_raw = meta.get("tools")
    if isinstance(tools_raw, str) and tools_raw.strip():
        allowed_tools = map_cc_tool_names(tools_raw.strip().split())
    elif isinstance(tools_raw, list):
        allowed_tools = map_cc_tool_names([str(t) for t in tools_raw if t])

    # Denied tools: same parsing
    denied_tools: list[str] | None = None
    disallowed_raw = meta.get("disallowedTools")
    if isinstance(disallowed_raw, str) and disallowed_raw.strip():
        denied_tools = map_cc_tool_names(disallowed_raw.strip().split())
    elif isinstance(disallowed_raw, list):
        denied_tools = map_cc_tool_names([str(t) for t in disallowed_raw if t])

    # Model: "inherit" → None
    model_raw = meta.get("model")
    model: str | None = None
    if isinstance(model_raw, str) and model_raw.strip() and model_raw.strip() != "inherit":
        model = model_raw.strip()

    # Capability gating. Accept either ``requires-capabilities`` (list) or
    # ``requires-capability`` (scalar string); merge if both are given.
    _raw_list = meta.get("requires-capabilities")
    list_form: list = _raw_list if isinstance(_raw_list, list) else []
    _raw_scalar = meta.get("requires-capability")
    scalar_form: str = _raw_scalar if isinstance(_raw_scalar, str) else ""
    requires_capabilities = parse_capabilities([*list_form, scalar_form])

    # Body: everything after the closing frontmatter ---
    body = raw[match.end() :]

    return AgentDef(
        name=name,
        description=description,
        source_path=str(path),
        source=source,
        body=body,
        # ``allowed_tools`` is three-state and an explicit ``tools: []`` is
        # PRESERVED as ``[]`` — an AgentDef declaring no tools must not parse
        # into the unrestricted state. ``denied_tools`` keeps ``or None``
        # deliberately: deny is purely subtractive, so an empty denylist and no
        # denylist are the same set, and there is no third state to lose.
        allowed_tools=allowed_tools,
        denied_tools=denied_tools or None,
        model=model,
        requires_capabilities=requires_capabilities,
    )

packages/mewbo_tools (tool integrations)

mewbo_tools.integration.mcp

MCP tool runner for integrating MCP servers into Mewbo.

MCPToolRunner

Wrapper to invoke MCP tools via langchain-mcp-adapters.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
class MCPToolRunner:
    """Wrapper to invoke MCP tools via langchain-mcp-adapters."""

    def __init__(
        self,
        server_name: str,
        tool_name: str,
        *,
        cwd: str | None = None,
        trust_cwd: bool = True,
    ) -> None:
        """Initialize the MCP tool runner for a specific server tool.

        Args:
            server_name: MCP server name from configuration.
            tool_name: Tool name to invoke on the server.
            cwd: Project working directory for merged config loading.
            trust_cwd: Whether *cwd* may contribute its own ``.mcp.json``.
                A runner RE-RESOLVES the merged config at invocation time, long
                after the registry was built — so a directory whose content
                arrived after that build (a clone, say) is first seen here.
                Carrying the trust decision on the runner is what keeps the
                boundary from holding at discovery and leaking at call time.
        """
        self.server_name = server_name
        self.tool_name = tool_name
        self._cwd = cwd
        self._trust_cwd = trust_cwd

    async def _invoke_async(self, input_payload: str | dict[str, Any]) -> str:
        """Invoke an MCP tool asynchronously and return its output.

        Prefers the connection pool for persistent, cached connections.
        Falls back to the one-shot client path when the pool is unavailable.

        Args:
            input_payload: Input payload to send to the MCP tool.

        Returns:
            Stringified tool response.

        Raises:
            RuntimeError: If MCP adapters are not installed.
            ValueError: If the server or tool is not configured.
        """
        try:
            return await self._invoke_via_pool(input_payload)
        except (OSError, json.JSONDecodeError):
            # The MCP config FILE itself is unreadable/malformed -- already
            # diagnosed loudly at the read site (`_read_json_config`). This is
            # a config problem, not a transport hiccup, so falling back to the
            # one-shot client would just re-hit the identical parse failure.
            raise
        except Exception as pool_exc:
            logging.debug(
                "Pool invocation failed for {}.{}, falling back: {}",
                self.server_name,
                self.tool_name,
                pool_exc,
            )
            return await self._invoke_legacy(input_payload)

    async def _invoke_via_pool(self, input_payload: str | dict[str, Any]) -> str:
        """Invoke via the persistent connection pool."""
        from mewbo_tools.integration.mcp_pool import get_mcp_pool

        pool = get_mcp_pool()

        # A server that is already live must never trigger
        # a full-config reload + reconnect mid-query — that re-dialed EVERY
        # server (incl. dead ones) on every tool call, re-incurring the stall.
        # Only when the target isn't connected do we sync config (so a config
        # change / first use is picked up), and even then with connect=False so
        # the eager dial is deferred to get_or_connect for THIS server alone.
        if not pool.is_connected(self.server_name):
            config = _normalize_mcp_config(
                _load_mcp_config(cwd=self._cwd, trust_cwd=self._trust_cwd)
            )
            await pool.refresh_if_config_changed(config, connect=False)

        state = await pool.get_or_connect(self.server_name)

        tool_map = {getattr(t, "name", ""): t for t in state.tools} if state.tools else {}
        tool = tool_map.get(self.tool_name)
        if tool is None:
            raise ValueError(
                f"Tool '{self.tool_name}' not found on MCP server '{self.server_name}'."
            )

        prepared = _prepare_mcp_input(tool, input_payload)
        result = await pool.call_tool(self.server_name, self.tool_name, prepared)
        return str(result)

    async def _invoke_legacy(self, input_payload: str | dict[str, Any]) -> str:
        """One-shot client path: connect, call, disconnect. No pooling."""
        try:
            from langchain_mcp_adapters.client import MultiServerMCPClient
        except Exception as exc:  # pragma: no cover - runtime dependency
            raise RuntimeError("langchain-mcp-adapters is required for MCP tools.") from exc

        config = _load_mcp_config(cwd=self._cwd, trust_cwd=self._trust_cwd)
        servers = config.get("servers", {})
        if not servers or self.server_name not in servers:
            if self.server_name in disabled_servers():
                raise ValueError(f"MCP server '{self.server_name}' is disabled.")
            raise ValueError(f"MCP server '{self.server_name}' not found in config.")

        client = MultiServerMCPClient({self.server_name: servers[self.server_name]})
        tools = await client.get_tools(server_name=self.server_name)
        tool_map = {tool.name: tool for tool in tools}
        tool = tool_map.get(self.tool_name)
        if tool is None:
            raise ValueError(
                f"Tool '{self.tool_name}' not found on MCP server '{self.server_name}'."
            )
        try:
            result = await tool.ainvoke(_prepare_mcp_input(tool, input_payload))
            return str(result)
        except Exception as exc:
            _log_runtime_failure(self.server_name, self.tool_name, exc)
            raise

    async def arun(self, action_step: ActionStep) -> MockSpeaker:
        """Async execution — preferred when called from an async context.

        Calls ``_invoke_async`` directly, avoiding the ``asyncio.run()``
        wrapper that would fail inside a running event loop.
        """
        if action_step is None:
            raise ValueError("Action step cannot be None.")
        MockSpeakerType = get_mock_speaker()
        result = await self._invoke_async(action_step.tool_input)
        return MockSpeakerType(content=result)

    def run(self, action_step: ActionStep) -> MockSpeaker:
        """Sync execution — for use from sync-only callers.

        Raises:
            ValueError: If action_step is None.
            RuntimeError: If called from inside a running event loop.
        """
        if action_step is None:
            raise ValueError("Action step cannot be None.")
        MockSpeakerType = get_mock_speaker()
        result = asyncio.run(self._invoke_async(action_step.tool_input))
        return MockSpeakerType(content=result)

__init__(server_name: str, tool_name: str, *, cwd: str | None = None, trust_cwd: bool = True) -> None

Initialize the MCP tool runner for a specific server tool.

Parameters:

Name Type Description Default
server_name str

MCP server name from configuration.

required
tool_name str

Tool name to invoke on the server.

required
cwd str | None

Project working directory for merged config loading.

None
trust_cwd bool

Whether cwd may contribute its own .mcp.json. A runner RE-RESOLVES the merged config at invocation time, long after the registry was built — so a directory whose content arrived after that build (a clone, say) is first seen here. Carrying the trust decision on the runner is what keeps the boundary from holding at discovery and leaking at call time.

True
Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
def __init__(
    self,
    server_name: str,
    tool_name: str,
    *,
    cwd: str | None = None,
    trust_cwd: bool = True,
) -> None:
    """Initialize the MCP tool runner for a specific server tool.

    Args:
        server_name: MCP server name from configuration.
        tool_name: Tool name to invoke on the server.
        cwd: Project working directory for merged config loading.
        trust_cwd: Whether *cwd* may contribute its own ``.mcp.json``.
            A runner RE-RESOLVES the merged config at invocation time, long
            after the registry was built — so a directory whose content
            arrived after that build (a clone, say) is first seen here.
            Carrying the trust decision on the runner is what keeps the
            boundary from holding at discovery and leaking at call time.
    """
    self.server_name = server_name
    self.tool_name = tool_name
    self._cwd = cwd
    self._trust_cwd = trust_cwd

arun(action_step: ActionStep) -> MockSpeaker async

Async execution — preferred when called from an async context.

Calls _invoke_async directly, avoiding the asyncio.run() wrapper that would fail inside a running event loop.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
616
617
618
619
620
621
622
623
624
625
626
async def arun(self, action_step: ActionStep) -> MockSpeaker:
    """Async execution — preferred when called from an async context.

    Calls ``_invoke_async`` directly, avoiding the ``asyncio.run()``
    wrapper that would fail inside a running event loop.
    """
    if action_step is None:
        raise ValueError("Action step cannot be None.")
    MockSpeakerType = get_mock_speaker()
    result = await self._invoke_async(action_step.tool_input)
    return MockSpeakerType(content=result)

run(action_step: ActionStep) -> MockSpeaker

Sync execution — for use from sync-only callers.

Raises:

Type Description
ValueError

If action_step is None.

RuntimeError

If called from inside a running event loop.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
628
629
630
631
632
633
634
635
636
637
638
639
def run(self, action_step: ActionStep) -> MockSpeaker:
    """Sync execution — for use from sync-only callers.

    Raises:
        ValueError: If action_step is None.
        RuntimeError: If called from inside a running event loop.
    """
    if action_step is None:
        raise ValueError("Action step cannot be None.")
    MockSpeakerType = get_mock_speaker()
    result = asyncio.run(self._invoke_async(action_step.tool_input))
    return MockSpeakerType(content=result)

disabled_servers() -> frozenset[str]

Server names switched off in config at the most recent load.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
110
111
112
def disabled_servers() -> frozenset[str]:
    """Server names switched off in config at the most recent load."""
    return frozenset(_DISABLED_SERVERS)

discover_mcp_tool_details(config: dict[str, Any]) -> dict[str, list[dict[str, Any]]]

Discover MCP tool names and schemas per server from configuration.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
438
439
440
def discover_mcp_tool_details(config: dict[str, Any]) -> dict[str, list[dict[str, Any]]]:
    """Discover MCP tool names and schemas per server from configuration."""
    return asyncio.run(_discover_mcp_tool_details_async(_normalize_mcp_config(config)))

discover_mcp_tool_details_with_failures(config: dict[str, Any]) -> tuple[dict[str, list[dict[str, Any]]], dict[str, Exception]]

Discover MCP tool names, schemas, and per-server failures.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
443
444
445
446
447
448
449
450
451
def discover_mcp_tool_details_with_failures(
    config: dict[str, Any],
) -> tuple[dict[str, list[dict[str, Any]]], dict[str, Exception]]:
    """Discover MCP tool names, schemas, and per-server failures."""
    discovered, failures = asyncio.run(
        _discover_mcp_tool_details_with_failures_async(_normalize_mcp_config(config))
    )
    _record_discovery_failures(failures)
    return discovered, failures

discover_mcp_tools(config: dict[str, Any]) -> dict[str, list[str]]

Discover MCP tool names per server from configuration.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
429
430
431
432
433
434
435
def discover_mcp_tools(config: dict[str, Any]) -> dict[str, list[str]]:
    """Discover MCP tool names per server from configuration."""
    details = discover_mcp_tool_details(config)
    return {
        server_name: [tool["name"] for tool in tools if tool.get("name")]
        for server_name, tools in details.items()
    }

get_last_config_error() -> dict[str, str] | None

Return the most recent malformed-MCP-config diagnostic, if any.

Read by MCPConnectionPool.status_snapshot so a broken config file is visible in /mcp alongside per-server connect states.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
208
209
210
211
212
213
214
def get_last_config_error() -> dict[str, str] | None:
    """Return the most recent malformed-MCP-config diagnostic, if any.

    Read by ``MCPConnectionPool.status_snapshot`` so a broken config file is
    visible in ``/mcp`` alongside per-server connect states.
    """
    return dict(_LAST_CONFIG_ERROR) if _LAST_CONFIG_ERROR is not None else None

get_last_discovery_failures() -> dict[str, str]

Return last MCP discovery failures per server (if any).

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
461
462
463
def get_last_discovery_failures() -> dict[str, str]:
    """Return last MCP discovery failures per server (if any)."""
    return dict(_LAST_DISCOVERY_FAILURES)

list_server_tool_schemas(server_name: str, *, cwd: str | None = None, trust_cwd: bool = True) -> list[dict[str, Any]]

Return one configured server's live tool schemas via the shared pool.

The public introspection seam for callers that need a server's advertised tool list (name / description / input schema) without binding the tools: loads the merged MCP config for cwd, refreshes the pool fingerprint, and connects on demand (the MCPToolRunner._invoke_via_pool pattern). Only schema-bearing attributes are read off each tool — never connection or auth material.

Raises:

Type Description
LookupError

server_name has no entry in the merged MCP config.

RuntimeError

the config could not be read, or the live introspection (pool connect / handshake) failed.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
def list_server_tool_schemas(
    server_name: str,
    *,
    cwd: str | None = None,
    trust_cwd: bool = True,
) -> list[dict[str, Any]]:
    """Return one configured server's live tool schemas via the shared pool.

    The public introspection seam for callers that need a server's advertised
    tool list (name / description / input schema) without binding the tools:
    loads the merged MCP config for *cwd*, refreshes the pool fingerprint, and
    connects on demand (the ``MCPToolRunner._invoke_via_pool`` pattern). Only
    schema-bearing attributes are read off each tool — never connection or
    auth material.

    Raises:
        LookupError: *server_name* has no entry in the merged MCP config.
        RuntimeError: the config could not be read, or the live introspection
            (pool connect / handshake) failed.
    """
    try:
        config = _load_mcp_config(cwd=cwd, trust_cwd=trust_cwd)
    except Exception as exc:
        raise RuntimeError(f"failed to read MCP config: {exc}") from exc
    servers = config.get("servers", {}) if isinstance(config, dict) else {}
    if server_name not in servers:
        raise LookupError(f"MCP server '{server_name}' is not configured.")
    return asyncio.run(_list_server_tool_schemas_async(server_name, config))

mark_tool_auto_approved(config: dict[str, Any], server_name: str, tool_name: str) -> dict[str, Any]

Record a tool as auto-approved in the MCP config.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
479
480
481
482
483
484
485
486
487
488
489
490
491
def mark_tool_auto_approved(
    config: dict[str, Any],
    server_name: str,
    tool_name: str,
) -> dict[str, Any]:
    """Record a tool as auto-approved in the MCP config."""
    servers = config.setdefault("servers", {})
    server_config = servers.setdefault(server_name, {})
    allowlist = server_config.setdefault("auto_approve_tools", [])
    if tool_name not in allowlist:
        allowlist.append(tool_name)
        server_config["auto_approve_tools"] = sorted(set(allowlist))
    return config

save_mcp_config(config: dict[str, Any], path: str | None = None) -> None

Persist an MCP configuration payload to disk.

Parameters:

Name Type Description Default
config dict[str, Any]

MCP configuration payload to write.

required
path str | None

Optional explicit file path (defaults to the configured MCP path).

None
Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
286
287
288
289
290
291
292
293
294
295
296
297
298
299
def save_mcp_config(config: dict[str, Any], path: str | None = None) -> None:
    """Persist an MCP configuration payload to disk.

    Args:
        config: MCP configuration payload to write.
        path: Optional explicit file path (defaults to the configured MCP path).
    """
    config_path = path or get_mcp_config_path()
    if not config_path:
        raise ValueError("MCP config path is not set.")
    config_path = os.path.abspath(config_path)
    with open(config_path, "w", encoding="utf-8") as handle:
        json.dump(config, handle, indent=2)
        handle.write("\n")

tool_auto_approved(config: dict[str, Any], server_name: str, tool_name: str) -> bool

Return True when a tool is marked as auto-approved.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/mcp.py
466
467
468
469
470
471
472
473
474
475
476
def tool_auto_approved(
    config: dict[str, Any],
    server_name: str,
    tool_name: str,
) -> bool:
    """Return True when a tool is marked as auto-approved."""
    server_config = config.get("servers", {}).get(server_name, {})
    if server_config.get("auto_approve_all"):
        return True
    allowlist = server_config.get("auto_approve_tools", [])
    return tool_name in allowlist

mewbo_tools.integration.homeassistant

Home Assistant integration tools and data models.

CacheHolder

Bases: Protocol

Protocol describing objects with a Home Assistant cache attribute.

Attributes:

Name Type Description
cache HomeAssistantCache

Home Assistant cache payload.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
46
47
48
49
50
51
52
53
54
@runtime_checkable
class CacheHolder(Protocol):
    """Protocol describing objects with a Home Assistant cache attribute.

    Attributes:
        cache: Home Assistant cache payload.
    """

    cache: HomeAssistantCache

HomeAssistant

Bases: AbstractTool

A service to manage and interact with Home Assistant.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
class HomeAssistant(AbstractTool):
    """A service to manage and interact with Home Assistant."""

    def __init__(self) -> None:
        """Initialize the Home Assistant tool with environment defaults."""
        super().__init__(
            name="Home Assistant",
            description="A service to manage and interact with Home Assistant",
        )
        self.base_url = get_config_value("home_assistant", "url")
        self._api_token = get_config_value("home_assistant", "token")
        self.cache: HomeAssistantCache = {
            "entity_ids": [],
            "sensor_ids": [],
            "entities": [],
            "services": [],
            "sensors": [],
            "allowed_domains": ["scene", "switch", "weather", "kodi", "automation"],
        }

        if not self.base_url or not self._api_token:
            raise ValueError("home_assistant.url and home_assistant.token must be set.")

        self.api_headers: dict[str, str] = {
            "Authorization": f"Bearer {self._api_token}",
            "Content-Type": "application/json",
        }

    @cache_monitor
    def update_services(self) -> bool:
        """Update the list of services from Home Assistant.

        Returns:
            True when services are fetched successfully.
        """
        url = f"{self.base_url}/services"
        try:
            response = requests.get(url, headers=self.api_headers, timeout=30)
            status_code = getattr(response, "status_code", None)
            if status_code in {401, 403}:
                logging.error("Home Assistant authorization failed with status {}.", status_code)
                raise PermissionError("Home Assistant authorization failed.")
            response.raise_for_status()
            self.cache["services"] = response.json()
            self._save_json(self.cache["services"], "services.json")
            return True
        except requests.exceptions.RequestException as e:
            logging.error("Error: {}", e)
            return False

    @cache_monitor
    def update_entities(self) -> bool:
        """Update the list of entities from Home Assistant.

        Returns:
            True when entities are fetched successfully.
        """
        url = f"{self.base_url}/states"
        try:
            response = requests.get(url, headers=self.api_headers, timeout=30)
            status_code = getattr(response, "status_code", None)
            if status_code in {401, 403}:
                logging.error("Home Assistant authorization failed with status {}.", status_code)
                raise PermissionError("Home Assistant authorization failed.")
            response.raise_for_status()
            self.cache["entities"] = response.json()
            return True
        except requests.exceptions.RequestException as e:
            logging.error("Error: {}", e)
            return False

    @cache_monitor
    def update_entity_ids(self) -> bool:
        """Update the list of entity IDs from Home Assistant.

        Returns:
            True when entity IDs are populated.

        Raises:
            ValueError: If no entities are available for ID extraction.
        """
        # TODO: Always assumes blacklist by default due to cache_monitor.
        self.update_entities()
        entities = self.cache["entities"]
        if not entities:
            raise ValueError("No entities found while updating entity IDs.")
        self.cache["entity_ids"] = [entity["entity_id"] for entity in entities]
        logging.info("Entity IDs updated.")
        return True

    @cache_monitor
    def update_cache(self) -> None:
        """Update the entire cache.

        Raises:
            ValueError: If entity IDs cannot be derived.
        """
        self.update_entity_ids()
        self.update_services()
        self._save_json(self.cache["entities"], "entities.json")
        self._save_json(self.cache["sensors"], "sensors.json")

    def call_service(
        self,
        domain: str,
        service: str,
        entity_id: str,
        data: dict | None = None,
    ) -> tuple[bool, list[dict[str, Any]]]:
        """Call a service in Home Assistant.

        Args:
            domain: Home Assistant domain name (e.g., "light").
            service: Service name within the domain (e.g., "turn_on").
            entity_id: Entity ID to target.
            data: Optional extra payload for the service call.

        Returns:
            Tuple of success flag and JSON response payload.

        Raises:
            ValueError: If the domain is not allowed.
        """
        if domain not in self.cache["allowed_domains"]:
            raise ValueError(f"Domain does not exist or blacklisted: {domain}")

        url = f"{self.base_url}/services/{domain}/{service}"
        payload = {"entity_id": entity_id}
        if data:
            payload.update(data)

        try:
            response = requests.post(url, headers=self.api_headers, json=payload, timeout=30)
            status_code = getattr(response, "status_code", None)
            if status_code in {401, 403}:
                logging.error("Home Assistant authorization failed with status {}.", status_code)
                raise PermissionError("Home Assistant authorization failed.")
            response.raise_for_status()
            logging.info(
                "Service <{}.{}> called on entity <{}> returned `{}`.",
                domain,
                service,
                entity_id,
                response.text,
            )
            return True, response.json()
        except requests.exceptions.RequestException as e:
            logging.error(
                "Unable to call service <{}.{}> on entity <{}>: {}", domain, service, entity_id, e
            )
            return False, []

    @staticmethod
    def _create_set_prompt(
        system_prompt: str,
        parser: PydanticOutputParser,
    ) -> ChatPromptTemplate:
        """Create the prompt template for a set-state operation.

        Args:
            system_prompt: System prompt content.
            parser: Pydantic output parser for HomeAssistantCall.

        Returns:
            ChatPromptTemplate configured for set-state tasks.
        """
        example = HomeAssistantCall(
            domain="scene", service="turn_on", entity_id="scene.lamp_power_on"
        )
        prompt = ChatPromptTemplate(
            messages=[
                SystemMessage(content=system_prompt),
                HumanMessage(content="Turn on the lamp lights."),
                AIMessage(content=example.model_dump_json()),
                HumanMessagePromptTemplate.from_template(
                    "The user asked you to `{action_step}`. You must use the information "
                    "provided to pick the right Home Assistant service call values only "
                    "considering the current user query.\n\n"
                    "## Format Instructions\n{format_instructions}\n\n"
                    "## Home Assistant Entities and Domain-Services\n```\n{context}```\n"
                ),
            ],
            partial_variables={"format_instructions": parser.get_format_instructions()},
            input_variables=["action_step"],
        )
        return prompt

    @staticmethod
    def _create_get_prompt(system_prompt: str) -> ChatPromptTemplate:
        """Create the prompt template for a get-state operation.

        Args:
            system_prompt: System prompt content.

        Returns:
            ChatPromptTemplate configured for get-state tasks.
        """
        prompt = ChatPromptTemplate(
            messages=[
                SystemMessage(content=system_prompt),
                HumanMessage(content="How is the air quality today?"),
                AIMessage(
                    content=(
                        "AccuWeather reported today's air quality in your home as good. "
                        "This level of air quality ensures that the environment is healthy, "
                        "supporting your daily activities and wellbeing without any air "
                        "quality-related risks."
                    )
                ),
                HumanMessagePromptTemplate.from_template(
                    "The user asked you to `{action_step}`. You must use the sensor "
                    "information to answer the user's query. Keep your answer "
                    "analytical, brief and useful.\n\n"
                    "## Home Assistant Sensors\n```\n{context}```\n"
                ),
            ],
            input_variables=["action_step"],
        )
        return prompt

    @staticmethod
    def _clean_answer(answer: str) -> str:
        """Clean the answer by removing/replacing characters.

        Args:
            answer: Raw answer string to normalize.

        Returns:
            Cleaned answer string.
        """
        replacements = {
            # Common entities
            "RealFeel": "Real Feel",
            # Confident Abbreviations
            "km/h": " kilometer per hour",
            "°C": " degrees celsius",
            "%": " percent",
            "mm/h": " millimeter per hour",
            "Gb/s": " gigabits per second",
            "Mb/s": " megabits per second",
            "Kb/s": " kilobits per second",
            "GHz": "Gigahertz",
            # Formatting
            '"': "",
        }

        for old, new in replacements.items():
            answer = answer.replace(old, new)

        # Remove extra spaces and new lines, condense all multiple spaces
        #   to a single space
        answer = re.sub(r"\s+", " ", answer).strip()

        return answer

    def _invoke_service_and_set_state(
        self,
        chain: SupportsInvoke,
        rag_documents: list[Document],
        action_step: ActionStep,
    ) -> MockSpeaker:
        """Invoke the service and set the state.

        Args:
            chain: Runnable chain that yields HomeAssistantCall.
            rag_documents: Context documents for the chain.
            action_step: Action step describing the request.

        Returns:
            MockSpeaker with a status message.
        """
        MockSpeaker = get_mock_speaker()

        try:
            action_step_curr = str(action_step.tool_input).strip()
            call_service_values = chain.invoke(
                {"action_step": action_step_curr, "context": rag_documents, "cache": self.cache},
            )
            logging.debug(
                "Call Service Values for `{}`: `{}`", action_step_curr, call_service_values
            )
            status_bool, response_json = self.call_service(
                domain=call_service_values.domain,
                service=call_service_values.service,
                entity_id=call_service_values.entity_id,
            )
            if status_bool:
                tmp_return_message = f"Successfully called service: `{response_json}`"
            else:
                tmp_return_message = f"Failed to call service: `{response_json}`"
        except Exception as err_mesaage:
            logging.error("Error: {}", err_mesaage)
            tmp_return_message = f"I received an error - `{err_mesaage}`"
        return MockSpeaker(content=tmp_return_message)

    def set_state(self, action_step: ActionStep | None = None) -> MockSpeaker:
        """Predict and call a service for a given action step.

        Args:
            action_step: Action step describing the desired change.

        Returns:
            MockSpeaker with a status message.

        Raises:
            ValueError: If action_step is None.
        """
        if action_step is None:
            raise ValueError("Action step cannot be None.")
        self.update_cache()
        rag_documents = self._load_rag_documents(["entities.json", "services.json"])
        system_prompt = ha_render_system_prompt(
            name="homeassistant-set-state", all_entities=self.cache["entity_ids"]
        )

        parser = PydanticOutputParser(pydantic_object=HomeAssistantCall)
        prompt = self._create_set_prompt(system_prompt, parser)
        if self.model is None:
            raise RuntimeError("LLM client not initialized for Home Assistant.")
        model = self.model
        chain: Any = prompt | model | parser

        logging.info(
            "Invoking `set` action chain using `{}` for `{}`.", self.model_name, action_step
        )
        # TODO: Interpret the response from call service.
        return self._invoke_service_and_set_state(chain, rag_documents, action_step)

    def get_state(self, action_step: ActionStep | None = None) -> MockSpeaker:
        """Generate response for a given action step based on sensors.

        Args:
            action_step: Action step describing the desired query.

        Returns:
            MockSpeaker with the generated response.

        Raises:
            ValueError: If action_step is None.
        """
        if action_step is None:
            raise ValueError("Action step cannot be None.")
        self.update_cache()
        rag_documents = self._load_rag_documents(["sensors.json"])

        system_prompt = ha_render_system_prompt(name="homeassistant-get-state")

        prompt = self._create_get_prompt(system_prompt)
        if self.model is None:
            raise RuntimeError("LLM client not initialized for Home Assistant.")
        model = self.model
        chain: Any = prompt | model

        logging.info("Invoking `get` action chain using `{}`.", self.model_name)
        message = chain.invoke(
            {
                "action_step": str(action_step.tool_input).strip(),
                "context": rag_documents,
            },
        )
        cleaned_message = self._clean_answer(str(message.content))
        MockSpeaker = get_mock_speaker()
        return MockSpeaker(content=cleaned_message)

__init__() -> None

Initialize the Home Assistant tool with environment defaults.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
def __init__(self) -> None:
    """Initialize the Home Assistant tool with environment defaults."""
    super().__init__(
        name="Home Assistant",
        description="A service to manage and interact with Home Assistant",
    )
    self.base_url = get_config_value("home_assistant", "url")
    self._api_token = get_config_value("home_assistant", "token")
    self.cache: HomeAssistantCache = {
        "entity_ids": [],
        "sensor_ids": [],
        "entities": [],
        "services": [],
        "sensors": [],
        "allowed_domains": ["scene", "switch", "weather", "kodi", "automation"],
    }

    if not self.base_url or not self._api_token:
        raise ValueError("home_assistant.url and home_assistant.token must be set.")

    self.api_headers: dict[str, str] = {
        "Authorization": f"Bearer {self._api_token}",
        "Content-Type": "application/json",
    }

call_service(domain: str, service: str, entity_id: str, data: dict | None = None) -> tuple[bool, list[dict[str, Any]]]

Call a service in Home Assistant.

Parameters:

Name Type Description Default
domain str

Home Assistant domain name (e.g., "light").

required
service str

Service name within the domain (e.g., "turn_on").

required
entity_id str

Entity ID to target.

required
data dict | None

Optional extra payload for the service call.

None

Returns:

Type Description
tuple[bool, list[dict[str, Any]]]

Tuple of success flag and JSON response payload.

Raises:

Type Description
ValueError

If the domain is not allowed.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
def call_service(
    self,
    domain: str,
    service: str,
    entity_id: str,
    data: dict | None = None,
) -> tuple[bool, list[dict[str, Any]]]:
    """Call a service in Home Assistant.

    Args:
        domain: Home Assistant domain name (e.g., "light").
        service: Service name within the domain (e.g., "turn_on").
        entity_id: Entity ID to target.
        data: Optional extra payload for the service call.

    Returns:
        Tuple of success flag and JSON response payload.

    Raises:
        ValueError: If the domain is not allowed.
    """
    if domain not in self.cache["allowed_domains"]:
        raise ValueError(f"Domain does not exist or blacklisted: {domain}")

    url = f"{self.base_url}/services/{domain}/{service}"
    payload = {"entity_id": entity_id}
    if data:
        payload.update(data)

    try:
        response = requests.post(url, headers=self.api_headers, json=payload, timeout=30)
        status_code = getattr(response, "status_code", None)
        if status_code in {401, 403}:
            logging.error("Home Assistant authorization failed with status {}.", status_code)
            raise PermissionError("Home Assistant authorization failed.")
        response.raise_for_status()
        logging.info(
            "Service <{}.{}> called on entity <{}> returned `{}`.",
            domain,
            service,
            entity_id,
            response.text,
        )
        return True, response.json()
    except requests.exceptions.RequestException as e:
        logging.error(
            "Unable to call service <{}.{}> on entity <{}>: {}", domain, service, entity_id, e
        )
        return False, []

get_state(action_step: ActionStep | None = None) -> MockSpeaker

Generate response for a given action step based on sensors.

Parameters:

Name Type Description Default
action_step ActionStep | None

Action step describing the desired query.

None

Returns:

Type Description
MockSpeaker

MockSpeaker with the generated response.

Raises:

Type Description
ValueError

If action_step is None.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
def get_state(self, action_step: ActionStep | None = None) -> MockSpeaker:
    """Generate response for a given action step based on sensors.

    Args:
        action_step: Action step describing the desired query.

    Returns:
        MockSpeaker with the generated response.

    Raises:
        ValueError: If action_step is None.
    """
    if action_step is None:
        raise ValueError("Action step cannot be None.")
    self.update_cache()
    rag_documents = self._load_rag_documents(["sensors.json"])

    system_prompt = ha_render_system_prompt(name="homeassistant-get-state")

    prompt = self._create_get_prompt(system_prompt)
    if self.model is None:
        raise RuntimeError("LLM client not initialized for Home Assistant.")
    model = self.model
    chain: Any = prompt | model

    logging.info("Invoking `get` action chain using `{}`.", self.model_name)
    message = chain.invoke(
        {
            "action_step": str(action_step.tool_input).strip(),
            "context": rag_documents,
        },
    )
    cleaned_message = self._clean_answer(str(message.content))
    MockSpeaker = get_mock_speaker()
    return MockSpeaker(content=cleaned_message)

set_state(action_step: ActionStep | None = None) -> MockSpeaker

Predict and call a service for a given action step.

Parameters:

Name Type Description Default
action_step ActionStep | None

Action step describing the desired change.

None

Returns:

Type Description
MockSpeaker

MockSpeaker with a status message.

Raises:

Type Description
ValueError

If action_step is None.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
def set_state(self, action_step: ActionStep | None = None) -> MockSpeaker:
    """Predict and call a service for a given action step.

    Args:
        action_step: Action step describing the desired change.

    Returns:
        MockSpeaker with a status message.

    Raises:
        ValueError: If action_step is None.
    """
    if action_step is None:
        raise ValueError("Action step cannot be None.")
    self.update_cache()
    rag_documents = self._load_rag_documents(["entities.json", "services.json"])
    system_prompt = ha_render_system_prompt(
        name="homeassistant-set-state", all_entities=self.cache["entity_ids"]
    )

    parser = PydanticOutputParser(pydantic_object=HomeAssistantCall)
    prompt = self._create_set_prompt(system_prompt, parser)
    if self.model is None:
        raise RuntimeError("LLM client not initialized for Home Assistant.")
    model = self.model
    chain: Any = prompt | model | parser

    logging.info(
        "Invoking `set` action chain using `{}` for `{}`.", self.model_name, action_step
    )
    # TODO: Interpret the response from call service.
    return self._invoke_service_and_set_state(chain, rag_documents, action_step)

update_cache() -> None

Update the entire cache.

Raises:

Type Description
ValueError

If entity IDs cannot be derived.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
385
386
387
388
389
390
391
392
393
394
395
@cache_monitor
def update_cache(self) -> None:
    """Update the entire cache.

    Raises:
        ValueError: If entity IDs cannot be derived.
    """
    self.update_entity_ids()
    self.update_services()
    self._save_json(self.cache["entities"], "entities.json")
    self._save_json(self.cache["sensors"], "sensors.json")

update_entities() -> bool

Update the list of entities from Home Assistant.

Returns:

Type Description
bool

True when entities are fetched successfully.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
@cache_monitor
def update_entities(self) -> bool:
    """Update the list of entities from Home Assistant.

    Returns:
        True when entities are fetched successfully.
    """
    url = f"{self.base_url}/states"
    try:
        response = requests.get(url, headers=self.api_headers, timeout=30)
        status_code = getattr(response, "status_code", None)
        if status_code in {401, 403}:
            logging.error("Home Assistant authorization failed with status {}.", status_code)
            raise PermissionError("Home Assistant authorization failed.")
        response.raise_for_status()
        self.cache["entities"] = response.json()
        return True
    except requests.exceptions.RequestException as e:
        logging.error("Error: {}", e)
        return False

update_entity_ids() -> bool

Update the list of entity IDs from Home Assistant.

Returns:

Type Description
bool

True when entity IDs are populated.

Raises:

Type Description
ValueError

If no entities are available for ID extraction.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
@cache_monitor
def update_entity_ids(self) -> bool:
    """Update the list of entity IDs from Home Assistant.

    Returns:
        True when entity IDs are populated.

    Raises:
        ValueError: If no entities are available for ID extraction.
    """
    # TODO: Always assumes blacklist by default due to cache_monitor.
    self.update_entities()
    entities = self.cache["entities"]
    if not entities:
        raise ValueError("No entities found while updating entity IDs.")
    self.cache["entity_ids"] = [entity["entity_id"] for entity in entities]
    logging.info("Entity IDs updated.")
    return True

update_services() -> bool

Update the list of services from Home Assistant.

Returns:

Type Description
bool

True when services are fetched successfully.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
@cache_monitor
def update_services(self) -> bool:
    """Update the list of services from Home Assistant.

    Returns:
        True when services are fetched successfully.
    """
    url = f"{self.base_url}/services"
    try:
        response = requests.get(url, headers=self.api_headers, timeout=30)
        status_code = getattr(response, "status_code", None)
        if status_code in {401, 403}:
            logging.error("Home Assistant authorization failed with status {}.", status_code)
            raise PermissionError("Home Assistant authorization failed.")
        response.raise_for_status()
        self.cache["services"] = response.json()
        self._save_json(self.cache["services"], "services.json")
        return True
    except requests.exceptions.RequestException as e:
        logging.error("Error: {}", e)
        return False

HomeAssistantCache

Bases: TypedDict

Cached Home Assistant entity and service metadata.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
34
35
36
37
38
39
40
41
42
43
class HomeAssistantCache(TypedDict):
    """Cached Home Assistant entity and service metadata."""

    entity_ids: list[str]
    sensor_ids: list[str]
    entities: list[dict[str, Any]]
    services: list[dict[str, Any]]
    sensors: list[dict[str, Any]]
    allowed_domains: list[str]
    sensor: NotRequired[list[dict[str, Any]]]

HomeAssistantCall

Bases: BaseModel

Structured Home Assistant service call extracted from the model output.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
class HomeAssistantCall(BaseModel):
    """Structured Home Assistant service call extracted from the model output."""

    # SkipJsonSchema: ``cache`` is injected at runtime for validation only, not
    # filled by the LLM. CacheHolder is a Protocol (no JSON schema), so without
    # this PydanticOutputParser(pydantic_object=HomeAssistantCall) raises
    # PydanticInvalidForJsonSchema and breaks the entire `set` action path.
    cache: SkipJsonSchema[CacheHolder | None] = Field(alias="_ha_cache", default=None)
    domain: str = Field(
        description=("The category of the service to call, such as 'light', 'switch', or 'scene'.")
    )
    service: str = Field(
        description=(
            "The specific action to perform within the domain, such as 'turn_on', "
            "'turn_off', or 'set_temperature'."
        )
    )
    entity_id: str = Field(
        description=(
            "The ID of the specific device or entity within the domain to apply the "
            "service to, such as 'scene.heater'."
        )
    )

    @field_validator("entity_id")
    @classmethod
    def validate_entity_id(cls: type[Any], entity_id: str, info: ValidationInfo) -> str:
        """Validate the entity_id against the cache when available.

        Args:
            cls: Pydantic model class.
            entity_id: Candidate entity identifier.
            info: Pydantic validation info with access to other field values.

        Returns:
            Validated entity identifier.

        Raises:
            ValueError: If the entity ID is not found in the cache.
        """
        # ! BUG: The entity_id may not be validated correctly as the cache
        # !     is not passed to the validator.
        ha_cache = info.data.get("cache")
        if ha_cache and entity_id not in ha_cache.cache["entity_ids"]:
            raise ValueError(f"Entity ID '{entity_id}' is not in the Home Assistant cache.")
        return entity_id

    @field_validator("domain")
    @classmethod
    def validate_domain(cls: type[Any], domain: str, info: ValidationInfo) -> str:
        """Validate the domain against the cache when available.

        Args:
            cls: Pydantic model class.
            domain: Domain string to validate.
            info: Pydantic validation info with access to other field values.

        Returns:
            Validated domain string.

        Raises:
            ValueError: If the domain is not found in the cache.
        """
        # ! BUG: The entity_id may not be validated correctly as the cache
        # !     is not passed to the validator.
        ha_cache = info.data.get("cache")
        if ha_cache and domain not in ha_cache.cache["allowed_domains"]:
            raise ValueError(f"Domain '{domain}' is not in the Home Assistant cache.")
        return domain

    model_config = ConfigDict(arbitrary_types_allowed=True)

validate_domain(domain: str, info: ValidationInfo) -> str classmethod

Validate the domain against the cache when available.

Parameters:

Name Type Description Default
cls type[Any]

Pydantic model class.

required
domain str

Domain string to validate.

required
info ValidationInfo

Pydantic validation info with access to other field values.

required

Returns:

Type Description
str

Validated domain string.

Raises:

Type Description
ValueError

If the domain is not found in the cache.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
@field_validator("domain")
@classmethod
def validate_domain(cls: type[Any], domain: str, info: ValidationInfo) -> str:
    """Validate the domain against the cache when available.

    Args:
        cls: Pydantic model class.
        domain: Domain string to validate.
        info: Pydantic validation info with access to other field values.

    Returns:
        Validated domain string.

    Raises:
        ValueError: If the domain is not found in the cache.
    """
    # ! BUG: The entity_id may not be validated correctly as the cache
    # !     is not passed to the validator.
    ha_cache = info.data.get("cache")
    if ha_cache and domain not in ha_cache.cache["allowed_domains"]:
        raise ValueError(f"Domain '{domain}' is not in the Home Assistant cache.")
    return domain

validate_entity_id(entity_id: str, info: ValidationInfo) -> str classmethod

Validate the entity_id against the cache when available.

Parameters:

Name Type Description Default
cls type[Any]

Pydantic model class.

required
entity_id str

Candidate entity identifier.

required
info ValidationInfo

Pydantic validation info with access to other field values.

required

Returns:

Type Description
str

Validated entity identifier.

Raises:

Type Description
ValueError

If the entity ID is not found in the cache.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
@field_validator("entity_id")
@classmethod
def validate_entity_id(cls: type[Any], entity_id: str, info: ValidationInfo) -> str:
    """Validate the entity_id against the cache when available.

    Args:
        cls: Pydantic model class.
        entity_id: Candidate entity identifier.
        info: Pydantic validation info with access to other field values.

    Returns:
        Validated entity identifier.

    Raises:
        ValueError: If the entity ID is not found in the cache.
    """
    # ! BUG: The entity_id may not be validated correctly as the cache
    # !     is not passed to the validator.
    ha_cache = info.data.get("cache")
    if ha_cache and entity_id not in ha_cache.cache["entity_ids"]:
        raise ValueError(f"Entity ID '{entity_id}' is not in the Home Assistant cache.")
    return entity_id

SupportsInvoke

Bases: Protocol

Protocol for runnable chains that return HomeAssistantCall.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
57
58
59
60
61
62
63
64
65
66
67
68
69
class SupportsInvoke(Protocol):
    """Protocol for runnable chains that return HomeAssistantCall."""

    def invoke(self, input_data: dict[str, Any]) -> HomeAssistantCall:
        """Invoke the chain with structured input.

        Args:
            input_data: Input payload for the chain.

        Returns:
            Parsed HomeAssistantCall.
        """
        ...

invoke(input_data: dict[str, Any]) -> HomeAssistantCall

Invoke the chain with structured input.

Parameters:

Name Type Description Default
input_data dict[str, Any]

Input payload for the chain.

required

Returns:

Type Description
HomeAssistantCall

Parsed HomeAssistantCall.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
60
61
62
63
64
65
66
67
68
69
def invoke(self, input_data: dict[str, Any]) -> HomeAssistantCall:
    """Invoke the chain with structured input.

    Args:
        input_data: Input payload for the chain.

    Returns:
        Parsed HomeAssistantCall.
    """
    ...

cache_monitor(func: Callable[Concatenate[SelfT, P], R]) -> Callable[Concatenate[SelfT, P], R]

Decorator to monitor and update the cache.

Parameters:

Name Type Description Default
func Callable[Concatenate[SelfT, P], R]

Method that updates a portion of the cache.

required

Returns:

Type Description
Callable[Concatenate[SelfT, P], R]

Wrapped function that normalizes cache contents after execution.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/homeassistant.py
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
def cache_monitor(func: Callable[Concatenate[SelfT, P], R]) -> Callable[Concatenate[SelfT, P], R]:
    """Decorator to monitor and update the cache.

    Args:
        func: Method that updates a portion of the cache.

    Returns:
        Wrapped function that normalizes cache contents after execution.
    """

    def sort_by_entity_id(dict_list: list[dict[str, Any]]) -> list[dict[str, Any]]:
        """Sort a list of entities by the entity_id field.

        Args:
            dict_list: List of entity dictionaries.

        Returns:
            Sorted list of entities.
        """
        return sorted(dict_list, key=lambda x: x["entity_id"])

    def clean_entities(
        self: CacheHolder,
        forbidden_prefixes: list[str],
        forbidden_substrings: list[str],
    ) -> HomeAssistantCache:
        """Filter and normalize entities while populating sensors.

        Args:
            self: Cache holder to mutate.
            forbidden_prefixes: Entity ID prefixes to exclude.
            forbidden_substrings: Entity ID substrings to exclude.

        Returns:
            Updated HomeAssistantCache payload.
        """
        # Build fresh lists rather than mutating the list being iterated:
        # the previous in-place .remove()/.pop(idx) under enumerate() skipped
        # adjacent entities (classic mutate-during-iteration bug).
        kept_entities: list[dict[str, Any]] = []
        for entity in list(self.cache["entities"]):
            if "context" in entity:
                entity.pop("context", None)
                entity.pop("last_changed", None)
                entity.pop("last_reported", None)
                entity.pop("last_updated", None)

            if "attributes" in entity:
                for attr in (
                    "icon",
                    "monitor_cert_days_remaining",
                    "monitor_cert_is_valid",
                    "monitor_hostname",
                    "monitor_port",
                ):
                    entity["attributes"].pop(attr, None)

            entity_id = entity["entity_id"]
            if any(entity_id.startswith(prefix) for prefix in forbidden_prefixes):
                continue
            if any(substring in entity_id for substring in forbidden_substrings):
                continue

            if entity_id.startswith("scene."):
                entity.pop("state", None)

            if entity_id.startswith("sensor.") or entity_id.startswith("binary_sensor."):
                self.cache["sensors"].append(entity)
                continue

            kept_entities.append(entity)

        self.cache["entities"] = sort_by_entity_id(kept_entities)
        self.cache["sensors"] = sort_by_entity_id(self.cache["sensors"])
        return self.cache

    def wrapper(self: SelfT, *args: P.args, **kwargs: P.kwargs) -> R:
        """Invoke the wrapped function and normalize cache content.

        Args:
            self: Cache holder instance.
            *args: Positional arguments forwarded to the wrapped function.
            **kwargs: Keyword arguments forwarded to the wrapped function.

        Returns:
            Result of the wrapped function.
        """
        result = func(self, *args, **kwargs)

        forbidden_prefixes = [
            "alarm_control_panel.",
            "automation.",
            "binary_sensor.remote_ui",
            "camera.",
            "climate",
            "conversation",
            "device_tracker.mediacenter_raspberry_pi_5",
            "media_player.example_player",
            "media_player.example_player_2",
            "media_player.chrome",
            "media_player.fire_tv_livingroom",
            "person.",
            "remote.",
            "script.higher",
            "sensor.hacs",
            "sensor.hacs",
            "sensor.mediacenter_raspberry_pi_5_",
            "sensor.sonarr_commands",
            "sensor.sun",
            "sensor.uptimekuma_",
            "stt.",
            "sun.",
            "switch.",
            "switch.example_switch",
            "switch.bedroom_camera_camera_motion_detection",
            "tts.",
            "update.",
            "zone.home",
        ]
        forbidden_substrings = ["example_camera_motion"]
        self.cache["sensor"] = []
        self.cache = clean_entities(self, forbidden_prefixes, forbidden_substrings)

        self.cache["services"] = [
            service
            for service in self.cache["services"]
            if service["domain"] in self.cache["allowed_domains"]
        ]

        # Retrieve entity and sensor IDs
        self.cache["entity_ids"] = sorted(self.cache["entity_ids"])
        self.cache["sensor_ids"] = sorted(self.cache["sensor_ids"])

        logging.info(
            (
                "`{}` modified cache to <(len) Entity IDs: {}; (len) Entities: {}; "
                "(len) Sensors: {}; (len) Services: {};>"
            ),
            func.__name__,
            len(self.cache["entity_ids"]),
            len(self.cache["entities"]),
            len(self.cache["sensors"]),
            len(self.cache["services"]),
        )

        return result

    return wrapper

mewbo_tools.integration.lsp

Native LSP integration for Mewbo.

Provides a single lsp_tool that the agent can optionally invoke for code diagnostics, go-to-definition, find-references, and hover info. Language servers are spawned lazily and scoped per-session.

get_lsp_manager(cwd: str) -> LSPServerManager

Return (or create) the LSP manager for cwd.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/__init__.py
78
79
80
81
82
83
84
85
86
87
def get_lsp_manager(cwd: str) -> LSPServerManager:
    """Return (or create) the LSP manager for *cwd*."""
    if cwd not in _managers:
        from mewbo_core.config import get_config

        from mewbo_tools.integration.lsp.manager import LSPServerManager

        lsp_cfg = getattr(get_config().agent, "lsp", None)
        _managers[cwd] = LSPServerManager(cwd=cwd, config=lsp_cfg)
    return _managers[cwd]

get_passive_diagnostics(file_path: str, cwd: str) -> str | None

Return formatted diagnostics for file_path, or None.

Called by the tool-use loop after file edits to provide passive feedback to the LLM. Returns None if LSP is unavailable or no errors/warnings were found.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/__init__.py
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
def get_passive_diagnostics(file_path: str, cwd: str) -> str | None:
    """Return formatted diagnostics for *file_path*, or ``None``.

    Called by the tool-use loop after file edits to provide passive
    feedback to the LLM.  Returns ``None`` if LSP is unavailable or
    no errors/warnings were found.
    """
    if not LSP_AVAILABLE or not _managers:
        return None
    manager = _managers.get(cwd)
    if manager is None:
        return None
    sdef = manager.server_for_file(file_path)
    if sdef is None:
        return None
    # Only proceed if this server is already running (don't start one
    # just for passive feedback — that would slow down edits).
    if sdef.id not in manager._clients:
        return None
    try:
        client = manager._clients[sdef.id]
        # Notify the server about the changed file
        run_lsp_async(manager.open_file(client, file_path))
        # Brief pause for the server to re-analyze
        time.sleep(PASSIVE_DIAGNOSTICS_SETTLE_S)
        diags = manager.get_cached_diagnostics(file_path)
        from lsprotocol.types import DiagnosticSeverity

        errors = [
            d
            for d in diags
            if d.severity in (DiagnosticSeverity.Error, DiagnosticSeverity.Warning, None)
        ]
        if not errors:
            return None
        # Format concisely
        from mewbo_tools.integration.lsp.tool import LSPTool

        return LSPTool._format_diagnostics(file_path, diags, manager)
    except Exception:
        return None

run_lsp_async(coro: Coroutine[object, object, T], *, timeout: float = 30) -> T

Run an async coroutine on the persistent LSP event loop.

This bridges the sync tool code to the async pygls client without creating/destroying event loops on each call.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/__init__.py
60
61
62
63
64
65
66
67
68
def run_lsp_async(coro: Coroutine[object, object, T], *, timeout: float = 30) -> T:
    """Run an async coroutine on the persistent LSP event loop.

    This bridges the sync tool code to the async pygls client without
    creating/destroying event loops on each call.
    """
    loop = _get_lsp_loop()
    future = asyncio.run_coroutine_threadsafe(coro, loop)
    return future.result(timeout=timeout)

shutdown_lsp_managers() -> None async

Shut down all LSP managers. Called on session end.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/__init__.py
133
134
135
136
137
138
139
140
141
142
143
async def shutdown_lsp_managers() -> None:
    """Shut down all LSP managers.  Called on session end."""
    for mgr in list(_managers.values()):
        await mgr.shutdown_all()
    _managers.clear()
    global _lsp_loop, _lsp_thread
    with _loop_lock:
        if _lsp_loop is not None and not _lsp_loop.is_closed():
            _lsp_loop.call_soon_threadsafe(_lsp_loop.stop)
            _lsp_loop = None
            _lsp_thread = None

mewbo_tools.integration.lsp.manager

LSP server manager — lazy startup, per-session lifecycle.

Uses pygls.lsp.client.BaseLanguageClient for all protocol handling. We only manage lifecycle and map file extensions to servers.

LSPServerManager

Per-session language server manager.

Servers are started lazily on first request for a matching file type. All servers are shut down via :meth:shutdown_all on session end.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/manager.py
 28
 29
 30
 31
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
class LSPServerManager:
    """Per-session language server manager.

    Servers are started lazily on first request for a matching file type.
    All servers are shut down via :meth:`shutdown_all` on session end.
    """

    def __init__(self, cwd: str, config: Any | None = None) -> None:  # noqa: D107
        self._cwd = cwd
        self._clients: dict[str, BaseLanguageClient] = {}
        self._diagnostics: dict[str, list[types.Diagnostic]] = {}  # uri → diags
        self._failed: set[str] = set()  # server IDs that failed to start

        # Build extension map from available servers
        overrides = getattr(config, "servers", {}) if config else {}
        self._servers = available_servers(overrides)
        self._extension_map: dict[str, ServerDef] = {}
        for sdef in self._servers:
            for ext in sdef.extensions:
                self._extension_map.setdefault(ext, sdef)

        if self._servers:
            logger.info(
                "LSP servers available: {}",
                ", ".join(s.id for s in self._servers),
            )

    # ------------------------------------------------------------------
    # Public API
    # ------------------------------------------------------------------

    def server_for_file(self, file_path: str) -> ServerDef | None:
        """Return the server definition for *file_path*, or ``None``."""
        ext = Path(file_path).suffix.lower()
        return self._extension_map.get(ext)

    async def ensure_server(self, file_path: str) -> BaseLanguageClient | None:
        """Start the appropriate server for *file_path* if not running.

        Returns ``None`` if no server is available or the server is marked
        failed.
        """
        sdef = self.server_for_file(file_path)
        if sdef is None or sdef.id in self._failed:
            return None

        if sdef.id in self._clients:
            return self._clients[sdef.id]

        return await self._start_server(sdef)

    def get_cached_diagnostics(self, file_path: str) -> list[types.Diagnostic]:
        """Return diagnostics pushed by the server for *file_path*."""
        uri = Path(file_path).resolve().as_uri()
        return self._diagnostics.get(uri, [])

    async def open_file(self, client: BaseLanguageClient, file_path: str) -> None:
        """Notify the server that a file has been opened."""
        resolved = Path(file_path).resolve()
        uri = resolved.as_uri()
        sdef = self.server_for_file(file_path)
        lang_id = (
            sdef.language_id
            if sdef
            else _EXTENSION_LANGUAGE_MAP.get(resolved.suffix.lower(), "plaintext")
        )
        try:
            text = resolved.read_text(encoding="utf-8", errors="replace")
        except OSError:
            return
        client.text_document_did_open(
            types.DidOpenTextDocumentParams(
                text_document=types.TextDocumentItem(
                    uri=uri,
                    language_id=lang_id,
                    version=1,
                    text=text,
                )
            )
        )

    async def shutdown_all(self) -> None:
        """Shut down all running language servers."""
        for sid, client in list(self._clients.items()):
            try:
                await asyncio.wait_for(client.shutdown_async(None), timeout=5)
                client.exit(None)
                await client.stop()
            except Exception as exc:
                logger.debug("Error shutting down LSP server '{}': {}", sid, exc)
        self._clients.clear()
        self._diagnostics.clear()
        logger.debug("All LSP servers shut down")

    # ------------------------------------------------------------------
    # Internal
    # ------------------------------------------------------------------

    async def _start_server(self, sdef: ServerDef) -> BaseLanguageClient | None:
        """Spawn a language server and perform the LSP initialize handshake."""
        try:
            client = BaseLanguageClient(
                name=f"mewbo-{sdef.id}",
                version="0.1.0",
            )

            # Register handlers before starting
            @client.feature(types.TEXT_DOCUMENT_PUBLISH_DIAGNOSTICS)
            def _on_diagnostics(params: types.PublishDiagnosticsParams) -> None:
                self._diagnostics[params.uri] = list(params.diagnostics)

            @client.feature(types.WINDOW_LOG_MESSAGE)
            def _on_log_message(params: types.LogMessageParams) -> None:
                pass  # Suppress noisy log messages from servers

            # Resolved BEFORE the spawn, not after: it is both the workspace
            # this server will serve and the root its sandbox is keyed on.
            root_path = self._find_root(sdef)
            command, kwargs = self._launch(sdef, root_path)
            await client.start_io(command[0], *command[1:], **kwargs)

            root_uri = Path(root_path).as_uri()

            await client.initialize_async(
                types.InitializeParams(
                    process_id=os.getpid(),
                    root_uri=root_uri,
                    capabilities=types.ClientCapabilities(
                        text_document=types.TextDocumentClientCapabilities(
                            publish_diagnostics=types.PublishDiagnosticsClientCapabilities(),
                            definition=types.DefinitionClientCapabilities(),
                            references=types.ReferenceClientCapabilities(),
                            hover=types.HoverClientCapabilities(),
                        ),
                    ),
                )
            )
            client.initialized(types.InitializedParams())

            self._clients[sdef.id] = client
            logger.info("Started LSP server '{}' (root: {})", sdef.id, root_path)
            return client

        except Exception as exc:
            logger.warning("Failed to start LSP server '{}': {}", sdef.id, exc)
            self._failed.add(sdef.id)
            return None

    @staticmethod
    def _launch(sdef: ServerDef, root_path: str) -> tuple[list[str], dict[str, Any]]:
        """The argv and ``start_io`` kwargs that spawn *sdef*, confined if asked.

        ``pygls`` spawns the server itself, so there is no ``preexec_fn`` to
        hand it — the command is prefixed with the launcher shim instead, which
        applies the ruleset to itself and is then replaced by the real server.

        Scope: *root_path*, the workspace this server exists to serve. Unlike
        the process-wide MCP pool, a manager is built per session around a fixed
        ``cwd``, so the served workspace is known at spawn time and stays the
        one this server is for.

        With the sandbox off it returns ``sdef.command`` and NO kwargs, so the
        spawn is byte-identical to an unsandboxed one — passing ``env`` at all
        would replace the inherited environment rather than extend it, which is
        also why the sandboxed arm carries ``os.environ`` over explicitly.
        """
        launcher = SandboxLauncher.for_root(root_path)
        if launcher is None:
            return list(sdef.command), {}
        return launcher.command(sdef.command), {"env": launcher.environ(os.environ)}

    def _find_root(self, sdef: ServerDef) -> str:
        """Walk up from CWD to find a directory containing a root marker."""
        current = Path(self._cwd).resolve()
        for _ in range(20):  # max depth
            for marker in sdef.root_markers:
                if (current / marker).exists():
                    return str(current)
            parent = current.parent
            if parent == current:
                break
            current = parent
        return self._cwd

ensure_server(file_path: str) -> BaseLanguageClient | None async

Start the appropriate server for file_path if not running.

Returns None if no server is available or the server is marked failed.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/manager.py
64
65
66
67
68
69
70
71
72
73
74
75
76
77
async def ensure_server(self, file_path: str) -> BaseLanguageClient | None:
    """Start the appropriate server for *file_path* if not running.

    Returns ``None`` if no server is available or the server is marked
    failed.
    """
    sdef = self.server_for_file(file_path)
    if sdef is None or sdef.id in self._failed:
        return None

    if sdef.id in self._clients:
        return self._clients[sdef.id]

    return await self._start_server(sdef)

get_cached_diagnostics(file_path: str) -> list[types.Diagnostic]

Return diagnostics pushed by the server for file_path.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/manager.py
79
80
81
82
def get_cached_diagnostics(self, file_path: str) -> list[types.Diagnostic]:
    """Return diagnostics pushed by the server for *file_path*."""
    uri = Path(file_path).resolve().as_uri()
    return self._diagnostics.get(uri, [])

open_file(client: BaseLanguageClient, file_path: str) -> None async

Notify the server that a file has been opened.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/manager.py
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
async def open_file(self, client: BaseLanguageClient, file_path: str) -> None:
    """Notify the server that a file has been opened."""
    resolved = Path(file_path).resolve()
    uri = resolved.as_uri()
    sdef = self.server_for_file(file_path)
    lang_id = (
        sdef.language_id
        if sdef
        else _EXTENSION_LANGUAGE_MAP.get(resolved.suffix.lower(), "plaintext")
    )
    try:
        text = resolved.read_text(encoding="utf-8", errors="replace")
    except OSError:
        return
    client.text_document_did_open(
        types.DidOpenTextDocumentParams(
            text_document=types.TextDocumentItem(
                uri=uri,
                language_id=lang_id,
                version=1,
                text=text,
            )
        )
    )

server_for_file(file_path: str) -> ServerDef | None

Return the server definition for file_path, or None.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/manager.py
59
60
61
62
def server_for_file(self, file_path: str) -> ServerDef | None:
    """Return the server definition for *file_path*, or ``None``."""
    ext = Path(file_path).suffix.lower()
    return self._extension_map.get(ext)

shutdown_all() -> None async

Shut down all running language servers.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/manager.py
109
110
111
112
113
114
115
116
117
118
119
120
async def shutdown_all(self) -> None:
    """Shut down all running language servers."""
    for sid, client in list(self._clients.items()):
        try:
            await asyncio.wait_for(client.shutdown_async(None), timeout=5)
            client.exit(None)
            await client.stop()
        except Exception as exc:
            logger.debug("Error shutting down LSP server '{}': {}", sid, exc)
    self._clients.clear()
    self._diagnostics.clear()
    logger.debug("All LSP servers shut down")

mewbo_tools.integration.lsp.servers

Built-in language server definitions.

ServerDef dataclass

A language server that can be spawned for files matching extensions.

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/servers.py
 9
10
11
12
13
14
15
16
17
@dataclass(frozen=True)
class ServerDef:
    """A language server that can be spawned for files matching *extensions*."""

    id: str
    extensions: tuple[str, ...]
    command: tuple[str, ...]
    root_markers: tuple[str, ...]
    language_id: str  # LSP languageId string

available_servers(overrides: dict[str, dict] | None = None) -> list[ServerDef]

Return servers whose command binary is installed on the system.

overrides can disable built-in servers ({"pyright": {"disabled": true}}) or add custom ones ({"my-lsp": {"command": [...], "extensions": [...], ...}}).

Source code in packages/mewbo_tools/src/mewbo_tools/integration/lsp/servers.py
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
def available_servers(
    overrides: dict[str, dict] | None = None,
) -> list[ServerDef]:
    """Return servers whose command binary is installed on the system.

    *overrides* can disable built-in servers (``{"pyright": {"disabled": true}}``)
    or add custom ones (``{"my-lsp": {"command": [...], "extensions": [...], ...}}``).
    """
    overrides = overrides or {}
    result: list[ServerDef] = []

    for server in BUILTIN_SERVERS:
        ovr = overrides.get(server.id, {})
        if ovr.get("disabled"):
            continue
        if shutil.which(server.command[0]):
            result.append(server)

    # User-defined servers
    for name, cfg in overrides.items():
        if any(s.id == name for s in BUILTIN_SERVERS):
            continue  # already handled above
        if cfg.get("disabled"):
            continue
        cmd = cfg.get("command")
        exts = cfg.get("extensions")
        if not cmd or not exts:
            continue
        cmd_tuple = tuple(cmd) if isinstance(cmd, list) else (cmd,)
        if not shutil.which(cmd_tuple[0]):
            continue
        result.append(
            ServerDef(
                id=name,
                extensions=tuple(exts),
                command=cmd_tuple,
                root_markers=tuple(cfg.get("root_markers", [])),
                language_id=cfg.get("language_id", name),
            )
        )

    return result

packages/mewbo_graph (knowledge-graph capability library)

The optional substrate shared by Agentic Wiki and Mewbo Search. Requires the library extras (treesitter, retrieval); absent when uninstalled.

Agentic Wiki substrate (mewbo_graph.wiki)

mewbo_graph.wiki.graph

Wiki code-graph module.

Two atomic classes share the same domain (the code knowledge graph) at different phases:

  • GraphIndex runs at indexing time — tree-sitter-driven AST extraction yielding flat GraphNode + GraphEdge lists that the store persists.
  • KnowledgeGraphView runs at view time — loads a slug's persisted nodes
  • edges, computes lightweight stats, and serialises a Cytoscape-friendly wire shape for the /v1/wiki/projects/<slug>/graph endpoint.

Keeping both in one module avoids splitting the same domain across files; each class owns its own state and behaviour over that state.

GraphIndex

AST-graph extractor.

Constructed once per wiki indexing job. Caches loaded languages and compiled queries so per-file parse is a hot-loop friendly call.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/graph.py
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
class GraphIndex:
    """AST-graph extractor.

    Constructed once per wiki indexing job. Caches loaded languages and
    compiled queries so per-file parse is a hot-loop friendly call.
    """

    def __init__(self) -> None:
        """Initialise caches and verify that the wiki extras are installed."""
        self._lang_cache: dict[str, object] = {}   # name → tree_sitter.Language
        self._query_cache: dict[str, object] = {}  # name → tree_sitter.Query
        # Anchored on the PACKAGE, not on this file's depth: ``graph_queries/``
        # belongs to ``mewbo_graph.wiki``, so a __file__-relative path would
        # silently follow this module if it ever moves and fail at query-load
        # time rather than at import.
        self._queries_dir = resources.files("mewbo_graph.wiki") / "graph_queries"
        # Defensive import so missing extras give a clean error.
        try:
            import tree_sitter  # noqa: F401
            import tree_sitter_language_pack  # noqa: F401
        except ImportError as exc:
            raise ImportError(
                "GraphIndex requires the 'wiki' extras: install with "
                "`uv sync --extra wiki` or `pip install mewbo-api[wiki]`."
            ) from exc

    def parse_file(
        self, slug: str, file_path: Path, *, repo_root: Path
    ) -> GraphParseResult:
        """Parse a single file.

        Returns an empty result — the path lands in ``skipped``, same as an
        unsupported extension — for a file with no supported extension, one
        under a vendored directory, or one that looks minified. This is the
        ONE seam every caller funnels through (``parse_repo``, the agent-driven
        ``wiki_build_graph`` tool, and the zero-LLM ``GraphOnlyIndexer`` each
        build their own file list independently and never share a walk), so
        gating it here is what makes the exclusion apply everywhere rather
        than needing to be re-applied at each caller's own file-collection
        site.
        """
        ext = file_path.suffix.lower()
        lang_name = _LANG_BY_EXT.get(ext)
        rel_path = file_path.relative_to(repo_root)
        if lang_name is None or _is_vendored_path(rel_path):
            return GraphParseResult(nodes=[], edges=[], skipped=[str(rel_path)])

        from tree_sitter import Parser

        source = file_path.read_bytes()
        if _is_minified(rel_path, source):
            return GraphParseResult(nodes=[], edges=[], skipped=[str(rel_path)])
        lang = self._lang_for(lang_name)
        tree = Parser(lang).parse(source)
        query = self._query_for(lang_name, lang)
        from tree_sitter import QueryCursor

        cursor = QueryCursor(query)
        # ``captures()`` groups nodes per capture NAME but the per-name lists are
        # NOT mutually index-aligned (order varies once matches nest), so
        # ``_extract`` re-aligns each parallel ``.def``/``.name`` pair by start
        # byte via ``_by_position`` before zipping — see that helper.
        captures: dict[str, list] = cursor.captures(tree.root_node)

        rel = str(file_path.relative_to(repo_root))
        return _extract(slug, rel, source, captures)

    def parse_repo(
        self,
        slug: str,
        repo_root: Path,
        files: list[Path],
        *,
        on_progress: Callable[[int, int, str], None] | None = None,
    ) -> GraphParseResult:
        """Parse every file. Files outside the supported set go to ``skipped``.

        ``on_progress(done, total, path)`` is called after each file when
        supplied. It is INJECTED rather than written here because progress is
        persisted against an indexing job, which this parser knows nothing about
        — and because the loop is the only place that knows how far along it is.
        The throttling is the callee's business (see ``PhaseProgress``); this
        loop reports every file and never decides what is worth writing.
        """
        result = GraphParseResult(nodes=[], edges=[], skipped=[])
        total = len(files)
        for done, fp in enumerate(files, start=1):
            result += self.parse_file(slug, fp, repo_root=repo_root)
            if on_progress is not None:
                on_progress(done, total, str(fp.relative_to(repo_root)))
        return result

    # ── helpers ───────────────────────────────────────────────────────────────

    def _lang_for(self, lang_name: str) -> Any:
        """Return a cached ``tree_sitter.Language`` for *lang_name*."""
        if lang_name not in self._lang_cache:
            self._lang_cache[lang_name] = _load_ts_language(lang_name)
        return self._lang_cache[lang_name]

    def _query_for(self, lang_name: str, lang: Any) -> Any:
        """Return a cached compiled ``tree_sitter.Query`` for *lang_name*."""
        if lang_name not in self._query_cache:
            from tree_sitter import Query

            filename = _SPEC_BY_NAME[lang_name].query_filename
            scm = (self._queries_dir / filename).read_text(encoding="utf-8")
            self._query_cache[lang_name] = Query(lang, scm)
        return self._query_cache[lang_name]

__init__() -> None

Initialise caches and verify that the wiki extras are installed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/graph.py
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
def __init__(self) -> None:
    """Initialise caches and verify that the wiki extras are installed."""
    self._lang_cache: dict[str, object] = {}   # name → tree_sitter.Language
    self._query_cache: dict[str, object] = {}  # name → tree_sitter.Query
    # Anchored on the PACKAGE, not on this file's depth: ``graph_queries/``
    # belongs to ``mewbo_graph.wiki``, so a __file__-relative path would
    # silently follow this module if it ever moves and fail at query-load
    # time rather than at import.
    self._queries_dir = resources.files("mewbo_graph.wiki") / "graph_queries"
    # Defensive import so missing extras give a clean error.
    try:
        import tree_sitter  # noqa: F401
        import tree_sitter_language_pack  # noqa: F401
    except ImportError as exc:
        raise ImportError(
            "GraphIndex requires the 'wiki' extras: install with "
            "`uv sync --extra wiki` or `pip install mewbo-api[wiki]`."
        ) from exc

parse_file(slug: str, file_path: Path, *, repo_root: Path) -> GraphParseResult

Parse a single file.

Returns an empty result — the path lands in skipped, same as an unsupported extension — for a file with no supported extension, one under a vendored directory, or one that looks minified. This is the ONE seam every caller funnels through (parse_repo, the agent-driven wiki_build_graph tool, and the zero-LLM GraphOnlyIndexer each build their own file list independently and never share a walk), so gating it here is what makes the exclusion apply everywhere rather than needing to be re-applied at each caller's own file-collection site.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/graph.py
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
def parse_file(
    self, slug: str, file_path: Path, *, repo_root: Path
) -> GraphParseResult:
    """Parse a single file.

    Returns an empty result — the path lands in ``skipped``, same as an
    unsupported extension — for a file with no supported extension, one
    under a vendored directory, or one that looks minified. This is the
    ONE seam every caller funnels through (``parse_repo``, the agent-driven
    ``wiki_build_graph`` tool, and the zero-LLM ``GraphOnlyIndexer`` each
    build their own file list independently and never share a walk), so
    gating it here is what makes the exclusion apply everywhere rather
    than needing to be re-applied at each caller's own file-collection
    site.
    """
    ext = file_path.suffix.lower()
    lang_name = _LANG_BY_EXT.get(ext)
    rel_path = file_path.relative_to(repo_root)
    if lang_name is None or _is_vendored_path(rel_path):
        return GraphParseResult(nodes=[], edges=[], skipped=[str(rel_path)])

    from tree_sitter import Parser

    source = file_path.read_bytes()
    if _is_minified(rel_path, source):
        return GraphParseResult(nodes=[], edges=[], skipped=[str(rel_path)])
    lang = self._lang_for(lang_name)
    tree = Parser(lang).parse(source)
    query = self._query_for(lang_name, lang)
    from tree_sitter import QueryCursor

    cursor = QueryCursor(query)
    # ``captures()`` groups nodes per capture NAME but the per-name lists are
    # NOT mutually index-aligned (order varies once matches nest), so
    # ``_extract`` re-aligns each parallel ``.def``/``.name`` pair by start
    # byte via ``_by_position`` before zipping — see that helper.
    captures: dict[str, list] = cursor.captures(tree.root_node)

    rel = str(file_path.relative_to(repo_root))
    return _extract(slug, rel, source, captures)

parse_repo(slug: str, repo_root: Path, files: list[Path], *, on_progress: Callable[[int, int, str], None] | None = None) -> GraphParseResult

Parse every file. Files outside the supported set go to skipped.

on_progress(done, total, path) is called after each file when supplied. It is INJECTED rather than written here because progress is persisted against an indexing job, which this parser knows nothing about — and because the loop is the only place that knows how far along it is. The throttling is the callee's business (see PhaseProgress); this loop reports every file and never decides what is worth writing.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/graph.py
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
def parse_repo(
    self,
    slug: str,
    repo_root: Path,
    files: list[Path],
    *,
    on_progress: Callable[[int, int, str], None] | None = None,
) -> GraphParseResult:
    """Parse every file. Files outside the supported set go to ``skipped``.

    ``on_progress(done, total, path)`` is called after each file when
    supplied. It is INJECTED rather than written here because progress is
    persisted against an indexing job, which this parser knows nothing about
    — and because the loop is the only place that knows how far along it is.
    The throttling is the callee's business (see ``PhaseProgress``); this
    loop reports every file and never decides what is worth writing.
    """
    result = GraphParseResult(nodes=[], edges=[], skipped=[])
    total = len(files)
    for done, fp in enumerate(files, start=1):
        result += self.parse_file(slug, fp, repo_root=repo_root)
        if on_progress is not None:
            on_progress(done, total, str(fp.relative_to(repo_root)))
    return result

GraphParseResult dataclass

Output of a single-file or repo parse.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/graph.py
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
@dataclass(frozen=True)
class GraphParseResult:
    """Output of a single-file or repo parse."""

    nodes: list[GraphNode]
    edges: list[GraphEdge]
    skipped: list[str]  # files whose extension isn't supported

    def __add__(self, other: GraphParseResult) -> GraphParseResult:
        """Merge two results by concatenating their node, edge, and skipped lists."""
        return GraphParseResult(
            nodes=self.nodes + other.nodes,
            edges=self.edges + other.edges,
            skipped=self.skipped + other.skipped,
        )

__add__(other: GraphParseResult) -> GraphParseResult

Merge two results by concatenating their node, edge, and skipped lists.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/graph.py
159
160
161
162
163
164
165
def __add__(self, other: GraphParseResult) -> GraphParseResult:
    """Merge two results by concatenating their node, edge, and skipped lists."""
    return GraphParseResult(
        nodes=self.nodes + other.nodes,
        edges=self.edges + other.edges,
        skipped=self.skipped + other.skipped,
    )

KnowledgeGraphView dataclass

Slug-scoped projection of the persisted MULTIPLEX graph for the viewer.

Three node layers share one viewer payload: the tree-sitter ast layer (File/Class/Function/… + synthesized External convergence nodes), the abstract entity layer, and the atomic-note memory layer. Edges carry a layer tag — ast (CONTAINS/IMPORTS/CALLS/EXTENDS/REFERENCES), entity (entity↔entity RELATES, open-vocab verb in label), memory (note RELATES) and cross (ANCHORS spanning layers).

Construction is exclusively via for_slug so the invariant — every node + edge belongs to the same slug, and every emitted edge endpoint is a real node in the payload — stays enforced in one place. Once built, the instance is immutable; safe to share across requests and trivially cacheable upstream.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/graph.py
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
@dataclass(frozen=True, slots=True)
class KnowledgeGraphView:
    """Slug-scoped projection of the persisted MULTIPLEX graph for the viewer.

    Three node layers share one viewer payload: the tree-sitter ``ast`` layer
    (File/Class/Function/… + synthesized ``External`` convergence nodes), the
    abstract ``entity`` layer, and the atomic-note ``memory`` layer. Edges carry
    a ``layer`` tag — ``ast`` (CONTAINS/IMPORTS/CALLS/EXTENDS/REFERENCES),
    ``entity`` (entity↔entity RELATES, open-vocab verb in ``label``), ``memory``
    (note RELATES) and ``cross`` (ANCHORS spanning layers).

    Construction is exclusively via ``for_slug`` so the invariant — every node +
    edge belongs to the same slug, and every emitted edge endpoint is a real
    node in the payload — stays enforced in one place. Once built, the instance
    is immutable; safe to share across requests and trivially cacheable upstream.
    """

    slug: str
    # AST layer (real in-repo nodes only) + synthesized External convergence
    # nodes (one per distinct unresolved cross-file symbol name).
    nodes: tuple[GraphNode, ...]
    external_nodes: tuple[GraphNode, ...]
    edges: tuple[GraphEdge, ...]  # ast-layer edges (endpoints all in payload)
    # Entity layer.
    entity_nodes: tuple[Entity, ...]
    entity_edges: tuple[EntityRelation, ...]  # entity↔entity RELATES only
    # Memory layer.
    memory_nodes: tuple[MemoryNode, ...]
    memory_edges: tuple[MemoryEdge, ...]  # note↔note RELATES only
    # Cross-layer ANCHORS, pre-resolved to ``(source_node_id, target_node_id)``.
    cross_edges: tuple[tuple[str, str], ...]
    total_nodes: int  # full AST node count (pre-cap), for the "showing N of M" banner
    total_edges: int  # full AST edge count (pre-cap)
    # Externals the payload COULD have carried, before the cap took its share.
    # Distinct from ``total_nodes`` because externals are synthesized at read
    # time rather than stored, and the cap can drop them while leaving every
    # real AST node in place — without this the banner reports "showing N of N"
    # for a payload that silently lost nodes.
    total_externals: int = 0
    # Directory scaffold for the "hierarchy" wire mode — ``None`` in the default
    # mode, so ``to_wire`` stays byte-identical when hierarchy is off.
    folder_tree: FolderTree | None = None

    # ── Construction ────────────────────────────────────────────────────

    @classmethod
    def _reanchor_entity_edges(
        cls,
        store: WikiStoreBase,
        slug: str,
        stranded: list[EntityRelation],
        live_nodes: list[GraphNode],
        payload_ast_ids: set[str],
    ) -> list[tuple[str, str]]:
        """Re-point entity ANCHORS that address a superseded generation's nodes.

        An entity anchors to a code node by RAW ``node_id``, and a node id
        embeds the symbol's ``start_byte`` (``_stable_id``). So every edit that
        shifts a symbol within its file re-keys it, and an anchor minted
        against an earlier index addresses an id the current generation does
        not contain. Unscoped that never showed, because the superseded node
        was still in the payload to point at — which is exactly the bug commit
        scoping fixes, and exactly why scoping alone would silently sever most
        of the entity→code bridge.

        The repair is by ``entity_key`` (``file#name``), which carries NO byte
        offset and so survives the edit that broke the id. Cost is bounded by
        the number of stranded anchors, not by the graph: the superseded nodes
        are fetched by explicit id rather than by reading the union.

        The memory layer needs none of this — a memory ANCHORS edge already
        stores an ``entity_key`` as its target rather than a node id, so it
        re-resolves against whatever generation is loaded.

        An anchor whose symbol genuinely no longer exists resolves to nothing
        and stays dropped. That is the correct outcome, not a loss: an entity
        pointing at deleted code is what this phase set out to stop rendering.
        """
        wanted = {rel.target_id for rel in stranded}
        if not wanted:
            return []
        live_by_key: dict[str, str] = {}
        for n in live_nodes:
            live_by_key.setdefault(entity_key_for_node(n), n.node_id)
        superseded = store.query_graph(
            slug, scope=CommitScope.every(), node_ids=wanted
        )
        repointed: dict[str, str] = {}
        for n in superseded:
            live_id = live_by_key.get(entity_key_for_node(n))
            if live_id is not None and live_id in payload_ast_ids:
                repointed[n.node_id] = live_id
        return [
            (rel.source_id, repointed[rel.target_id])
            for rel in stranded
            if rel.target_id in repointed
        ]

    @classmethod
    def for_slug(
        cls,
        store: WikiStoreBase,
        slug: str,
        *,
        node_limit: int | None = None,
        hierarchy: bool = False,
        scope: CommitScope | None = None,
    ) -> KnowledgeGraphView:
        """Load the full multiplex (ast + entity + memory layers) for *slug*.

        **Commit-scoped.** *scope* defaults to ``store.live_scope(slug)`` — the
        project's own commit — so the viewer shows the code as it is now. Before
        this the AST layer was read unscoped, which is why the payload was the
        union of every generation ever indexed: deleted code rendered alongside
        live code with nothing distinguishing them, and the node count grew
        monotonically with re-indexes rather than with the repository.

        Pass an explicit scope to read a specific generation; pass
        ``CommitScope.every()`` to restore the pre-scoping union (which the
        entity-anchor repair below relies on, and nothing else should).

        AST connectivity: a cross-file IMPORTS/CALLS/EXTENDS edge carries the
        raw ``target_name``; if that name resolves to a real in-repo node it is
        re-pointed there (genuinely connecting File-clusters through shared
        symbols), otherwise every reference to the same external name converges
        on ONE synthesized ``External`` view-node.

        ``node_limit`` (when set and exceeded) degree-prunes the **whole AST
        payload** — real nodes AND the view-synthesized ``External`` nodes
        together, not the persisted layer alone. The AST nodes are pruned
        first (unchanged from before: highest-degree kept, ties on
        ``node_id``); External nodes then get whatever budget is left,
        pruned by the same degree-then-``node_id`` rule over their surviving
        edges. A ``node_limit`` that the AST layer alone does not exceed can
        still cap externals down, and an AST layer that exhausts the whole
        budget leaves externals at zero — both are the cap "governing the
        whole payload" rather than only the persisted one. Entity + memory
        layers are always fully included (they're small) and never counted
        against this cap. ``total_nodes``/``total_edges`` always reflect the
        full AST graph so the wire response can honestly say "showing N of M".

        Cross-layer ANCHORS are reconciled to real node ids in O(nodes+edges):
        memory ANCHORS targets (``EntityKey`` / ``entity:<id>``) batch-resolve
        through the existing ``CodeStructureProvider`` + ``EntityAnchorResolver``;
        an anchor that resolves to nothing is dropped (no dangling edges).

        ``hierarchy`` (default off) synthesises a directory scaffold via
        ``FolderTree`` from the kept File nodes' paths — folder nodes + folder
        ``CONTAINS`` edges, and a single ``parentId`` per node — and stamps it
        onto the wire in ``to_wire``. Off ⇒ the wire is byte-identical to the
        default mode (the SCG / Agentic Search reuse path is undisturbed).
        """
        from mewbo_graph.entities.anchor import EntityAnchorResolver  # noqa: PLC0415

        from .structure_provider import CodeStructureProvider  # noqa: PLC0415

        # ── AST layer ────────────────────────────────────────────────────
        scope = store.live_scope(slug) if scope is None else scope
        all_nodes = store.query_graph(slug, scope=scope)
        all_edges = list(store.list_edges(slug, scope=scope))
        total_nodes = len(all_nodes)

        # Resolve cross-file edge targets by name → real in-repo node. Build the
        # name index once (O(nodes)); externals converge on one synthesized node.
        by_name: dict[str, str] = {}
        for n in all_nodes:
            by_name.setdefault(n.name, n.node_id)
        resolved_edges, external_nodes = cls._resolve_ast_edges(
            slug, all_edges, by_name, {n.node_id for n in all_nodes}
        )
        # ``total_edges`` is the FULL emittable AST edge count — measured AFTER
        # resolution, because that pass drops orphan structural edges (a
        # ``target_name=None`` edge with a missing endpoint). Using the raw
        # ``len(all_edges)`` would overstate "M" by the orphans that never reach
        # the payload. (Resolution adds External nodes but never adds/drops
        # cross-file edges, so this count is cap-independent.)
        total_edges = len(resolved_edges)

        # Degree over the full resolved edge set — needed by the AST prune
        # below (when it fires) and by the external cap that follows it, so
        # it's computed once whenever a limit is in play at all. Left EMPTY
        # when nothing is capped: both readers are behind the same
        # ``node_limit is not None`` guard, and a Counter answers 0 for an
        # absent key, so an unlimited read can never see a wrong rank.
        degree: Counter[str] = Counter()
        if node_limit is not None:
            for e in resolved_edges:
                degree[e.source] += 1
                degree[e.target] += 1

        # Degree-prune the AST layer (externals follow their surviving edge,
        # then get their own cap below).
        if node_limit is None or total_nodes <= node_limit:
            nodes = list(all_nodes)
        else:
            nodes = sorted(
                all_nodes, key=lambda n: (-degree[n.node_id], n.node_id)
            )[:node_limit]

        kept_ast_ids = {n.node_id for n in nodes}
        ext_by_id = {n.node_id: n for n in external_nodes}
        edges = [
            e
            for e in resolved_edges
            if e.source in kept_ast_ids
            and (e.target in kept_ast_ids or e.target in ext_by_id)
        ]
        # Keep only externals still referenced by a surviving edge.
        live_ext_ids = {e.target for e in edges if e.target in ext_by_id}

        # ``node_limit`` caps the WHOLE payload, not just the persisted AST
        # layer — a view-synthesized External node still counts against it.
        # Externals get whatever budget the AST prune above left behind (zero
        # when that prune already spent the full cap), ranked by the SAME
        # degree-then-``node_id`` rule so the tie-break policy is one rule,
        # not two. Final tuple is re-sorted by ``node_id`` for stable output,
        # matching the AST layer's own ordering guarantee.
        if node_limit is None:
            live_externals = [ext_by_id[i] for i in live_ext_ids]
        else:
            budget = max(0, node_limit - len(nodes))
            live_externals = sorted(
                (ext_by_id[i] for i in live_ext_ids),
                key=lambda n: (-degree[n.node_id], n.node_id),
            )[:budget]
        kept_externals = tuple(sorted(live_externals, key=lambda n: n.node_id))
        kept_ext_ids = {n.node_id for n in kept_externals}
        if kept_ext_ids != live_ext_ids:
            # The external cap dropped some of the externals ``edges`` above
            # was built against — drop the now-dangling edges pointing at
            # them so no edge survives whose target was capped away.
            edges = [
                e
                for e in edges
                if e.target in kept_ast_ids or e.target in kept_ext_ids
            ]
        payload_ast_ids = kept_ast_ids | kept_ext_ids

        # ── Entity layer ─────────────────────────────────────────────────
        entity_nodes = store.query_entities(slug)
        entity_ids = {e.id for e in entity_nodes}
        entity_rels: list[EntityRelation] = []
        entity_cross: list[tuple[str, str]] = []
        stranded: list[EntityRelation] = []
        for rel in store.list_entity_edges(slug):
            if rel.target_id in entity_ids and rel.source_id in entity_ids:
                # entity ↔ entity → RELATES (verb in label)
                entity_rels.append(rel)
            elif rel.source_id in entity_ids and rel.target_id in payload_ast_ids:
                # entity → AST node → cross-layer ANCHORS
                entity_cross.append((rel.source_id, rel.target_id))
            elif rel.source_id in entity_ids:
                # Anchored at an AST node that is not in this payload. Under a
                # commit scope that is usually not a dangling anchor but a
                # SUPERSEDED one — see the repair below.
                stranded.append(rel)
            # else: dangling (target absent from this payload) → dropped
        entity_cross.extend(
            cls._reanchor_entity_edges(store, slug, stranded, nodes, payload_ast_ids)
        )

        # ── Memory layer ─────────────────────────────────────────────────
        memory_nodes = store.query_memory(slug)
        memory_ids = {n.node_id for n in memory_nodes}
        mem_rels: list[MemoryEdge] = []
        mem_anchor_edges: list[MemoryEdge] = []
        for me in store.list_memory_edges(slug, include_invalidated=False):
            if me.type == "RELATES" and me.source in memory_ids:
                mem_rels.append(me)
            elif me.type == "ANCHORS" and me.source in memory_ids:
                mem_anchor_edges.append(me)

        # Batch-resolve memory ANCHORS targets to real node ids (one pass each).
        code_keys = [e.target for e in mem_anchor_edges if not e.target.startswith("entity:")]
        ent_keys = [e.target for e in mem_anchor_edges if e.target.startswith("entity:")]
        code_map = CodeStructureProvider(store).resolve_many(slug, code_keys)
        ent_map = EntityAnchorResolver(store).resolve_many(slug, ent_keys)
        memory_cross: list[tuple[str, str]] = []
        for me in mem_anchor_edges:
            if me.target.startswith("entity:"):
                ent = ent_map.get(me.target)
                tid = ent.id if ent is not None and ent.id in entity_ids else None
            else:
                node = code_map.get(me.target)
                tid = node.node_id if node is not None and node.node_id in payload_ast_ids else None
            if tid is not None:
                memory_cross.append((me.source, tid))

        # ── Directory scaffold (hierarchy wire mode only) ────────────────
        # Built over the KEPT AST layer so folder nodes only scaffold files
        # actually in the payload; the symbol→container parentId reads off the
        # same filtered CONTAINS edges the wire emits.
        folder_tree = (
            FolderTree.build(slug, nodes, edges, external_nodes=kept_externals)
            if hierarchy
            else None
        )

        return cls(
            slug=slug,
            nodes=tuple(nodes),
            external_nodes=kept_externals,
            edges=tuple(edges),
            entity_nodes=tuple(entity_nodes),
            entity_edges=tuple(entity_rels),
            memory_nodes=tuple(memory_nodes),
            memory_edges=tuple(mem_rels),
            cross_edges=tuple(entity_cross + memory_cross),
            total_nodes=total_nodes,
            total_edges=total_edges,
            total_externals=len(live_ext_ids),
            folder_tree=folder_tree,
        )

    @staticmethod
    def _resolve_ast_edges(
        slug: str,
        all_edges: list[GraphEdge],
        by_name: dict[str, str],
        node_ids: set[str],
    ) -> tuple[list[GraphEdge], list[GraphNode]]:
        """Re-point cross-file edges to real nodes; synthesize External nodes.

        An edge with a ``target_name`` (cross-file IMPORTS/CALLS/EXTENDS) is
        re-pointed at the in-repo node of that name when one exists; otherwise
        every reference to the same name converges on one synthesized
        ``External`` node (deterministic id over the name). Edges WITHOUT a
        ``target_name`` (CONTAINS etc.) pass through only when both endpoints are
        real nodes — orphan hygiene unchanged from the prior behaviour.
        """
        resolved: list[GraphEdge] = []
        externals: dict[str, GraphNode] = {}
        for e in all_edges:
            if e.target_name is None:
                # In-repo structural edge; keep only if both endpoints are real.
                if e.source in node_ids and e.target in node_ids:
                    resolved.append(e)
                continue
            real = by_name.get(e.target_name)
            if real is not None and real != e.source:
                resolved.append(e.model_copy(update={"target": real}))
            else:
                ext_id = _stable_id(slug, "External", e.target_name, "<external>", 0)
                externals.setdefault(
                    ext_id,
                    ExternalNode(
                        slug=slug,
                        node_id=ext_id,
                        name=e.target_name,
                        file="",
                        range=(0, 0),
                        docstring=None,
                    ),
                )
                resolved.append(e.model_copy(update={"target": ext_id}))
        return resolved, list(externals.values())

    # ── Derived state ───────────────────────────────────────────────────

    @property
    def node_count(self) -> int:
        """AST nodes in this view (real + External), drives the truncation banner."""
        return len(self.nodes) + len(self.external_nodes)

    @property
    def edge_count(self) -> int:
        """AST-layer edges in this view (post-filter)."""
        return len(self.edges)

    @property
    def kinds(self) -> dict[str, int]:
        """Per-kind histogram over EVERY node this view puts on the wire.

        Counts the Entity, Memory and Folder classes alongside the AST kinds
        because the FE derives its whole legend from this one map: it lists a
        kind only when the tally is positive, and sums the same map to decide
        which LAYERS to offer. Counting the AST layer alone therefore did not
        merely under-report — it drew those three classes on the canvas with no
        legend entry and no toggle, so a user could neither identify them nor
        turn them off.

        Folder is counted only in the hierarchy wire mode, which is the only
        mode that emits Folder nodes.
        """
        counts = (
            Counter(n.type for n in self.nodes)
            + Counter(n.type for n in self.external_nodes)
            + Counter({"Entity": len(self.entity_nodes)} if self.entity_nodes else {})
            + Counter({"Memory": len(self.memory_nodes)} if self.memory_nodes else {})
        )
        if self.folder_tree is not None and self.folder_tree.folder_nodes:
            counts += Counter({"Folder": len(self.folder_tree.folder_nodes)})
        return dict(counts)

    # ── Serialisation ───────────────────────────────────────────────────

    def to_wire(self) -> dict[str, Any]:
        """Wire shape: ``{slug, nodes, edges, stats}`` over all three layers.

        Each node/edge is Cytoscape-ready (``{data: {...}}``) and carries a
        ``layer`` tag so the FE can style/filter per layer. The FE hands the
        arrays straight to ``cy.add(elements)`` with no intermediate transform.

        When the view was built with ``hierarchy=True`` the directory scaffold
        is merged in: synthesised ``Folder`` nodes + their ``CONTAINS`` edges
        are appended, and every emitted node gains a ``parentId`` /
        ``folderPath`` (``folderCount`` lands in ``stats``). With hierarchy off
        the ``folder_tree`` is ``None`` and the payload is byte-identical to the
        default mode.
        """
        tree = self.folder_tree
        nodes = (
            [self._node_to_wire(n, "ast", tree) for n in self.nodes]
            + [self._node_to_wire(n, "ast", tree) for n in self.external_nodes]
            + [self._entity_node_to_wire(e, tree) for e in self.entity_nodes]
            + [self._memory_node_to_wire(m, tree) for m in self.memory_nodes]
        )
        edges = (
            [self._edge_to_wire(e, "ast") for e in self.edges]
            + [self._entity_edge_to_wire(e) for e in self.entity_edges]
            + [self._memory_edge_to_wire(e) for e in self.memory_edges]
            + [self._cross_edge_to_wire(s, t) for s, t in self.cross_edges]
        )
        if tree is not None:
            # Folder nodes carry their own parentId/folderPath off the tree.
            nodes += [self._node_to_wire(f, "ast", tree) for f in tree.folder_nodes]
            edges += [self._edge_to_wire(e, "ast") for e in tree.folder_edges]
        stats: dict[str, Any] = {
            # AST-only counters — the flat shape wire consumers read.
            "nodeCount": self.node_count,
            "edgeCount": self.edge_count,
            "kinds": self.kinds,
            # FULL multiplex counts ("M" in the FE "showing N of M" banner).
            # Only the AST layer is ever capped, so entity + memory + the
            # view-only External nodes contribute their in-view counts; the
            # pre-cap AST total is ``self.total_nodes``. Uncapped ⇒ M == N.
            "totalNodes": (
                self.total_nodes
                + self.total_externals
                + len(self.entity_nodes)
                + len(self.memory_nodes)
            ),
            "totalEdges": (
                self.total_edges
                + len(self.entity_edges)
                + len(self.memory_edges)
                + len(self.cross_edges)
            ),
            # ``truncated`` answers "did the cap drop anything the payload
            # would otherwise carry", and the cap governs BOTH populations —
            # so both are compared against their own pre-cap total. The two
            # counts stay separate rather than summed: summing lets a surplus
            # on one side mask a shortfall on the other, which is how a
            # payload that had lost a quarter of its nodes still reported
            # itself complete. Entity + memory layers are never capped, and
            # an edge dropped by orphan hygiene is not truncation.
            "truncated": (
                len(self.nodes) < self.total_nodes
                or len(self.external_nodes) < self.total_externals
            ),
            "perLayer": {
                "ast": len(self.nodes) + len(self.external_nodes),
                "entity": len(self.entity_nodes),
                "memory": len(self.memory_nodes),
            },
        }
        # ``folderCount`` only in hierarchy mode — the default wire stays
        # byte-identical (the SCG / Agentic Search reuse path is undisturbed).
        if tree is not None:
            stats["folderCount"] = len(tree.folder_nodes)
        return {
            "slug": self.slug,
            "nodes": nodes,
            "edges": edges,
            "stats": stats,
        }

    # ── Static helpers (per-record formatters) ──────────────────────────

    @staticmethod
    def _stamp_hierarchy(
        data: dict[str, Any], node_id: str, tree: FolderTree | None
    ) -> dict[str, Any]:
        """Add ``parentId``/``folderPath`` to a node's data when in hierarchy mode.

        No-op (identity) off-mode so the default wire keeps NEITHER key and
        stays byte-identical. ``folderPath`` is set only for Folder + File node
        ids (per the tree's map); every other node carries an explicit ``None``.
        """
        if tree is None:
            return data
        data["parentId"] = tree.parent_of(node_id)
        data["folderPath"] = tree.folder_path_of(node_id)
        return data

    @classmethod
    def _node_to_wire(
        cls, n: GraphNode, layer: str, tree: FolderTree | None = None
    ) -> dict[str, Any]:
        data: dict[str, Any] = {
            "id": n.node_id,
            "label": n.name,
            "kind": n.type,
            "layer": layer,
            "file": n.file,
            "range": list(n.range),
            "docstring": n.docstring or "",
        }
        if n.subkind is not None:
            data["subkind"] = n.subkind
        return {"data": cls._stamp_hierarchy(data, n.node_id, tree)}

    @classmethod
    def _entity_node_to_wire(
        cls, e: Entity, tree: FolderTree | None = None
    ) -> dict[str, Any]:
        return {
            "data": cls._stamp_hierarchy(
                {
                    "id": e.id,
                    "label": e.name,
                    "kind": "Entity",
                    "layer": "entity",
                    "entityType": e.type,
                    "labels": list(e.labels),
                },
                e.id,
                tree,
            ),
        }

    @classmethod
    def _memory_node_to_wire(
        cls, m: MemoryNode, tree: FolderTree | None = None
    ) -> dict[str, Any]:
        content = m.content.strip()
        label = content[:_MEMORY_LABEL_CHARS]
        if len(content) > _MEMORY_LABEL_CHARS:
            label += "…"
        return {
            "data": cls._stamp_hierarchy(
                {
                    "id": m.node_id,
                    "label": label,
                    "kind": "Memory",
                    "layer": "memory",
                    "snippet": content[:_MEMORY_SNIPPET_CHARS],
                    "labels": list(m.labels),
                },
                m.node_id,
                tree,
            ),
        }

    @staticmethod
    def _edge_to_wire(e: GraphEdge, layer: str) -> dict[str, Any]:
        return {
            "data": {
                "id": f"{e.source}__{e.type}__{e.target}",
                "source": e.source,
                "target": e.target,
                "kind": e.type,
                "layer": layer,
            },
        }

    @staticmethod
    def _entity_edge_to_wire(e: EntityRelation) -> dict[str, Any]:
        # Open-vocab relation verb rides ``label`` (NOT ``kind``) so the FE's
        # closed kind-union stays closed — ``RELATES`` is the only entity kind.
        return {
            "data": {
                "id": f"{e.source_id}__RELATES__{e.target_id}__{e.type}",
                "source": e.source_id,
                "target": e.target_id,
                "kind": "RELATES",
                "layer": "entity",
                "label": e.type,
            },
        }

    @staticmethod
    def _memory_edge_to_wire(e: MemoryEdge) -> dict[str, Any]:
        return {
            "data": {
                "id": f"{e.source}__RELATES__{e.target}",
                "source": e.source,
                "target": e.target,
                "kind": "RELATES",
                "layer": "memory",
            },
        }

    @staticmethod
    def _cross_edge_to_wire(source: str, target: str) -> dict[str, Any]:
        return {
            "data": {
                "id": f"{source}__ANCHORS__{target}",
                "source": source,
                "target": target,
                "kind": "ANCHORS",
                "layer": "cross",
            },
        }

edge_count: int property

AST-layer edges in this view (post-filter).

kinds: dict[str, int] property

Per-kind histogram over EVERY node this view puts on the wire.

Counts the Entity, Memory and Folder classes alongside the AST kinds because the FE derives its whole legend from this one map: it lists a kind only when the tally is positive, and sums the same map to decide which LAYERS to offer. Counting the AST layer alone therefore did not merely under-report — it drew those three classes on the canvas with no legend entry and no toggle, so a user could neither identify them nor turn them off.

Folder is counted only in the hierarchy wire mode, which is the only mode that emits Folder nodes.

node_count: int property

AST nodes in this view (real + External), drives the truncation banner.

for_slug(store: WikiStoreBase, slug: str, *, node_limit: int | None = None, hierarchy: bool = False, scope: CommitScope | None = None) -> KnowledgeGraphView classmethod

Load the full multiplex (ast + entity + memory layers) for slug.

Commit-scoped. scope defaults to store.live_scope(slug) — the project's own commit — so the viewer shows the code as it is now. Before this the AST layer was read unscoped, which is why the payload was the union of every generation ever indexed: deleted code rendered alongside live code with nothing distinguishing them, and the node count grew monotonically with re-indexes rather than with the repository.

Pass an explicit scope to read a specific generation; pass CommitScope.every() to restore the pre-scoping union (which the entity-anchor repair below relies on, and nothing else should).

AST connectivity: a cross-file IMPORTS/CALLS/EXTENDS edge carries the raw target_name; if that name resolves to a real in-repo node it is re-pointed there (genuinely connecting File-clusters through shared symbols), otherwise every reference to the same external name converges on ONE synthesized External view-node.

node_limit (when set and exceeded) degree-prunes the whole AST payload — real nodes AND the view-synthesized External nodes together, not the persisted layer alone. The AST nodes are pruned first (unchanged from before: highest-degree kept, ties on node_id); External nodes then get whatever budget is left, pruned by the same degree-then-node_id rule over their surviving edges. A node_limit that the AST layer alone does not exceed can still cap externals down, and an AST layer that exhausts the whole budget leaves externals at zero — both are the cap "governing the whole payload" rather than only the persisted one. Entity + memory layers are always fully included (they're small) and never counted against this cap. total_nodes/total_edges always reflect the full AST graph so the wire response can honestly say "showing N of M".

Cross-layer ANCHORS are reconciled to real node ids in O(nodes+edges): memory ANCHORS targets (EntityKey / entity:<id>) batch-resolve through the existing CodeStructureProvider + EntityAnchorResolver; an anchor that resolves to nothing is dropped (no dangling edges).

hierarchy (default off) synthesises a directory scaffold via FolderTree from the kept File nodes' paths — folder nodes + folder CONTAINS edges, and a single parentId per node — and stamps it onto the wire in to_wire. Off ⇒ the wire is byte-identical to the default mode (the SCG / Agentic Search reuse path is undisturbed).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/graph.py
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
@classmethod
def for_slug(
    cls,
    store: WikiStoreBase,
    slug: str,
    *,
    node_limit: int | None = None,
    hierarchy: bool = False,
    scope: CommitScope | None = None,
) -> KnowledgeGraphView:
    """Load the full multiplex (ast + entity + memory layers) for *slug*.

    **Commit-scoped.** *scope* defaults to ``store.live_scope(slug)`` — the
    project's own commit — so the viewer shows the code as it is now. Before
    this the AST layer was read unscoped, which is why the payload was the
    union of every generation ever indexed: deleted code rendered alongside
    live code with nothing distinguishing them, and the node count grew
    monotonically with re-indexes rather than with the repository.

    Pass an explicit scope to read a specific generation; pass
    ``CommitScope.every()`` to restore the pre-scoping union (which the
    entity-anchor repair below relies on, and nothing else should).

    AST connectivity: a cross-file IMPORTS/CALLS/EXTENDS edge carries the
    raw ``target_name``; if that name resolves to a real in-repo node it is
    re-pointed there (genuinely connecting File-clusters through shared
    symbols), otherwise every reference to the same external name converges
    on ONE synthesized ``External`` view-node.

    ``node_limit`` (when set and exceeded) degree-prunes the **whole AST
    payload** — real nodes AND the view-synthesized ``External`` nodes
    together, not the persisted layer alone. The AST nodes are pruned
    first (unchanged from before: highest-degree kept, ties on
    ``node_id``); External nodes then get whatever budget is left,
    pruned by the same degree-then-``node_id`` rule over their surviving
    edges. A ``node_limit`` that the AST layer alone does not exceed can
    still cap externals down, and an AST layer that exhausts the whole
    budget leaves externals at zero — both are the cap "governing the
    whole payload" rather than only the persisted one. Entity + memory
    layers are always fully included (they're small) and never counted
    against this cap. ``total_nodes``/``total_edges`` always reflect the
    full AST graph so the wire response can honestly say "showing N of M".

    Cross-layer ANCHORS are reconciled to real node ids in O(nodes+edges):
    memory ANCHORS targets (``EntityKey`` / ``entity:<id>``) batch-resolve
    through the existing ``CodeStructureProvider`` + ``EntityAnchorResolver``;
    an anchor that resolves to nothing is dropped (no dangling edges).

    ``hierarchy`` (default off) synthesises a directory scaffold via
    ``FolderTree`` from the kept File nodes' paths — folder nodes + folder
    ``CONTAINS`` edges, and a single ``parentId`` per node — and stamps it
    onto the wire in ``to_wire``. Off ⇒ the wire is byte-identical to the
    default mode (the SCG / Agentic Search reuse path is undisturbed).
    """
    from mewbo_graph.entities.anchor import EntityAnchorResolver  # noqa: PLC0415

    from .structure_provider import CodeStructureProvider  # noqa: PLC0415

    # ── AST layer ────────────────────────────────────────────────────
    scope = store.live_scope(slug) if scope is None else scope
    all_nodes = store.query_graph(slug, scope=scope)
    all_edges = list(store.list_edges(slug, scope=scope))
    total_nodes = len(all_nodes)

    # Resolve cross-file edge targets by name → real in-repo node. Build the
    # name index once (O(nodes)); externals converge on one synthesized node.
    by_name: dict[str, str] = {}
    for n in all_nodes:
        by_name.setdefault(n.name, n.node_id)
    resolved_edges, external_nodes = cls._resolve_ast_edges(
        slug, all_edges, by_name, {n.node_id for n in all_nodes}
    )
    # ``total_edges`` is the FULL emittable AST edge count — measured AFTER
    # resolution, because that pass drops orphan structural edges (a
    # ``target_name=None`` edge with a missing endpoint). Using the raw
    # ``len(all_edges)`` would overstate "M" by the orphans that never reach
    # the payload. (Resolution adds External nodes but never adds/drops
    # cross-file edges, so this count is cap-independent.)
    total_edges = len(resolved_edges)

    # Degree over the full resolved edge set — needed by the AST prune
    # below (when it fires) and by the external cap that follows it, so
    # it's computed once whenever a limit is in play at all. Left EMPTY
    # when nothing is capped: both readers are behind the same
    # ``node_limit is not None`` guard, and a Counter answers 0 for an
    # absent key, so an unlimited read can never see a wrong rank.
    degree: Counter[str] = Counter()
    if node_limit is not None:
        for e in resolved_edges:
            degree[e.source] += 1
            degree[e.target] += 1

    # Degree-prune the AST layer (externals follow their surviving edge,
    # then get their own cap below).
    if node_limit is None or total_nodes <= node_limit:
        nodes = list(all_nodes)
    else:
        nodes = sorted(
            all_nodes, key=lambda n: (-degree[n.node_id], n.node_id)
        )[:node_limit]

    kept_ast_ids = {n.node_id for n in nodes}
    ext_by_id = {n.node_id: n for n in external_nodes}
    edges = [
        e
        for e in resolved_edges
        if e.source in kept_ast_ids
        and (e.target in kept_ast_ids or e.target in ext_by_id)
    ]
    # Keep only externals still referenced by a surviving edge.
    live_ext_ids = {e.target for e in edges if e.target in ext_by_id}

    # ``node_limit`` caps the WHOLE payload, not just the persisted AST
    # layer — a view-synthesized External node still counts against it.
    # Externals get whatever budget the AST prune above left behind (zero
    # when that prune already spent the full cap), ranked by the SAME
    # degree-then-``node_id`` rule so the tie-break policy is one rule,
    # not two. Final tuple is re-sorted by ``node_id`` for stable output,
    # matching the AST layer's own ordering guarantee.
    if node_limit is None:
        live_externals = [ext_by_id[i] for i in live_ext_ids]
    else:
        budget = max(0, node_limit - len(nodes))
        live_externals = sorted(
            (ext_by_id[i] for i in live_ext_ids),
            key=lambda n: (-degree[n.node_id], n.node_id),
        )[:budget]
    kept_externals = tuple(sorted(live_externals, key=lambda n: n.node_id))
    kept_ext_ids = {n.node_id for n in kept_externals}
    if kept_ext_ids != live_ext_ids:
        # The external cap dropped some of the externals ``edges`` above
        # was built against — drop the now-dangling edges pointing at
        # them so no edge survives whose target was capped away.
        edges = [
            e
            for e in edges
            if e.target in kept_ast_ids or e.target in kept_ext_ids
        ]
    payload_ast_ids = kept_ast_ids | kept_ext_ids

    # ── Entity layer ─────────────────────────────────────────────────
    entity_nodes = store.query_entities(slug)
    entity_ids = {e.id for e in entity_nodes}
    entity_rels: list[EntityRelation] = []
    entity_cross: list[tuple[str, str]] = []
    stranded: list[EntityRelation] = []
    for rel in store.list_entity_edges(slug):
        if rel.target_id in entity_ids and rel.source_id in entity_ids:
            # entity ↔ entity → RELATES (verb in label)
            entity_rels.append(rel)
        elif rel.source_id in entity_ids and rel.target_id in payload_ast_ids:
            # entity → AST node → cross-layer ANCHORS
            entity_cross.append((rel.source_id, rel.target_id))
        elif rel.source_id in entity_ids:
            # Anchored at an AST node that is not in this payload. Under a
            # commit scope that is usually not a dangling anchor but a
            # SUPERSEDED one — see the repair below.
            stranded.append(rel)
        # else: dangling (target absent from this payload) → dropped
    entity_cross.extend(
        cls._reanchor_entity_edges(store, slug, stranded, nodes, payload_ast_ids)
    )

    # ── Memory layer ─────────────────────────────────────────────────
    memory_nodes = store.query_memory(slug)
    memory_ids = {n.node_id for n in memory_nodes}
    mem_rels: list[MemoryEdge] = []
    mem_anchor_edges: list[MemoryEdge] = []
    for me in store.list_memory_edges(slug, include_invalidated=False):
        if me.type == "RELATES" and me.source in memory_ids:
            mem_rels.append(me)
        elif me.type == "ANCHORS" and me.source in memory_ids:
            mem_anchor_edges.append(me)

    # Batch-resolve memory ANCHORS targets to real node ids (one pass each).
    code_keys = [e.target for e in mem_anchor_edges if not e.target.startswith("entity:")]
    ent_keys = [e.target for e in mem_anchor_edges if e.target.startswith("entity:")]
    code_map = CodeStructureProvider(store).resolve_many(slug, code_keys)
    ent_map = EntityAnchorResolver(store).resolve_many(slug, ent_keys)
    memory_cross: list[tuple[str, str]] = []
    for me in mem_anchor_edges:
        if me.target.startswith("entity:"):
            ent = ent_map.get(me.target)
            tid = ent.id if ent is not None and ent.id in entity_ids else None
        else:
            node = code_map.get(me.target)
            tid = node.node_id if node is not None and node.node_id in payload_ast_ids else None
        if tid is not None:
            memory_cross.append((me.source, tid))

    # ── Directory scaffold (hierarchy wire mode only) ────────────────
    # Built over the KEPT AST layer so folder nodes only scaffold files
    # actually in the payload; the symbol→container parentId reads off the
    # same filtered CONTAINS edges the wire emits.
    folder_tree = (
        FolderTree.build(slug, nodes, edges, external_nodes=kept_externals)
        if hierarchy
        else None
    )

    return cls(
        slug=slug,
        nodes=tuple(nodes),
        external_nodes=kept_externals,
        edges=tuple(edges),
        entity_nodes=tuple(entity_nodes),
        entity_edges=tuple(entity_rels),
        memory_nodes=tuple(memory_nodes),
        memory_edges=tuple(mem_rels),
        cross_edges=tuple(entity_cross + memory_cross),
        total_nodes=total_nodes,
        total_edges=total_edges,
        total_externals=len(live_ext_ids),
        folder_tree=folder_tree,
    )

to_wire() -> dict[str, Any]

Wire shape: {slug, nodes, edges, stats} over all three layers.

Each node/edge is Cytoscape-ready ({data: {...}}) and carries a layer tag so the FE can style/filter per layer. The FE hands the arrays straight to cy.add(elements) with no intermediate transform.

When the view was built with hierarchy=True the directory scaffold is merged in: synthesised Folder nodes + their CONTAINS edges are appended, and every emitted node gains a parentId / folderPath (folderCount lands in stats). With hierarchy off the folder_tree is None and the payload is byte-identical to the default mode.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/graph.py
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
def to_wire(self) -> dict[str, Any]:
    """Wire shape: ``{slug, nodes, edges, stats}`` over all three layers.

    Each node/edge is Cytoscape-ready (``{data: {...}}``) and carries a
    ``layer`` tag so the FE can style/filter per layer. The FE hands the
    arrays straight to ``cy.add(elements)`` with no intermediate transform.

    When the view was built with ``hierarchy=True`` the directory scaffold
    is merged in: synthesised ``Folder`` nodes + their ``CONTAINS`` edges
    are appended, and every emitted node gains a ``parentId`` /
    ``folderPath`` (``folderCount`` lands in ``stats``). With hierarchy off
    the ``folder_tree`` is ``None`` and the payload is byte-identical to the
    default mode.
    """
    tree = self.folder_tree
    nodes = (
        [self._node_to_wire(n, "ast", tree) for n in self.nodes]
        + [self._node_to_wire(n, "ast", tree) for n in self.external_nodes]
        + [self._entity_node_to_wire(e, tree) for e in self.entity_nodes]
        + [self._memory_node_to_wire(m, tree) for m in self.memory_nodes]
    )
    edges = (
        [self._edge_to_wire(e, "ast") for e in self.edges]
        + [self._entity_edge_to_wire(e) for e in self.entity_edges]
        + [self._memory_edge_to_wire(e) for e in self.memory_edges]
        + [self._cross_edge_to_wire(s, t) for s, t in self.cross_edges]
    )
    if tree is not None:
        # Folder nodes carry their own parentId/folderPath off the tree.
        nodes += [self._node_to_wire(f, "ast", tree) for f in tree.folder_nodes]
        edges += [self._edge_to_wire(e, "ast") for e in tree.folder_edges]
    stats: dict[str, Any] = {
        # AST-only counters — the flat shape wire consumers read.
        "nodeCount": self.node_count,
        "edgeCount": self.edge_count,
        "kinds": self.kinds,
        # FULL multiplex counts ("M" in the FE "showing N of M" banner).
        # Only the AST layer is ever capped, so entity + memory + the
        # view-only External nodes contribute their in-view counts; the
        # pre-cap AST total is ``self.total_nodes``. Uncapped ⇒ M == N.
        "totalNodes": (
            self.total_nodes
            + self.total_externals
            + len(self.entity_nodes)
            + len(self.memory_nodes)
        ),
        "totalEdges": (
            self.total_edges
            + len(self.entity_edges)
            + len(self.memory_edges)
            + len(self.cross_edges)
        ),
        # ``truncated`` answers "did the cap drop anything the payload
        # would otherwise carry", and the cap governs BOTH populations —
        # so both are compared against their own pre-cap total. The two
        # counts stay separate rather than summed: summing lets a surplus
        # on one side mask a shortfall on the other, which is how a
        # payload that had lost a quarter of its nodes still reported
        # itself complete. Entity + memory layers are never capped, and
        # an edge dropped by orphan hygiene is not truncation.
        "truncated": (
            len(self.nodes) < self.total_nodes
            or len(self.external_nodes) < self.total_externals
        ),
        "perLayer": {
            "ast": len(self.nodes) + len(self.external_nodes),
            "entity": len(self.entity_nodes),
            "memory": len(self.memory_nodes),
        },
    }
    # ``folderCount`` only in hierarchy mode — the default wire stays
    # byte-identical (the SCG / Agentic Search reuse path is undisturbed).
    if tree is not None:
        stats["folderCount"] = len(tree.folder_nodes)
    return {
        "slug": self.slug,
        "nodes": nodes,
        "edges": edges,
        "stats": stats,
    }

LanguageSpec dataclass

One tree-sitter-backed language the code-graph extractor supports.

query_file defaults to <name>.scm — set it only when a language's query file diverges from its language name, as tsx does: it is a SEPARATE grammar from typescript (the two disagree about whether <T> opens a type assertion or a JSX element) but the node types the graph captures are identical, so one query file serves both.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/graph.py
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
@dataclass(frozen=True)
class LanguageSpec:
    """One tree-sitter-backed language the code-graph extractor supports.

    ``query_file`` defaults to ``<name>.scm`` — set it only when a language's
    query file diverges from its language name, as ``tsx`` does: it is a
    SEPARATE grammar from ``typescript`` (the two disagree about whether
    ``<T>`` opens a type assertion or a JSX element) but the node types the
    graph captures are identical, so one query file serves both.
    """

    name: str
    extensions: tuple[str, ...]
    query_file: str | None = None

    @property
    def query_filename(self) -> str:
        """Resolved query file name — ``query_file`` override, else ``<name>.scm``."""
        return self.query_file or f"{self.name}.scm"

query_filename: str property

Resolved query file name — query_file override, else <name>.scm.

mewbo_graph.wiki.structure_provider

StructureProvider — the code↔multiplex-key bridge.

entity_key (path/to/file.py#Qualified.Name, no byte offsets) is the shared identity that joins the memory and docs layers to the code graph. This module owns the only derivation of an entity_key from a GraphNode and the resolution back to a live node.

StructureProvider is a Protocol so the structural layer is pluggable (corpus-agnostic seam — code today; PDF sections / DB schemas later). v1 ships exactly one implementation, CodeStructureProvider, which composes over the wiki store. Keep it stateless: a refresh mutates the graph, so a cached map would go stale.

CodeStructureProvider

StructureProvider over the tree-sitter code graph (v1).

The two directions read different commit scopes, on purpose. resolve/resolve_many answer "where does this key live now", so they read the live generation: an entity_key carries no byte offset, so it re-resolves cleanly onto whatever generation is current, and returning a superseded node would hand the caller a node id no live payload contains. entity_key_of answers the opposite question about a node id recorded in the PAST — a QA provenance ref from an earlier session — so it must read the union or it would fail to label exactly the historical refs it exists to label.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/structure_provider.py
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
class CodeStructureProvider:
    """``StructureProvider`` over the tree-sitter code graph (v1).

    **The two directions read different commit scopes, on purpose.**
    ``resolve``/``resolve_many`` answer "where does this key live *now*", so
    they read the live generation: an ``entity_key`` carries no byte offset, so
    it re-resolves cleanly onto whatever generation is current, and returning a
    superseded node would hand the caller a node id no live payload contains.
    ``entity_key_of`` answers the opposite question about a node id recorded in
    the PAST — a QA provenance ref from an earlier session — so it must read
    the union or it would fail to label exactly the historical refs it exists
    to label.
    """

    def __init__(self, store: WikiStoreBase, *, scope: CommitScope | None = None) -> None:
        """Compose over a wiki store; *scope* defaults to the slug's live commit."""
        self._store = store
        self._scope = scope

    def _live(self, slug: str) -> CommitScope:
        """The scope forward resolution reads (injected one wins)."""
        return self._scope if self._scope is not None else self._store.live_scope(slug)

    def resolve(self, slug: str, entity_key: EntityKey) -> GraphNode | None:
        """Return the node addressed by *entity_key*, or None if absent."""
        for node in self._store.query_graph(slug, scope=self._live(slug)):
            if entity_key_for_node(node) == entity_key:
                return node
        return None

    def resolve_many(
        self, slug: str, entity_keys: list[EntityKey]
    ) -> dict[EntityKey, GraphNode]:
        """Resolve a batch in one graph pass; misses are omitted."""
        wanted = set(entity_keys)
        out: dict[EntityKey, GraphNode] = {}
        if not wanted:
            return out
        for node in self._store.query_graph(slug, scope=self._live(slug)):
            key = entity_key_for_node(node)
            if key in wanted and key not in out:
                out[key] = node
                if len(out) == len(wanted):
                    break
        return out

    def entity_key_of(self, slug: str, node_id: str) -> EntityKey | None:
        """Return the ``entity_key`` for a code ``node_id``, or None.

        Reads EVERY generation deliberately — see the class docstring. The
        callers hand it node ids captured during earlier sessions, and a node
        id embeds the symbol's byte offset, so any edit since then re-keyed it
        out of the live generation. Scoping this would turn a resolvable
        historical citation into an ``unknown(...)`` label.
        """
        for node in self._store.query_graph(slug, scope=CommitScope.every()):
            if node.node_id == node_id:
                return entity_key_for_node(node)
        return None

__init__(store: WikiStoreBase, *, scope: CommitScope | None = None) -> None

Compose over a wiki store; scope defaults to the slug's live commit.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/structure_provider.py
72
73
74
75
def __init__(self, store: WikiStoreBase, *, scope: CommitScope | None = None) -> None:
    """Compose over a wiki store; *scope* defaults to the slug's live commit."""
    self._store = store
    self._scope = scope

entity_key_of(slug: str, node_id: str) -> EntityKey | None

Return the entity_key for a code node_id, or None.

Reads EVERY generation deliberately — see the class docstring. The callers hand it node ids captured during earlier sessions, and a node id embeds the symbol's byte offset, so any edit since then re-keyed it out of the live generation. Scoping this would turn a resolvable historical citation into an unknown(...) label.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/structure_provider.py
104
105
106
107
108
109
110
111
112
113
114
115
116
def entity_key_of(self, slug: str, node_id: str) -> EntityKey | None:
    """Return the ``entity_key`` for a code ``node_id``, or None.

    Reads EVERY generation deliberately — see the class docstring. The
    callers hand it node ids captured during earlier sessions, and a node
    id embeds the symbol's byte offset, so any edit since then re-keyed it
    out of the live generation. Scoping this would turn a resolvable
    historical citation into an ``unknown(...)`` label.
    """
    for node in self._store.query_graph(slug, scope=CommitScope.every()):
        if node.node_id == node_id:
            return entity_key_for_node(node)
    return None

resolve(slug: str, entity_key: EntityKey) -> GraphNode | None

Return the node addressed by entity_key, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/structure_provider.py
81
82
83
84
85
86
def resolve(self, slug: str, entity_key: EntityKey) -> GraphNode | None:
    """Return the node addressed by *entity_key*, or None if absent."""
    for node in self._store.query_graph(slug, scope=self._live(slug)):
        if entity_key_for_node(node) == entity_key:
            return node
    return None

resolve_many(slug: str, entity_keys: list[EntityKey]) -> dict[EntityKey, GraphNode]

Resolve a batch in one graph pass; misses are omitted.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/structure_provider.py
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
def resolve_many(
    self, slug: str, entity_keys: list[EntityKey]
) -> dict[EntityKey, GraphNode]:
    """Resolve a batch in one graph pass; misses are omitted."""
    wanted = set(entity_keys)
    out: dict[EntityKey, GraphNode] = {}
    if not wanted:
        return out
    for node in self._store.query_graph(slug, scope=self._live(slug)):
        key = entity_key_for_node(node)
        if key in wanted and key not in out:
            out[key] = node
            if len(out) == len(wanted):
                break
    return out

StructureProvider

Bases: Protocol

Resolves between entity_key and the underlying structural unit.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/structure_provider.py
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
@runtime_checkable
class StructureProvider(Protocol):
    """Resolves between ``entity_key`` and the underlying structural unit."""

    def resolve(self, slug: str, entity_key: EntityKey) -> GraphNode | None:
        """Return the node addressed by *entity_key*, or None if absent."""
        ...

    def resolve_many(
        self, slug: str, entity_keys: list[EntityKey]
    ) -> dict[EntityKey, GraphNode]:
        """Resolve a batch in one pass; misses are omitted from the result."""
        ...

    def entity_key_of(self, slug: str, node_id: str) -> EntityKey | None:
        """Return the ``entity_key`` for a code ``node_id``, or None."""
        ...

entity_key_of(slug: str, node_id: str) -> EntityKey | None

Return the entity_key for a code node_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/structure_provider.py
53
54
55
def entity_key_of(self, slug: str, node_id: str) -> EntityKey | None:
    """Return the ``entity_key`` for a code ``node_id``, or None."""
    ...

resolve(slug: str, entity_key: EntityKey) -> GraphNode | None

Return the node addressed by entity_key, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/structure_provider.py
43
44
45
def resolve(self, slug: str, entity_key: EntityKey) -> GraphNode | None:
    """Return the node addressed by *entity_key*, or None if absent."""
    ...

resolve_many(slug: str, entity_keys: list[EntityKey]) -> dict[EntityKey, GraphNode]

Resolve a batch in one pass; misses are omitted from the result.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/structure_provider.py
47
48
49
50
51
def resolve_many(
    self, slug: str, entity_keys: list[EntityKey]
) -> dict[EntityKey, GraphNode]:
    """Resolve a batch in one pass; misses are omitted from the result."""
    ...

entity_key_for_node(node: GraphNode) -> EntityKey

Derive the multiplex entity_key for a code node.

File nodes key on their bare path; every other symbol keys on file#name. name is the tree-sitter symbol name — class→method qualification lands with the graph's scoping work, so two same-named methods in one file currently collapse to one key (an accepted v1 over-approximation: a false anchor is wasted work, never data loss).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/structure_provider.py
25
26
27
28
29
30
31
32
33
34
35
36
def entity_key_for_node(node: GraphNode) -> EntityKey:
    """Derive the multiplex ``entity_key`` for a code node.

    File nodes key on their bare path; every other symbol keys on
    ``file#name``. ``name`` is the tree-sitter symbol name — class→method
    qualification lands with the graph's scoping work, so two same-named
    methods in one file currently collapse to one key (an accepted v1
    over-approximation: a false anchor is wasted work, never data loss).
    """
    if node.type == "File":
        return node.file
    return f"{node.file}#{node.name}"

mewbo_graph.wiki.embedder

Embedder — thin wrapper around litellm.embedding.

LiteLLM is the project's canonical LLM client (chat completions already route through it), so embeddings ride the same proxy plumbing for free. This module's job is to:

  1. Read the configured embedding model from wiki.embedding.model.
  2. Normalise it with the proxy prefix (openai/<model>) so LiteLLM sends the request to our OpenAI-compatible LiteLLM proxy instead of trying to dispatch directly to a provider SDK.
  3. Wrap litellm.embedding so its return value materialises into our typed Embedding records (with slug + node_id + dim).
  4. Provide cosine and search static helpers.

Embedding is I/O-bound, not CPU-bound: a large indexing pass spends its time waiting on the network, one blocking call after another, while the process itself sits near idle. _EmbeddingPacer is what turns that into a bounded, rate-limit-aware pool of concurrent requests instead of a serial loop — see its docstring for the pacing/backoff contract.

KISS: no third-party LangChain abstraction layer, no batching wrappers — LiteLLM already handles batching and provider routing; this module adds only the concurrency/pacing layer LiteLLM does not provide.

Embedder

Thin facade: litellm.embedding + typed Embedding records.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
class Embedder:
    """Thin facade: ``litellm.embedding`` + typed ``Embedding`` records."""

    # Project convention: chat models go through the LiteLLM proxy as
    # ``openai/<model>`` so LiteLLM uses its OpenAI-compatible client
    # against ``llm.api_base`` instead of routing to a provider SDK.
    # Same rule applies to embedding model names.
    _PROXY_PREFIX = "openai/"

    def __init__(
        self,
        *,
        model: str | None = None,
        batch_size: int | None = None,
        concurrency: int | None = None,
        requests_per_minute: int | None = None,
        tokens_per_minute: int | None = None,
        max_retries: int | None = None,
        sleeper: Callable[[float], None] | None = None,
    ) -> None:
        """Construct the Embedder from config + kwargs.

        ``sleeper`` is the retry wait, injected as a collaborator so a caller
        observing the backoff observes only THIS object's waiting. Patching
        ``time.sleep`` on the stdlib module instead reaches every module and
        every thread in the process: the observation picks up unrelated
        waiting, and neutering it turns any concurrent poll loop into a spin.
        """
        cfg = get_config()
        raw_model = model or get_config_value(
            "wiki", "embedding", "model", default="openai/text-embedding-3-small"
        )
        self.model = self._normalise_model(raw_model)
        self.batch_size = batch_size or int(
            get_config_value("wiki", "embedding", "batch_size", default=64)
        )
        self._concurrency = concurrency or int(
            get_config_value("wiki", "embedding", "concurrency", default=4)
        )
        self._requests_per_minute = requests_per_minute or get_config_value(
            "wiki", "embedding", "requests_per_minute", default=None
        )
        self._tokens_per_minute = tokens_per_minute or get_config_value(
            "wiki", "embedding", "tokens_per_minute", default=None
        )
        self._max_retries = (
            max_retries
            if max_retries is not None
            else int(get_config_value("wiki", "embedding", "max_retries", default=5))
        )
        self._sleep = sleeper or time.sleep
        self._api_base = cfg.llm.api_base or None
        self._api_key = cfg.llm.api_key or "missing"
        self._pacer = _EmbeddingPacer(
            concurrency=self._concurrency,
            requests_per_minute=self._requests_per_minute,
            tokens_per_minute=self._tokens_per_minute,
        )

    @staticmethod
    def enabled() -> bool:
        """True when node embedding is switched on (``wiki.embedding.enabled``).

        The switch lives with the class that owns embedding rather than beside
        one of its callers: every path that embeds — the full index and the
        scoped refresh — has to read the same knob, and an operator who turns
        embedding off has no way to tell which caller re-implemented the read.
        Defaults to on, so an absent key never silently disables retrieval.
        """
        return bool(get_config_value("wiki", "embedding", "enabled", default=True))

    @classmethod
    def _normalise_model(cls, model: str) -> str:
        """Ensure the model name carries a provider prefix LiteLLM understands.

        Bare names like ``gemini-embedding-001`` route directly to a
        provider SDK and bypass our proxy. Prepending ``openai/`` forces
        the OpenAI-compatible path against ``api_base``.
        """
        return model if "/" in model else f"{cls._PROXY_PREFIX}{model}"

    def embed_nodes(
        self,
        items: list[tuple[str, str]],
        *,
        slug: str = "",
    ) -> list[Embedding]:
        """Embed ``(node_id, text)`` pairs and return ``Embedding`` records."""
        if not items:
            return []
        texts = [text for _, text in items]
        vectors = self._embed(texts)
        return [
            Embedding(
                slug=slug,
                node_id=nid,
                vector=list(vec),
                model=self.model,
                dim=len(vec),
            )
            for (nid, _), vec in zip(items, vectors)
        ]

    def embed_query(self, text: str) -> list[float]:
        """Embed a single query string. Stays cheap: never spins up a pool."""
        vectors = self._embed([text])
        return vectors[0] if vectors else []

    def _embed(self, texts: list[str]) -> list[list[float]]:
        """Issue one or more embedding calls, batching to ``batch_size``.

        Cost class: ``O(len(texts) / batch_size)`` requests, issued through a
        bounded thread pool — never more than ``concurrency`` in flight, and
        paced against ``requests_per_minute`` / ``tokens_per_minute`` when
        configured. Order is a hard contract: every caller depends on
        ``_embed`` returning one vector per input text, in input order, so
        each batch is submitted with its position and results are reassembled
        by index rather than by completion order.

        A single batch (the common case for ``embed_query`` and any job whose
        node count fits under ``batch_size``) is issued directly, with no
        thread pool spun up.
        """
        if not texts:
            return []
        batches = [
            texts[start : start + self.batch_size]
            for start in range(0, len(texts), self.batch_size)
        ]
        if len(batches) == 1:
            return self._call_batch(batches[0])

        results: list[list[list[float]] | None] = [None] * len(batches)
        workers = min(self._pacer_concurrency_hint(), len(batches))
        with ThreadPoolExecutor(max_workers=workers) as pool:
            futures = {
                pool.submit(self._call_batch, batch): idx
                for idx, batch in enumerate(batches)
            }
            for future in futures:
                idx = futures[future]
                results[idx] = future.result()

        out: list[list[float]] = []
        for vectors in results:
            assert vectors is not None  # every future resolved or raised above
            out.extend(vectors)
        return out

    def _pacer_concurrency_hint(self) -> int:
        """Thread-pool sizing hint — the pacer enforces the real bound.

        Sized off the CONFIGURED ceiling, not the pacer's live (possibly
        throttled-down) one, so a run that gets throttled keeps its existing
        worker threads — each blocks in ``_EmbeddingPacer.acquire`` until the
        lower ceiling admits it — rather than needing new threads spun up.
        """
        return max(1, self._concurrency)

    def _call_batch(self, batch: list[str]) -> list[list[float]]:
        """Issue one batched embedding call, retrying through 429s.

        A 429 is retried up to ``max_retries`` times: the provider's
        ``Retry-After`` header wins when present, otherwise exponential
        backoff with jitter. Sustained rate-limiting also shrinks the
        pacer's concurrency ceiling, so a rate-limited run degrades to
        slower rather than to failed. Any other exception propagates
        immediately — a different credential or a slower pace cannot fix a
        transport failure or a bad request.
        """
        estimated_tokens = self._estimate_tokens(batch)
        attempt = 0
        while True:
            self._pacer.acquire(estimated_tokens)
            try:
                resp = litellm.embedding(
                    model=self.model,
                    input=batch,
                    api_base=self._api_base,
                    api_key=self._api_key,
                )
            except litellm.RateLimitError as exc:
                self._pacer.release()
                attempt += 1
                if attempt > self._max_retries:
                    raise
                self._pacer.throttle()
                delay = self._retry_delay(exc, attempt)
                logging.warning(
                    "Embedding request rate-limited, retrying "
                    f"(attempt {attempt}/{self._max_retries}, waiting {delay:.1f}s)"
                )
                self._sleep(delay)
                continue
            except Exception:
                self._pacer.release()
                raise
            else:
                self._pacer.release()
                return [self._vector_of(row) for row in resp.data]

    @staticmethod
    def _vector_of(row: Any) -> list[float]:
        # litellm returns either a dict ({'embedding': [...], 'index': N})
        # or an EmbeddingResponse pydantic object — handle both.
        vec = row["embedding"] if isinstance(row, dict) else row.embedding
        return list(vec)

    @staticmethod
    def _estimate_tokens(batch: list[str]) -> int:
        """Rough token estimate for TPM pacing — not exact accounting."""
        return sum(max(1, len(text) // _CHARS_PER_TOKEN_ESTIMATE) for text in batch)

    @staticmethod
    def _retry_delay(exc: litellm.RateLimitError, attempt: int) -> float:
        retry_after = Embedder._parse_retry_after(exc)
        if retry_after is not None:
            return retry_after
        backoff = min(2 ** (attempt - 1), _MAX_BACKOFF_SECONDS)
        return backoff + random.uniform(0, backoff * 0.5)

    @staticmethod
    def _parse_retry_after(exc: litellm.RateLimitError) -> float | None:
        response = getattr(exc, "response", None)
        headers = getattr(response, "headers", None)
        if not headers:
            return None
        raw = headers.get("retry-after")
        if raw is None:
            return None
        try:
            return max(0.0, float(raw))
        except (TypeError, ValueError):
            return None

    # ── Vector math (provider-agnostic, no embedder state needed) ──────

    @staticmethod
    def cosine(a: list[float], b: list[float]) -> float:
        """Cosine similarity. Returns 0.0 if either vector is zero-length.

        Delegates to the shared, dependency-free ``mewbo_graph._util.cosine`` so
        the wiki vector math and the entity resolution ladder can never desync.
        """
        return _cosine(a, b)

    @staticmethod
    def search(
        qvec: list[float],
        vectors: list[list[float]],
        k: int = 10,
    ) -> list[tuple[int, float]]:
        """Return ``(index, cosine_score)`` for the top-k matches, sorted desc."""
        if not vectors:
            return []
        scored = [(i, Embedder.cosine(qvec, v)) for i, v in enumerate(vectors)]
        scored.sort(key=lambda t: t[1], reverse=True)
        return scored[:k]

__init__(*, model: str | None = None, batch_size: int | None = None, concurrency: int | None = None, requests_per_minute: int | None = None, tokens_per_minute: int | None = None, max_retries: int | None = None, sleeper: Callable[[float], None] | None = None) -> None

Construct the Embedder from config + kwargs.

sleeper is the retry wait, injected as a collaborator so a caller observing the backoff observes only THIS object's waiting. Patching time.sleep on the stdlib module instead reaches every module and every thread in the process: the observation picks up unrelated waiting, and neutering it turns any concurrent poll loop into a spin.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
def __init__(
    self,
    *,
    model: str | None = None,
    batch_size: int | None = None,
    concurrency: int | None = None,
    requests_per_minute: int | None = None,
    tokens_per_minute: int | None = None,
    max_retries: int | None = None,
    sleeper: Callable[[float], None] | None = None,
) -> None:
    """Construct the Embedder from config + kwargs.

    ``sleeper`` is the retry wait, injected as a collaborator so a caller
    observing the backoff observes only THIS object's waiting. Patching
    ``time.sleep`` on the stdlib module instead reaches every module and
    every thread in the process: the observation picks up unrelated
    waiting, and neutering it turns any concurrent poll loop into a spin.
    """
    cfg = get_config()
    raw_model = model or get_config_value(
        "wiki", "embedding", "model", default="openai/text-embedding-3-small"
    )
    self.model = self._normalise_model(raw_model)
    self.batch_size = batch_size or int(
        get_config_value("wiki", "embedding", "batch_size", default=64)
    )
    self._concurrency = concurrency or int(
        get_config_value("wiki", "embedding", "concurrency", default=4)
    )
    self._requests_per_minute = requests_per_minute or get_config_value(
        "wiki", "embedding", "requests_per_minute", default=None
    )
    self._tokens_per_minute = tokens_per_minute or get_config_value(
        "wiki", "embedding", "tokens_per_minute", default=None
    )
    self._max_retries = (
        max_retries
        if max_retries is not None
        else int(get_config_value("wiki", "embedding", "max_retries", default=5))
    )
    self._sleep = sleeper or time.sleep
    self._api_base = cfg.llm.api_base or None
    self._api_key = cfg.llm.api_key or "missing"
    self._pacer = _EmbeddingPacer(
        concurrency=self._concurrency,
        requests_per_minute=self._requests_per_minute,
        tokens_per_minute=self._tokens_per_minute,
    )

cosine(a: list[float], b: list[float]) -> float staticmethod

Cosine similarity. Returns 0.0 if either vector is zero-length.

Delegates to the shared, dependency-free mewbo_graph._util.cosine so the wiki vector math and the entity resolution ladder can never desync.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
466
467
468
469
470
471
472
473
@staticmethod
def cosine(a: list[float], b: list[float]) -> float:
    """Cosine similarity. Returns 0.0 if either vector is zero-length.

    Delegates to the shared, dependency-free ``mewbo_graph._util.cosine`` so
    the wiki vector math and the entity resolution ladder can never desync.
    """
    return _cosine(a, b)

embed_nodes(items: list[tuple[str, str]], *, slug: str = '') -> list[Embedding]

Embed (node_id, text) pairs and return Embedding records.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
def embed_nodes(
    self,
    items: list[tuple[str, str]],
    *,
    slug: str = "",
) -> list[Embedding]:
    """Embed ``(node_id, text)`` pairs and return ``Embedding`` records."""
    if not items:
        return []
    texts = [text for _, text in items]
    vectors = self._embed(texts)
    return [
        Embedding(
            slug=slug,
            node_id=nid,
            vector=list(vec),
            model=self.model,
            dim=len(vec),
        )
        for (nid, _), vec in zip(items, vectors)
    ]

embed_query(text: str) -> list[float]

Embed a single query string. Stays cheap: never spins up a pool.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
332
333
334
335
def embed_query(self, text: str) -> list[float]:
    """Embed a single query string. Stays cheap: never spins up a pool."""
    vectors = self._embed([text])
    return vectors[0] if vectors else []

enabled() -> bool staticmethod

True when node embedding is switched on (wiki.embedding.enabled).

The switch lives with the class that owns embedding rather than beside one of its callers: every path that embeds — the full index and the scoped refresh — has to read the same knob, and an operator who turns embedding off has no way to tell which caller re-implemented the read. Defaults to on, so an absent key never silently disables retrieval.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
288
289
290
291
292
293
294
295
296
297
298
@staticmethod
def enabled() -> bool:
    """True when node embedding is switched on (``wiki.embedding.enabled``).

    The switch lives with the class that owns embedding rather than beside
    one of its callers: every path that embeds — the full index and the
    scoped refresh — has to read the same knob, and an operator who turns
    embedding off has no way to tell which caller re-implemented the read.
    Defaults to on, so an absent key never silently disables retrieval.
    """
    return bool(get_config_value("wiki", "embedding", "enabled", default=True))

search(qvec: list[float], vectors: list[list[float]], k: int = 10) -> list[tuple[int, float]] staticmethod

Return (index, cosine_score) for the top-k matches, sorted desc.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
475
476
477
478
479
480
481
482
483
484
485
486
@staticmethod
def search(
    qvec: list[float],
    vectors: list[list[float]],
    k: int = 10,
) -> list[tuple[int, float]]:
    """Return ``(index, cosine_score)`` for the top-k matches, sorted desc."""
    if not vectors:
        return []
    scored = [(i, Embedder.cosine(qvec, v)) for i, v in enumerate(vectors)]
    scored.sort(key=lambda t: t[1], reverse=True)
    return scored[:k]

EmbedderProtocol

Bases: Protocol

The duck-typed embedder surface retriever/ingestor depend on.

Embedder (litellm-backed) and _NullEmbedder (BM25-fallback null object) both satisfy this; typing against it instead of Any catches wiring errors at definition.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
@runtime_checkable
class EmbedderProtocol(Protocol):
    """The duck-typed embedder surface retriever/ingestor depend on.

    ``Embedder`` (litellm-backed) and ``_NullEmbedder`` (BM25-fallback null
    object) both satisfy this; typing against it instead of ``Any`` catches
    wiring errors at definition.
    """

    def embed_nodes(
        self, items: list[tuple[str, str]], *, slug: str = ""
    ) -> list[Embedding]:
        """Embed ``(node_id, text)`` pairs into ``Embedding`` records."""
        ...

    def embed_query(self, text: str) -> list[float]:
        """Embed a single query string into a vector."""
        ...

embed_nodes(items: list[tuple[str, str]], *, slug: str = '') -> list[Embedding]

Embed (node_id, text) pairs into Embedding records.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
65
66
67
68
69
def embed_nodes(
    self, items: list[tuple[str, str]], *, slug: str = ""
) -> list[Embedding]:
    """Embed ``(node_id, text)`` pairs into ``Embedding`` records."""
    ...

embed_query(text: str) -> list[float]

Embed a single query string into a vector.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
71
72
73
def embed_query(self, text: str) -> list[float]:
    """Embed a single query string into a vector."""
    ...

make_embedder() -> Embedder

Construct the wiki Embedder using the configured proxy + model.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
76
77
78
def make_embedder() -> Embedder:
    """Construct the wiki Embedder using the configured proxy + model."""
    return Embedder()

make_embedder_for(store: Any, slug: str) -> Embedder

Build the Embedder bound to slug's own embedding model.

THE reason this exists rather than each caller reading config: the write side and the read side have to agree. Two embedding models rarely share a vector width, and vector_search scores cosine over whatever is stored — so a query embedded with a different model than the vectors it is scored against returns wrong neighbours rather than an error. Resolving both sides through one function is what makes that agreement structural instead of a convention every new retrieval site has to remember.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
104
105
106
107
108
109
110
111
112
113
114
115
def make_embedder_for(store: Any, slug: str) -> Embedder:
    """Build the Embedder bound to *slug*'s own embedding model.

    THE reason this exists rather than each caller reading config: the write
    side and the read side have to agree. Two embedding models rarely share a
    vector width, and ``vector_search`` scores cosine over whatever is stored —
    so a query embedded with a different model than the vectors it is scored
    against returns wrong neighbours rather than an error. Resolving both sides
    through one function is what makes that agreement structural instead of a
    convention every new retrieval site has to remember.
    """
    return Embedder(model=project_embedding_model(store, slug))

make_embedder_for_or_none(store: Any, slug: str) -> Embedder | None

Build slug's Embedder, or None when the caller may fall back to BM25.

The graceful twin of :func:make_embedder_for, for the write paths where a missing embedding backend must degrade retrieval rather than fail an index. It resolves the project's model and then goes through :func:make_embedder_or_none rather than constructing directly, so there stays exactly ONE graceful construction path however the model was chosen.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
118
119
120
121
122
123
124
125
126
127
def make_embedder_for_or_none(store: Any, slug: str) -> Embedder | None:
    """Build *slug*'s Embedder, or ``None`` when the caller may fall back to BM25.

    The graceful twin of :func:`make_embedder_for`, for the write paths where a
    missing embedding backend must degrade retrieval rather than fail an index.
    It resolves the project's model and then goes through
    :func:`make_embedder_or_none` rather than constructing directly, so there
    stays exactly ONE graceful construction path however the model was chosen.
    """
    return make_embedder_or_none(project_embedding_model(store, slug))

make_embedder_or_none(model: str | None = None) -> Embedder | None

Build an Embedder, or None if it can't be constructed (BM25-only).

The single construction path for callers that must degrade gracefully when no embedding backend is configured — used by insight ingestion so a missing proxy never fails a write.

model is the project's own embedding model, or None to inherit the deployment default. It is a defaulted parameter rather than a second function because every caller degrades identically; splitting them would give the graceful path two implementations, and a test patching one of them would let the other build a live Embedder and reach the network.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
def make_embedder_or_none(model: str | None = None) -> Embedder | None:
    """Build an Embedder, or None if it can't be constructed (BM25-only).

    The single construction path for callers that must degrade gracefully
    when no embedding backend is configured — used by insight ingestion so a
    missing proxy never fails a write.

    *model* is the project's own embedding model, or ``None`` to inherit the
    deployment default. It is a defaulted parameter rather than a second
    function because every caller degrades identically; splitting them would
    give the graceful path two implementations, and a test patching one of them
    would let the other build a live Embedder and reach the network.
    """
    try:
        return Embedder(model=model)
    except Exception:
        return None

project_embedding_model(store: Any, slug: str) -> str | None

The embedding model slug is indexed and searched with, or None.

None means "inherit wiki.embedding.model" — the answer for every project indexed before the override existed, and for every project whose operator never set one.

Best-effort by construction: a store that cannot be read answers None rather than raising. Falling back to the deployment default is what the caller would have done anyway, so a store hiccup degrades to today's behaviour instead of failing an index or a search.

Cost class: O(one record) — a single slug-keyed settings read.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/embedder.py
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
def project_embedding_model(store: Any, slug: str) -> str | None:
    """The embedding model *slug* is indexed and searched with, or ``None``.

    ``None`` means "inherit ``wiki.embedding.model``" — the answer for every
    project indexed before the override existed, and for every project whose
    operator never set one.

    Best-effort by construction: a store that cannot be read answers ``None``
    rather than raising. Falling back to the deployment default is what the
    caller would have done anyway, so a store hiccup degrades to today's
    behaviour instead of failing an index or a search.

    Cost class: ``O(one record)`` — a single slug-keyed settings read.
    """
    if not slug:
        return None
    try:
        settings = store.get_project_settings(slug)
    except Exception:  # noqa: BLE001 - see the best-effort contract above
        return None
    return getattr(settings, "embedding_model", None) if settings else None

mewbo_graph.wiki.retriever

HybridRetriever — BM25 + cosine + RRF fusion + graph/memory expansion.

Operates over two base candidate sets: wiki pages (text bodies) and graph nodes (name + docstring text). BM25 always runs over both. Vector cosine only over graph nodes (pages aren't embedded in v1). Final ranking is reciprocal-rank-fusion (RRF, k=60).

The memory multiplex layer is an additive overlay: with memory_expand, MultiplexExpander seeds atomic memory notes by cosine, then follows each note's ANCHORS edges back to code entities (+ their 1-hop structural neighbours), additive-fusing a small w_ppr booster (GAAMA's 0.1·ppr + 1.0·sim). Hubs are degree-damped. memory_expand=False skips the overlay entirely, leaving the RRF ranking above untouched.

HybridHit dataclass

Single ranked result returned by HybridRetriever.search.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/retriever.py
38
39
40
41
42
43
44
45
46
@dataclass(frozen=True)
class HybridHit:
    """Single ranked result returned by HybridRetriever.search."""

    kind: Literal["page", "node", "memory"]
    id: str
    score: float
    snippet: str
    metadata: dict[str, Any] = field(default_factory=dict)

HybridRetriever

Atomic retriever — store + embedder at construction; one public search method.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/retriever.py
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
class HybridRetriever:
    """Atomic retriever — store + embedder at construction; one public search method."""

    def __init__(
        self,
        *,
        store: WikiStoreBase,
        embedder: EmbedderProtocol,
        expander: MultiplexExpander | None = None,
    ) -> None:
        """Initialise with a store backend, an embedder, and optional expander."""
        self.store = store
        self.embedder = embedder
        self._expander = expander

    def search(
        self,
        slug: str,
        query: str,
        *,
        k: int = 10,
        types: list[str] | None = None,
        graph_expand: bool = False,
        sources: Literal["pages", "graph", "both"] = "both",
        memory_expand: bool = False,
        memory_filters: MemoryFilter | None = None,
    ) -> list[HybridHit]:
        """Return up to *k* results fused from BM25, cosine, graph/memory expansion.

        Args:
            slug: Project slug (e.g. "org/repo").
            query: Free-text search query.
            k: Maximum number of results to return.
            types: Filter graph candidates to these node types (e.g. ["Class"]).
            graph_expand: When True, 1-hop neighbours of top graph hits are added.
            sources: Which candidate pools to include — "pages", "graph", or "both".
            memory_expand: When True, additive-fuse the memory multiplex layer.
            memory_filters: Optional ``MemoryFilter`` applied to memory seeding.
        """
        # Resolve the commit scope ONCE per search and thread it down. The graph
        # helpers below run inside per-hit loops, so resolving it there would
        # turn one project read into one per candidate.
        scope = self.store.live_scope(slug)

        # 1. Collect candidates.
        page_candidates = (
            self._page_candidates(slug) if sources in {"pages", "both"} else []
        )
        node_candidates = (
            self._graph_candidates(slug, types, scope)
            if sources in {"graph", "both"}
            else []
        )

        corpus = page_candidates + node_candidates
        if corpus:
            # 2. BM25 score over the whole corpus (pages + nodes).
            bm25_ranks = _bm25_ranks(query, [c["text"] for c in corpus])

            # 3. Cosine score over graph nodes only (pages aren't embedded in v1).
            cos_ranks: dict[int, int] = {}
            if node_candidates:
                qvec = self.embedder.embed_query(query)
                top_embs = self.store.vector_search(slug, qvec=qvec, k=max(k * 3, 30))
                id_to_rank = {emb.node_id: r for r, emb in enumerate(top_embs)}
                for i, c in enumerate(corpus):
                    if c["kind"] == "node" and c["id"] in id_to_rank:
                        cos_ranks[i] = id_to_rank[c["id"]]

            # 4. RRF fusion — sum 1/(rrf_k + rank) across rankers.
            fused: dict[int, float] = {}
            for i, r in enumerate(bm25_ranks):
                fused[i] = fused.get(i, 0.0) + 1.0 / (_RRF_K + r + 1)
            for i, r in cos_ranks.items():
                fused[i] = fused.get(i, 0.0) + 1.0 / (_RRF_K + r + 1)

            # 5. Sort descending by fused score; take top-k.
            ranked = sorted(fused.items(), key=lambda t: t[1], reverse=True)
            hits = [_to_hit(corpus[i], score=s) for i, s in ranked[:k]]

            # 6. Optional 1-hop graph expansion over top-3 graph hits.
            if graph_expand:
                hits = self._expand_neighbors(slug, hits, k=k, scope=scope)
        else:
            hits = []

        # 7. Optional memory multiplex overlay (additive fusion).
        if memory_expand:
            hits = self._fuse_memory(slug, query, hits, k=k, filt=memory_filters)

        return hits

    def _fuse_memory(
        self,
        slug: str,
        query: str,
        base: list[HybridHit],
        *,
        k: int,
        filt: MemoryFilter | None,
    ) -> list[HybridHit]:
        """Seed memory by cosine, expand to anchored code, additive-fuse into *base*."""
        expander = self._expander
        if expander is None:
            expander = self._expander = MultiplexExpander.from_store(self.store)
        qvec = self.embedder.embed_query(query)
        extra = expander.expand(slug, qvec, k=k, filt=filt)
        return _merge_hits(base, extra, k=k)

    # -- Private helpers -------------------------------------------------------

    def _page_candidates(self, slug: str) -> list[dict]:
        out = []
        for p in self.store.list_pages(slug):
            text = (p.body or "")[:_PAGE_BODY_CAP]
            snippet = (text[:200] + "…") if len(text) > 200 else text
            out.append({
                "kind": "page",
                "id": p.id,
                "text": text,
                "snippet": snippet,
                "metadata": {"title": p.title},
            })
        return out

    def _graph_candidates(
        self, slug: str, types: list[str] | None, scope: CommitScope
    ) -> list[dict]:
        if types:
            nodes = []
            for t in types:
                nodes.extend(self.store.query_graph(slug, scope=scope, node_type=t))
        else:
            nodes = self.store.query_graph(slug, scope=scope)
        out = []
        for n in nodes:
            text = (n.name + " " + (n.docstring or "")).strip()
            out.append({
                "kind": "node",
                "id": n.node_id,
                "text": text,
                "snippet": text[:200],
                "metadata": {"type": n.type, "name": n.name, "file": n.file},
            })
        return out

    def _expand_neighbors(
        self, slug: str, hits: list[HybridHit], *, k: int, scope: CommitScope
    ) -> list[HybridHit]:
        """Add 1-hop graph neighbours for the top-3 node hits, with a score bonus."""
        seen: set[tuple[str, str]] = {(h.kind, h.id) for h in hits}
        bonus = (hits[0].score * 0.5) if hits else 0.0
        added: list[HybridHit] = []
        for h in hits[:3]:  # budget: expand only top-3 to avoid fanout
            if h.kind != "node":
                continue
            for n in self.store.query_graph(slug, scope=scope, neighbors_of=h.id):
                key = ("node", n.node_id)
                if key in seen:
                    continue
                text = (n.name + " " + (n.docstring or "")).strip()
                added.append(HybridHit(
                    kind="node",
                    id=n.node_id,
                    score=bonus,
                    snippet=text[:200],
                    metadata={
                        "type": n.type,
                        "name": n.name,
                        "file": n.file,
                        "expanded": True,
                    },
                ))
                seen.add(key)
        combined = hits + added
        combined.sort(key=lambda h: h.score, reverse=True)
        return combined[:k + len(added)]

__init__(*, store: WikiStoreBase, embedder: EmbedderProtocol, expander: MultiplexExpander | None = None) -> None

Initialise with a store backend, an embedder, and optional expander.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/retriever.py
52
53
54
55
56
57
58
59
60
61
62
def __init__(
    self,
    *,
    store: WikiStoreBase,
    embedder: EmbedderProtocol,
    expander: MultiplexExpander | None = None,
) -> None:
    """Initialise with a store backend, an embedder, and optional expander."""
    self.store = store
    self.embedder = embedder
    self._expander = expander

search(slug: str, query: str, *, k: int = 10, types: list[str] | None = None, graph_expand: bool = False, sources: Literal['pages', 'graph', 'both'] = 'both', memory_expand: bool = False, memory_filters: MemoryFilter | None = None) -> list[HybridHit]

Return up to k results fused from BM25, cosine, graph/memory expansion.

Parameters:

Name Type Description Default
slug str

Project slug (e.g. "org/repo").

required
query str

Free-text search query.

required
k int

Maximum number of results to return.

10
types list[str] | None

Filter graph candidates to these node types (e.g. ["Class"]).

None
graph_expand bool

When True, 1-hop neighbours of top graph hits are added.

False
sources Literal['pages', 'graph', 'both']

Which candidate pools to include — "pages", "graph", or "both".

'both'
memory_expand bool

When True, additive-fuse the memory multiplex layer.

False
memory_filters MemoryFilter | None

Optional MemoryFilter applied to memory seeding.

None
Source code in packages/mewbo_graph/src/mewbo_graph/wiki/retriever.py
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
def search(
    self,
    slug: str,
    query: str,
    *,
    k: int = 10,
    types: list[str] | None = None,
    graph_expand: bool = False,
    sources: Literal["pages", "graph", "both"] = "both",
    memory_expand: bool = False,
    memory_filters: MemoryFilter | None = None,
) -> list[HybridHit]:
    """Return up to *k* results fused from BM25, cosine, graph/memory expansion.

    Args:
        slug: Project slug (e.g. "org/repo").
        query: Free-text search query.
        k: Maximum number of results to return.
        types: Filter graph candidates to these node types (e.g. ["Class"]).
        graph_expand: When True, 1-hop neighbours of top graph hits are added.
        sources: Which candidate pools to include — "pages", "graph", or "both".
        memory_expand: When True, additive-fuse the memory multiplex layer.
        memory_filters: Optional ``MemoryFilter`` applied to memory seeding.
    """
    # Resolve the commit scope ONCE per search and thread it down. The graph
    # helpers below run inside per-hit loops, so resolving it there would
    # turn one project read into one per candidate.
    scope = self.store.live_scope(slug)

    # 1. Collect candidates.
    page_candidates = (
        self._page_candidates(slug) if sources in {"pages", "both"} else []
    )
    node_candidates = (
        self._graph_candidates(slug, types, scope)
        if sources in {"graph", "both"}
        else []
    )

    corpus = page_candidates + node_candidates
    if corpus:
        # 2. BM25 score over the whole corpus (pages + nodes).
        bm25_ranks = _bm25_ranks(query, [c["text"] for c in corpus])

        # 3. Cosine score over graph nodes only (pages aren't embedded in v1).
        cos_ranks: dict[int, int] = {}
        if node_candidates:
            qvec = self.embedder.embed_query(query)
            top_embs = self.store.vector_search(slug, qvec=qvec, k=max(k * 3, 30))
            id_to_rank = {emb.node_id: r for r, emb in enumerate(top_embs)}
            for i, c in enumerate(corpus):
                if c["kind"] == "node" and c["id"] in id_to_rank:
                    cos_ranks[i] = id_to_rank[c["id"]]

        # 4. RRF fusion — sum 1/(rrf_k + rank) across rankers.
        fused: dict[int, float] = {}
        for i, r in enumerate(bm25_ranks):
            fused[i] = fused.get(i, 0.0) + 1.0 / (_RRF_K + r + 1)
        for i, r in cos_ranks.items():
            fused[i] = fused.get(i, 0.0) + 1.0 / (_RRF_K + r + 1)

        # 5. Sort descending by fused score; take top-k.
        ranked = sorted(fused.items(), key=lambda t: t[1], reverse=True)
        hits = [_to_hit(corpus[i], score=s) for i, s in ranked[:k]]

        # 6. Optional 1-hop graph expansion over top-3 graph hits.
        if graph_expand:
            hits = self._expand_neighbors(slug, hits, k=k, scope=scope)
    else:
        hits = []

    # 7. Optional memory multiplex overlay (additive fusion).
    if memory_expand:
        hits = self._fuse_memory(slug, query, hits, k=k, filt=memory_filters)

    return hits

MultiplexExpander

Cross-layer retrieval: memory seeds → anchored code → structural hops.

Atomic, injectable (store + optional structure provider). Seeds atomic memory notes by cosine (memory_vector_search, MemoryFilter-aware, invalidated-excluded), then for each note follows its ANCHORS edges back to live code entities and their ≤expansion_hops structural neighbours, scoring an additive w_ppr booster and damping hub nodes (degree > hub_degree).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/retriever.py
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
class MultiplexExpander:
    """Cross-layer retrieval: memory seeds → anchored code → structural hops.

    Atomic, injectable (store + optional structure provider). Seeds atomic
    memory notes by cosine (``memory_vector_search``, MemoryFilter-aware,
    invalidated-excluded), then for each note follows its ``ANCHORS`` edges
    back to live code entities and their ≤``expansion_hops`` structural
    neighbours, scoring an additive ``w_ppr`` booster and damping hub nodes
    (degree > ``hub_degree``).
    """

    def __init__(
        self,
        *,
        store: WikiStoreBase,
        provider: StructureProvider | None = None,
        w_ppr: float = 0.1,
        hub_degree: int = 50,
        expansion_hops: int = 1,
        rrf_k: int = _RRF_K,
    ) -> None:
        """Inject store + (lazy) structure provider and fusion knobs."""
        self.store = store
        self._provider = provider
        self.w_ppr = w_ppr
        self.hub_degree = hub_degree
        self.expansion_hops = expansion_hops
        self.rrf_k = rrf_k

    @classmethod
    def from_store(
        cls,
        store: WikiStoreBase,
        *,
        provider: StructureProvider | None = None,
        w_ppr: float | None = None,
        hub_degree: int | None = None,
        expansion_hops: int | None = None,
    ) -> MultiplexExpander:
        """Build an expander with the ``wiki.memory.*`` fusion knobs applied.

        The one composition root a production caller reaches (``HybridRetriever``
        builds its default expander through it), so the knobs are read HERE and
        passed down as arguments rather than inside the class. Each constructor
        default already equals its config default, which is exactly what made
        the gap invisible: an expander built with none of them behaves like a
        correctly configured one on a default deployment, so an operator who
        retunes ``hub_degree`` sees no error and no effect. Keeping the read out
        of the class body also keeps it pure DI — a test constructs it directly
        and is never at the mercy of a deployment setting it cannot see.

        Config decides the DEFAULT, never "whether": an explicitly passed knob
        (or an expander injected into ``HybridRetriever``) still wins.
        """
        return cls(
            store=store,
            provider=provider,
            w_ppr=w_ppr if w_ppr is not None else float(cls._setting("fusion_w_ppr", 0.1)),
            hub_degree=(
                hub_degree if hub_degree is not None else int(cls._setting("hub_degree", 50))
            ),
            expansion_hops=(
                expansion_hops
                if expansion_hops is not None
                else int(cls._setting("expansion_hops", 1))
            ),
        )

    @staticmethod
    def _setting(field: str, default: float) -> float:
        """Read one ``wiki.memory.*`` knob, falling back to *default*.

        The default at each call site is the constructor's own, so a
        config-less caller lands on exactly the behaviour it had before these
        knobs were wired.
        """
        value = get_config_value("wiki", "memory", field, default=default)
        return default if value is None else value

    @property
    def provider(self) -> StructureProvider:
        """Lazily build the default code structure provider."""
        if self._provider is None:
            from .structure_provider import CodeStructureProvider
            self._provider = CodeStructureProvider(self.store)
        return self._provider

    def expand(
        self,
        slug: str,
        query_vec: list[float],
        *,
        k: int = 10,
        filt: MemoryFilter | None = None,
    ) -> list[HybridHit]:
        """Return memory-seed hits + their anchored/expanded code hits."""
        # One project read per expansion — ``_damp`` and ``_expand_neighbours``
        # below run per anchored code node.
        scope = self.store.live_scope(slug)
        seeds = self.store.memory_vector_search(
            slug, query_vec, k=k, filt=filt or MemoryFilter()
        )
        # Gather every seed's live ANCHORS edges, then resolve all targets in
        # ONE graph pass (resolve_many) instead of a scan per anchor — bounds
        # expansion to a single O(N) provider hit regardless of seed/anchor fanout.
        seed_nodes: list[tuple[float, MemoryNode]] = []
        anchors_by_seed: dict[str, list[str]] = {}
        for rank, emb in enumerate(seeds):
            node = self.store.get_memory_node(slug, emb.node_id)
            if node is None:
                continue
            seed_nodes.append((1.0 / (self.rrf_k + rank + 1), node))
            anchors_by_seed[node.node_id] = [
                e.target
                for e in self.store.list_memory_edges(slug, node_id=node.node_id)
                if e.type == "ANCHORS"
            ]
        all_targets = {t for targets in anchors_by_seed.values() for t in targets}
        resolved = self.provider.resolve_many(slug, list(all_targets))

        hits: list[HybridHit] = []
        for seed_score, node in seed_nodes:
            hits.append(
                HybridHit(
                    kind="memory",
                    id=node.node_id,
                    score=seed_score,
                    snippet=node.content[:200],
                    metadata={
                        "kind": node.kind,
                        "labels": node.labels,
                        "source": node.provenance.source,
                    },
                )
            )
            for target in anchors_by_seed[node.node_id]:
                code = resolved.get(target)
                if code is None:
                    continue
                damp = self._damp(slug, code.node_id, scope)
                base = self.w_ppr * seed_score * damp
                hits.append(self._node_hit(code, base, via=node.node_id))
                hits.extend(
                    self._expand_neighbours(
                        slug, code.node_id, base * 0.5, node.node_id, scope
                    )
                )
        return hits

    def _expand_neighbours(
        self, slug: str, start_id: str, score: float, via: str, scope: CommitScope
    ) -> list[HybridHit]:
        """Bounded BFS over ≤``expansion_hops`` structural neighbours."""
        out: list[HybridHit] = []
        seen = {start_id}
        frontier = [start_id]
        for _ in range(self.expansion_hops):
            nxt: list[str] = []
            for nid in frontier:
                for nb in self.store.query_graph(slug, scope=scope, neighbors_of=nid):
                    if nb.node_id in seen:
                        continue
                    seen.add(nb.node_id)
                    nxt.append(nb.node_id)
                    out.append(self._node_hit(nb, score, via=via, expanded=True))
            frontier = nxt
        return out

    def _damp(self, slug: str, node_id: str, scope: CommitScope) -> float:
        """Hub damping: 1.0 below threshold, else ``hub_degree / degree``."""
        degree = len(self.store.query_graph(slug, scope=scope, neighbors_of=node_id))
        if degree <= self.hub_degree:
            return 1.0
        return self.hub_degree / degree

    @staticmethod
    def _node_hit(
        node: GraphNode, score: float, *, via: str, expanded: bool = False
    ) -> HybridHit:
        text = (node.name + " " + (node.docstring or "")).strip()
        meta: dict[str, Any] = {
            "type": node.type,
            "name": node.name,
            "file": node.file,
            "via_memory": via,
        }
        if expanded:
            meta["expanded"] = True
        return HybridHit(
            kind="node", id=node.node_id, score=score, snippet=text[:200], metadata=meta
        )

provider: StructureProvider property

Lazily build the default code structure provider.

__init__(*, store: WikiStoreBase, provider: StructureProvider | None = None, w_ppr: float = 0.1, hub_degree: int = 50, expansion_hops: int = 1, rrf_k: int = _RRF_K) -> None

Inject store + (lazy) structure provider and fusion knobs.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/retriever.py
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
def __init__(
    self,
    *,
    store: WikiStoreBase,
    provider: StructureProvider | None = None,
    w_ppr: float = 0.1,
    hub_degree: int = 50,
    expansion_hops: int = 1,
    rrf_k: int = _RRF_K,
) -> None:
    """Inject store + (lazy) structure provider and fusion knobs."""
    self.store = store
    self._provider = provider
    self.w_ppr = w_ppr
    self.hub_degree = hub_degree
    self.expansion_hops = expansion_hops
    self.rrf_k = rrf_k

expand(slug: str, query_vec: list[float], *, k: int = 10, filt: MemoryFilter | None = None) -> list[HybridHit]

Return memory-seed hits + their anchored/expanded code hits.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/retriever.py
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
def expand(
    self,
    slug: str,
    query_vec: list[float],
    *,
    k: int = 10,
    filt: MemoryFilter | None = None,
) -> list[HybridHit]:
    """Return memory-seed hits + their anchored/expanded code hits."""
    # One project read per expansion — ``_damp`` and ``_expand_neighbours``
    # below run per anchored code node.
    scope = self.store.live_scope(slug)
    seeds = self.store.memory_vector_search(
        slug, query_vec, k=k, filt=filt or MemoryFilter()
    )
    # Gather every seed's live ANCHORS edges, then resolve all targets in
    # ONE graph pass (resolve_many) instead of a scan per anchor — bounds
    # expansion to a single O(N) provider hit regardless of seed/anchor fanout.
    seed_nodes: list[tuple[float, MemoryNode]] = []
    anchors_by_seed: dict[str, list[str]] = {}
    for rank, emb in enumerate(seeds):
        node = self.store.get_memory_node(slug, emb.node_id)
        if node is None:
            continue
        seed_nodes.append((1.0 / (self.rrf_k + rank + 1), node))
        anchors_by_seed[node.node_id] = [
            e.target
            for e in self.store.list_memory_edges(slug, node_id=node.node_id)
            if e.type == "ANCHORS"
        ]
    all_targets = {t for targets in anchors_by_seed.values() for t in targets}
    resolved = self.provider.resolve_many(slug, list(all_targets))

    hits: list[HybridHit] = []
    for seed_score, node in seed_nodes:
        hits.append(
            HybridHit(
                kind="memory",
                id=node.node_id,
                score=seed_score,
                snippet=node.content[:200],
                metadata={
                    "kind": node.kind,
                    "labels": node.labels,
                    "source": node.provenance.source,
                },
            )
        )
        for target in anchors_by_seed[node.node_id]:
            code = resolved.get(target)
            if code is None:
                continue
            damp = self._damp(slug, code.node_id, scope)
            base = self.w_ppr * seed_score * damp
            hits.append(self._node_hit(code, base, via=node.node_id))
            hits.extend(
                self._expand_neighbours(
                    slug, code.node_id, base * 0.5, node.node_id, scope
                )
            )
    return hits

from_store(store: WikiStoreBase, *, provider: StructureProvider | None = None, w_ppr: float | None = None, hub_degree: int | None = None, expansion_hops: int | None = None) -> MultiplexExpander classmethod

Build an expander with the wiki.memory.* fusion knobs applied.

The one composition root a production caller reaches (HybridRetriever builds its default expander through it), so the knobs are read HERE and passed down as arguments rather than inside the class. Each constructor default already equals its config default, which is exactly what made the gap invisible: an expander built with none of them behaves like a correctly configured one on a default deployment, so an operator who retunes hub_degree sees no error and no effect. Keeping the read out of the class body also keeps it pure DI — a test constructs it directly and is never at the mercy of a deployment setting it cannot see.

Config decides the DEFAULT, never "whether": an explicitly passed knob (or an expander injected into HybridRetriever) still wins.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/retriever.py
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
@classmethod
def from_store(
    cls,
    store: WikiStoreBase,
    *,
    provider: StructureProvider | None = None,
    w_ppr: float | None = None,
    hub_degree: int | None = None,
    expansion_hops: int | None = None,
) -> MultiplexExpander:
    """Build an expander with the ``wiki.memory.*`` fusion knobs applied.

    The one composition root a production caller reaches (``HybridRetriever``
    builds its default expander through it), so the knobs are read HERE and
    passed down as arguments rather than inside the class. Each constructor
    default already equals its config default, which is exactly what made
    the gap invisible: an expander built with none of them behaves like a
    correctly configured one on a default deployment, so an operator who
    retunes ``hub_degree`` sees no error and no effect. Keeping the read out
    of the class body also keeps it pure DI — a test constructs it directly
    and is never at the mercy of a deployment setting it cannot see.

    Config decides the DEFAULT, never "whether": an explicitly passed knob
    (or an expander injected into ``HybridRetriever``) still wins.
    """
    return cls(
        store=store,
        provider=provider,
        w_ppr=w_ppr if w_ppr is not None else float(cls._setting("fusion_w_ppr", 0.1)),
        hub_degree=(
            hub_degree if hub_degree is not None else int(cls._setting("hub_degree", 50))
        ),
        expansion_hops=(
            expansion_hops
            if expansion_hops is not None
            else int(cls._setting("expansion_hops", 1))
        ),
    )

mewbo_graph.wiki.memory

Insight ingestion — the DRY write core behind all memory surfaces.

One InsightIngestor backs the SessionTool, the REST endpoint, and the MCP tool. Given a claim (or raw text to condense) plus optional anchors, it: condenses → embeds → resolves/auto-resolves anchors → runs the 3-tier dedup/merge ladder → upserts node + embedding + edges. Every collaborator (store, embedder, structure provider, deduper, condenser, clock) is constructor-injected so the core is unit-testable with stubs only at the LLM/embedding I/O boundary.

Atomicity is load-bearing: a condensed blob becomes several ≤200-char notes, and a merge keeps the crisper of two overlapping notes (Molecular Facts / AtomicRAG) rather than concatenating them.

AnchorResolver

Bases: Protocol

The node-agnostic anchor seam the ingestor depends on.

The ingestor only needs to map a node_id back to its entity_key and to test which anchor keys resolve to a live unit — it never inspects the resolved node itself. Typing that surface as Mapping[..., object] (a covariant value) lets any corpus's provider satisfy it: the wiki CodeStructureProvider (code graph) and the SCG ScgAnchorResolver (connector graph) both conform without a cast, so an alternate corpus plugs its own node type in cleanly. StructureProvider is a structural subtype.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
@runtime_checkable
class AnchorResolver(Protocol):
    """The node-agnostic anchor seam the ingestor depends on.

    The ingestor only needs to map a ``node_id`` back to its ``entity_key`` and
    to test which anchor keys resolve to a *live* unit — it never inspects the
    resolved node itself. Typing that surface as ``Mapping[..., object]`` (a
    covariant value) lets any corpus's provider satisfy it: the wiki
    ``CodeStructureProvider`` (code graph) and the SCG ``ScgAnchorResolver``
    (connector graph) both conform without a cast, so an alternate corpus plugs
    its own node type in cleanly. ``StructureProvider`` is a structural subtype.
    """

    def resolve_many(
        self, slug: str, entity_keys: list[EntityKey]
    ) -> Mapping[EntityKey, object]:
        """Resolve a batch of anchor keys; misses are omitted from the result."""
        ...

    def entity_key_of(self, slug: str, node_id: str) -> EntityKey | None:
        """Return the ``entity_key`` for a resolved ``node_id``, or None."""
        ...

entity_key_of(slug: str, node_id: str) -> EntityKey | None

Return the entity_key for a resolved node_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
64
65
66
def entity_key_of(self, slug: str, node_id: str) -> EntityKey | None:
    """Return the ``entity_key`` for a resolved ``node_id``, or None."""
    ...

resolve_many(slug: str, entity_keys: list[EntityKey]) -> Mapping[EntityKey, object]

Resolve a batch of anchor keys; misses are omitted from the result.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
58
59
60
61
62
def resolve_many(
    self, slug: str, entity_keys: list[EntityKey]
) -> Mapping[EntityKey, object]:
    """Resolve a batch of anchor keys; misses are omitted from the result."""
    ...

DedupDecision dataclass

Verdict from the dedup ladder for a candidate note.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
160
161
162
163
164
165
166
@dataclass(frozen=True)
class DedupDecision:
    """Verdict from the dedup ladder for a candidate note."""

    action: Literal["new", "merge", "link"]
    target_node_id: str | None = None
    tier: str | None = None

IngestResult

Bases: BaseModel

Aggregate result of one ingest call (one+ claims).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
116
117
118
119
120
121
122
123
124
125
126
class IngestResult(BaseModel):
    """Aggregate result of one ``ingest`` call (one+ claims)."""

    model_config = _CFG

    claims: list[IngestedClaim]

    @property
    def ok(self) -> bool:
        """True if at least one claim was stored (not all rejected)."""
        return any(c.action != "rejected" for c in self.claims)

ok: bool property

True if at least one claim was stored (not all rejected).

IngestedClaim

Bases: BaseModel

Outcome for a single atomic claim processed by the ingestor.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
103
104
105
106
107
108
109
110
111
112
113
class IngestedClaim(BaseModel):
    """Outcome for a single atomic claim processed by the ingestor."""

    model_config = _CFG

    action: Literal["created", "merged", "linked", "rejected"]
    content: str
    node_id: str | None = None
    tier: str | None = None
    anchors: list[EntityKey] = []
    warnings: list[str] = []

InsightCondenser

LLM decomposition of raw text into atomic claims (raw path only).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
class InsightCondenser:
    """LLM decomposition of raw text into atomic claims (raw path only)."""

    def __init__(self, llm: Any) -> None:
        """Inject a chat model (``.invoke`` → text)."""
        self._llm = llm

    def condense(self, raw: str) -> list[str]:
        """Return atomic claims for *raw*. May raise — callers treat as non-fatal."""
        text = llm_text(self._llm, _CONDENSE_PROMPT.format(cap=MAX_INSIGHT_CHARS, raw=raw))
        claims: list[str] = []
        for line in text.splitlines():
            cleaned = line.strip().lstrip("-*0123456789.) ").strip()
            if cleaned:
                claims.append(cleaned)
        return claims

__init__(llm: Any) -> None

Inject a chat model (.invoke → text).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
142
143
144
def __init__(self, llm: Any) -> None:
    """Inject a chat model (``.invoke`` → text)."""
    self._llm = llm

condense(raw: str) -> list[str]

Return atomic claims for raw. May raise — callers treat as non-fatal.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
146
147
148
149
150
151
152
153
154
def condense(self, raw: str) -> list[str]:
    """Return atomic claims for *raw*. May raise — callers treat as non-fatal."""
    text = llm_text(self._llm, _CONDENSE_PROMPT.format(cap=MAX_INSIGHT_CHARS, raw=raw))
    claims: list[str] = []
    for line in text.splitlines():
        cleaned = line.strip().lstrip("-*0123456789.) ").strip()
        if cleaned:
            claims.append(cleaned)
    return claims

InsightDeduper

3-tier dedup/merge: exact node_id → fuzzy Jaccard → LLM over cosine-kNN.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
class InsightDeduper:
    """3-tier dedup/merge: exact node_id → fuzzy Jaccard → LLM over cosine-kNN."""

    def __init__(
        self,
        *,
        store: WikiStoreBase,
        llm: Any = None,
        fuzzy_jaccard: float = 0.85,
        dedup_k: int = 5,
        dedup_cosine: float = 0.6,
        min_token_len: int = 3,
    ) -> None:
        """Inject the store; *llm* optional (tier-3 degrades to NEW).

        Similarity uses the stateless ``Embedder.cosine`` + the store's
        ``memory_vector_search``, so no embedder instance is needed here.
        """
        self._store = store
        self._llm = llm
        self._fuzzy_jaccard = fuzzy_jaccard
        self._dedup_k = dedup_k
        self._dedup_cosine = dedup_cosine
        self._min_token_len = min_token_len

    def classify(
        self, slug: str, candidate: MemoryNode, *, candidate_vec: list[float] | None = None
    ) -> DedupDecision:
        """Classify *candidate* against existing notes (NONE-default → NEW).

        SCALE: tiers 2+3 run over the cosine-kNN candidate set (the single
        ``memory_vector_search`` ANN seam), NOT the whole store — so dedup is
        ``O(dedup_k)`` similarity work per ingest, and upgrading that one seam
        to a real ANN index makes the entire ladder sublinear. (A lexical
        near-duplicate is also an embedding near-duplicate, so scoping fuzzy
        to the kNN set never misses one.) Only when no embedding is available
        (BM25-only) does it fall back to a bounded full scan.

        DRY: tiers 2+3 delegate the block→score→decide step to the ONE generic
        ``ResolutionLadder`` shared with entity resolution. The injected
        strategies hold the tier thresholds EXACTLY — ``fuzzy_jaccard``
        ⇒ the merge band, ``dedup_cosine`` ⇒ the LLM (flag) band, NONE-default
        ⇒ NEW — so the ladder's bands and the tiers are one rule (see
        ``_build_ladder``).
        """
        # Tier 1 — exact: identical normalized content ⇒ identical node_id
        # (indexed point lookup, never a scan). Kept inline — the ladder is for
        # the *similarity* tiers; an exact id hit is a cheaper short-circuit.
        if self._store.get_memory_node(slug, candidate.node_id) is not None:
            return DedupDecision("merge", candidate.node_id, tier="exact")

        ladder, by_id = self._build_ladder(slug, candidate, candidate_vec)
        decision = ladder.decide(candidate.node_id)

        # A merge-band verdict ⇒ tier 2, the fuzzy (jaccard band) merge.
        if decision.action == "merge" and decision.target_id:
            return DedupDecision("merge", decision.target_id, tier="fuzzy")

        # A flag-band verdict ⇒ tier 3, the LLM call over the nearest
        # candidate above the cosine floor (only when a vector + LLM exist).
        if (
            decision.action == "flag"
            and decision.target_id
            and candidate_vec is not None
            and self._llm is not None
        ):
            target, _cos = by_id[decision.target_id]
            verdict = self._decide(candidate, target)
            if verdict == "merge":
                return DedupDecision("merge", target.node_id, tier="llm")
            if verdict == "link":
                return DedupDecision("link", target.node_id, tier="llm")
        return DedupDecision("new", tier="new")

    def _build_ladder(
        self, slug: str, candidate: MemoryNode, candidate_vec: list[float] | None
    ) -> tuple[ResolutionLadder, dict[str, tuple[MemoryNode, float]]]:
        """Build the shared ``ResolutionLadder`` over the cosine-kNN candidate set.

        The injected strategies encode the three tiers EXACTLY:

        - **block**: the kNN set from ``_nearest`` (bounded ANN seam), in cosine
          order — so flag-band ties break to the nearest, matching tier-3's
          ``above[0]`` selection.
        - **score**: ``max(jaccard_promoted, cosine_clamped)`` where a jaccard
          ``>= fuzzy_jaccard`` promotes to the auto band (tier-2: jaccard ALONE
          merges) and cosine is clamped strictly below the auto threshold so a
          high cosine can only ever reach the flag band (tier-3: cosine NEVER
          auto-merges — it routes to the LLM). Cosine below ``dedup_cosine``
          stays below the flag threshold ⇒ NEW.
        - **identity**: node_id (dedup consults no recommendation priors).

        Thresholds map ``auto_merge = fuzzy_jaccard`` and ``flag = dedup_cosine``.
        Returns the ladder plus the ``node_id -> (node, cosine)`` index the
        LLM-band branch needs to recover the target node.
        """
        ranked = self._nearest(slug, candidate, candidate_vec)
        by_id: dict[str, tuple[MemoryNode, float]] = {
            n.node_id: (n, c) for n, c in ranked if n.node_id != candidate.node_id
        }
        cand_tokens = _tokens(candidate.content, self._min_token_len)
        # Clamp cosine strictly below the auto-merge threshold so cosine alone
        # can reach the flag band but NEVER the merge band (the dedup invariant).
        cosine_cap = self._fuzzy_jaccard - _CLAMP_EPS

        def block(_key: str) -> list[tuple[str, str]]:
            return [(nid, node.content) for nid, (node, _c) in by_id.items()]

        def score(_key: str, nid: str) -> float:
            node, cos = by_id[nid]
            jac = (
                self._fuzzy_jaccard
                if cand_tokens
                and _jaccard(cand_tokens, _tokens(node.content, self._min_token_len))
                >= self._fuzzy_jaccard
                else 0.0
            )
            return max(jac, min(cos, cosine_cap))

        ladder = ResolutionLadder(
            block=block,
            score=score,
            identity=lambda nid: nid,
            auto_merge=self._fuzzy_jaccard,
            flag=self._dedup_cosine,
        )
        return ladder, by_id

    def _nearest(
        self, slug: str, candidate: MemoryNode, candidate_vec: list[float] | None
    ) -> list[tuple[MemoryNode, float]]:
        """Return ``(node, cosine)`` candidates for the dedup tiers.

        With an embedding: the cosine-kNN set via the ANN seam (bounded). Without
        one (BM25-only): a full scan paired with a 0.0 score so the fuzzy tier
        still runs and tier-3 (which needs a vector) is naturally skipped.
        """
        if candidate_vec is None:
            return [
                (n, 0.0)
                for n in self._store.query_memory(slug, filt=MemoryFilter())
                if n.node_id != candidate.node_id
            ]
        out: list[tuple[MemoryNode, float]] = []
        for emb in self._store.memory_vector_search(
            slug, candidate_vec, k=self._dedup_k, filt=MemoryFilter()
        ):
            if emb.node_id == candidate.node_id:
                continue
            node = self._store.get_memory_node(slug, emb.node_id)
            if node is not None:
                out.append((node, Embedder.cosine(candidate_vec, emb.vector)))
        return out

    def _decide(self, candidate: MemoryNode, target: MemoryNode) -> str:
        """Ask the LLM merge|link|new; any failure defaults to ``new``."""
        try:
            text = llm_text(
                self._llm,
                _DEDUP_PROMPT.format(a=candidate.content, b=target.content),
            ).strip().lower()
        except Exception:
            return "new"
        if "merge" in text:
            return "merge"
        if "link" in text:
            return "link"
        return "new"

__init__(*, store: WikiStoreBase, llm: Any = None, fuzzy_jaccard: float = 0.85, dedup_k: int = 5, dedup_cosine: float = 0.6, min_token_len: int = 3) -> None

Inject the store; llm optional (tier-3 degrades to NEW).

Similarity uses the stateless Embedder.cosine + the store's memory_vector_search, so no embedder instance is needed here.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
def __init__(
    self,
    *,
    store: WikiStoreBase,
    llm: Any = None,
    fuzzy_jaccard: float = 0.85,
    dedup_k: int = 5,
    dedup_cosine: float = 0.6,
    min_token_len: int = 3,
) -> None:
    """Inject the store; *llm* optional (tier-3 degrades to NEW).

    Similarity uses the stateless ``Embedder.cosine`` + the store's
    ``memory_vector_search``, so no embedder instance is needed here.
    """
    self._store = store
    self._llm = llm
    self._fuzzy_jaccard = fuzzy_jaccard
    self._dedup_k = dedup_k
    self._dedup_cosine = dedup_cosine
    self._min_token_len = min_token_len

classify(slug: str, candidate: MemoryNode, *, candidate_vec: list[float] | None = None) -> DedupDecision

Classify candidate against existing notes (NONE-default → NEW).

SCALE: tiers 2+3 run over the cosine-kNN candidate set (the single memory_vector_search ANN seam), NOT the whole store — so dedup is O(dedup_k) similarity work per ingest, and upgrading that one seam to a real ANN index makes the entire ladder sublinear. (A lexical near-duplicate is also an embedding near-duplicate, so scoping fuzzy to the kNN set never misses one.) Only when no embedding is available (BM25-only) does it fall back to a bounded full scan.

DRY: tiers 2+3 delegate the block→score→decide step to the ONE generic ResolutionLadder shared with entity resolution. The injected strategies hold the tier thresholds EXACTLY — fuzzy_jaccard ⇒ the merge band, dedup_cosine ⇒ the LLM (flag) band, NONE-default ⇒ NEW — so the ladder's bands and the tiers are one rule (see _build_ladder).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
def classify(
    self, slug: str, candidate: MemoryNode, *, candidate_vec: list[float] | None = None
) -> DedupDecision:
    """Classify *candidate* against existing notes (NONE-default → NEW).

    SCALE: tiers 2+3 run over the cosine-kNN candidate set (the single
    ``memory_vector_search`` ANN seam), NOT the whole store — so dedup is
    ``O(dedup_k)`` similarity work per ingest, and upgrading that one seam
    to a real ANN index makes the entire ladder sublinear. (A lexical
    near-duplicate is also an embedding near-duplicate, so scoping fuzzy
    to the kNN set never misses one.) Only when no embedding is available
    (BM25-only) does it fall back to a bounded full scan.

    DRY: tiers 2+3 delegate the block→score→decide step to the ONE generic
    ``ResolutionLadder`` shared with entity resolution. The injected
    strategies hold the tier thresholds EXACTLY — ``fuzzy_jaccard``
    ⇒ the merge band, ``dedup_cosine`` ⇒ the LLM (flag) band, NONE-default
    ⇒ NEW — so the ladder's bands and the tiers are one rule (see
    ``_build_ladder``).
    """
    # Tier 1 — exact: identical normalized content ⇒ identical node_id
    # (indexed point lookup, never a scan). Kept inline — the ladder is for
    # the *similarity* tiers; an exact id hit is a cheaper short-circuit.
    if self._store.get_memory_node(slug, candidate.node_id) is not None:
        return DedupDecision("merge", candidate.node_id, tier="exact")

    ladder, by_id = self._build_ladder(slug, candidate, candidate_vec)
    decision = ladder.decide(candidate.node_id)

    # A merge-band verdict ⇒ tier 2, the fuzzy (jaccard band) merge.
    if decision.action == "merge" and decision.target_id:
        return DedupDecision("merge", decision.target_id, tier="fuzzy")

    # A flag-band verdict ⇒ tier 3, the LLM call over the nearest
    # candidate above the cosine floor (only when a vector + LLM exist).
    if (
        decision.action == "flag"
        and decision.target_id
        and candidate_vec is not None
        and self._llm is not None
    ):
        target, _cos = by_id[decision.target_id]
        verdict = self._decide(candidate, target)
        if verdict == "merge":
            return DedupDecision("merge", target.node_id, tier="llm")
        if verdict == "link":
            return DedupDecision("link", target.node_id, tier="llm")
    return DedupDecision("new", tier="new")

InsightIngestor

DRY write core: condense → embed → anchor → dedup/merge → upsert.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
class InsightIngestor:
    """DRY write core: condense → embed → anchor → dedup/merge → upsert."""

    def __init__(
        self,
        *,
        store: WikiStoreBase,
        embedder: EmbedderProtocol,
        provider: AnchorResolver,
        deduper: InsightDeduper,
        condenser: InsightCondenser | None = None,
        clock: Any = None,
        max_anchors: int = 8,
        max_chars: int = MAX_INSIGHT_CHARS,
        auto_anchor_k: int = 3,
    ) -> None:
        """Wire collaborators (all injected); *condenser* is optional."""
        self._store = store
        self._embedder = embedder
        self._provider = provider
        self._deduper = deduper
        self._condenser = condenser
        self._clock = clock or utc_now_iso
        self._max_anchors = max_anchors
        self._max_chars = max_chars
        self._auto_anchor_k = auto_anchor_k

    @classmethod
    def from_store(
        cls,
        store: WikiStoreBase,
        *,
        embedder: EmbedderProtocol | None = None,
        slug: str | None = None,
        llm: Any = None,
        condenser: InsightCondenser | None = None,
        clock: Any = None,
        provider: AnchorResolver | None = None,
        deduper: InsightDeduper | None = None,
        max_anchors: int | None = None,
        max_chars: int | None = None,
    ) -> InsightIngestor:
        """Build an ingestor with the standard collaborators (DRY across surfaces).

        The single construction path shared by the SessionTool, the REST
        endpoint, and the MCP tool, so every surface dedups/anchors
        identically. ``embedder`` defaults to the wiki ``Embedder`` (BM25-only
        ``_NullEmbedder`` when none can be built); ``llm``/``condenser`` are
        opt-in (the in-session tool leaves them off — agents pre-atomize).
        ``provider`` overrides the default ``CodeStructureProvider`` so an
        alternate corpus (e.g. the SCG connector graph) can resolve its own
        anchors; ``None`` keeps the code-graph default (backward-compatible).

        The ``wiki.memory.*`` knobs are read HERE, at the one composition root
        every production caller goes through, and passed down as arguments. The
        failure mode that made this necessary is invisible: this class and
        ``InsightDeduper`` take each knob as a keyword argument whose default
        EQUALS the config default, so an ingestor built with none of them
        behaves exactly like a correctly configured one on a default
        deployment — an operator who retunes ``dedup_k`` or ``max_anchors``
        sees no error and no effect. Reading them here keeps both classes pure
        DI: a class that read config itself could not be constructed in a test
        without one, and a directly-constructed one would silently start
        obeying a deployment setting the test had no way to see.

        Config decides the DEFAULT, never "whether" — an explicitly passed
        ``deduper``/``max_anchors``/``max_chars`` still wins.
        """
        from .structure_provider import CodeStructureProvider

        if embedder is None:
            from .embedder import make_embedder_for_or_none

            embedder = make_embedder_for_or_none(store, slug or "") or _NullEmbedder()
        if deduper is None:
            deduper = InsightDeduper(
                store=store,
                llm=llm,
                fuzzy_jaccard=float(cls._setting("fuzzy_jaccard", 0.85)),
                dedup_k=int(cls._setting("dedup_k", 5)),
                dedup_cosine=float(cls._setting("dedup_cosine", 0.6)),
            )
        return cls(
            store=store,
            embedder=embedder,
            provider=provider or CodeStructureProvider(store),
            deduper=deduper,
            condenser=condenser,
            clock=clock,
            max_anchors=(
                max_anchors if max_anchors is not None else int(cls._setting("max_anchors", 8))
            ),
            max_chars=max_chars if max_chars is not None else cls._configured_max_chars(),
        )

    @staticmethod
    def _setting(field: str, default: float) -> float:
        """Read one ``wiki.memory.*`` knob, falling back to *default*.

        The default at each call site is the constructor's own, so a
        config-less caller lands on exactly the behaviour it had before these
        knobs were wired.
        """
        value = get_config_value("wiki", "memory", field, default=default)
        return default if value is None else value

    @staticmethod
    def _configured_max_chars() -> int:
        """``wiki.memory.max_insight_chars``, capped by the model's own ceiling.

        ``MemoryNode.content`` is declared ``max_length=MAX_INSIGHT_CHARS``, so
        200 is a hard ceiling no setting can lift: a larger cap would let a
        claim past this class's own length check and then raise a
        ``ValidationError`` out of :meth:`ingest`. Lowering it is the direction
        that works, and the one operators actually use.
        """
        raw = int(InsightIngestor._setting("max_insight_chars", MAX_INSIGHT_CHARS))
        return min(raw, MAX_INSIGHT_CHARS)

    def ingest(
        self,
        slug: str,
        content: str | None = None,
        *,
        raw: str | None = None,
        anchors: list[EntityKey] | None = None,
        links: list[str] | None = None,
        kind: MemoryKind = "propositional",
        labels: list[str] | None = None,
        corpus: str = "code",
        condense: bool = False,
        source: MemorySource = "indexer",
        author_agent: str = "insight",
        session_id: str | None = None,
    ) -> IngestResult:
        """Ingest a claim (or condensed raw blob) into the memory graph."""
        base_anchors = list(anchors or [])
        links = list(links or [])
        labels = list(labels or [])
        use_raw = raw is not None or condense
        claims, condense_warnings = self._claims_for(raw if raw is not None else content, use_raw)

        outcomes: list[IngestedClaim] = []
        if not claims:
            return IngestResult(
                claims=[
                    IngestedClaim(
                        action="rejected",
                        content=(raw or content or ""),
                        warnings=condense_warnings,
                    )
                ]
            )
        for claim in claims:
            outcomes.append(
                self._ingest_claim(
                    slug,
                    claim,
                    base_anchors=base_anchors,
                    links=links,
                    kind=kind,
                    labels=labels,
                    corpus=corpus,
                    source=source,
                    author_agent=author_agent,
                    session_id=session_id,
                    auto_anchor=use_raw,
                    seed_warnings=condense_warnings,
                )
            )
        return IngestResult(claims=outcomes)

    # -- claim derivation ----------------------------------------------------

    def _claims_for(self, text: str | None, use_raw: bool) -> tuple[list[str], list[str]]:
        """Split input into atomic claims; return (claims, warnings)."""
        text = (text or "").strip()
        if not text:
            return [], ["empty insight"]
        if not use_raw:
            return [text], []
        if self._condenser is not None:
            try:
                claims = [c for c in self._condenser.condense(text) if c.strip()]
                if claims:
                    return claims, []
            except Exception:
                pass  # non-fatal — fall through to single-claim fallback
        if len(text) <= self._max_chars:
            return [text], ["condenser unavailable; stored raw as a single claim"]
        return [], [f"insight exceeds {self._max_chars} chars and no condenser is available"]

    # -- per-claim pipeline --------------------------------------------------

    def _ingest_claim(
        self,
        slug: str,
        claim: str,
        *,
        base_anchors: list[EntityKey],
        links: list[str],
        kind: MemoryKind,
        labels: list[str],
        corpus: str,
        source: MemorySource,
        author_agent: str,
        session_id: str | None,
        auto_anchor: bool,
        seed_warnings: list[str],
    ) -> IngestedClaim:
        warnings = list(seed_warnings)
        claim = claim.strip()
        if not claim:
            return IngestedClaim(
                action="rejected", content=claim, warnings=warnings + ["empty claim"]
            )
        if len(claim) > self._max_chars:
            return IngestedClaim(
                action="rejected",
                content=claim,
                warnings=warnings + [f"claim exceeds {self._max_chars} chars"],
            )

        now = self._clock()
        candidate = MemoryNode(
            slug=slug,
            content=claim,
            kind=kind,
            labels=labels,
            corpus=corpus,
            provenance=MemoryProvenance(
                author_agent=author_agent,
                source=source,
                session_id=session_id,
                created_at=now,
                updated_at=now,
            ),
        )

        anchor_keys = list(base_anchors)
        if auto_anchor and not anchor_keys:
            anchor_keys = self._auto_anchor(slug, claim, warnings)
        if len(anchor_keys) > self._max_anchors:
            warnings.append(f"anchors capped to {self._max_anchors}")
            anchor_keys = anchor_keys[: self._max_anchors]
        resolved = self._resolve_anchors(slug, anchor_keys, warnings)

        vec = self._embed(candidate, warnings)
        decision = self._deduper.classify(slug, candidate, candidate_vec=vec)

        if decision.action == "merge" and decision.target_node_id:
            return self._apply_merge(slug, candidate, decision, resolved, links, now, warnings)
        if decision.action == "link" and decision.target_node_id:
            return self._apply_new(
                slug, candidate, resolved, links, vec, now, warnings,
                action="linked", relate_to=[decision.target_node_id], tier=decision.tier,
            )
        return self._apply_new(
            slug, candidate, resolved, links, vec, now, warnings, action="created"
        )

    # -- helpers -------------------------------------------------------------

    def _auto_anchor(self, slug: str, claim: str, warnings: list[str]) -> list[EntityKey]:
        """Embed the claim, NN-search code embeddings, map hits → entity_keys."""
        try:
            qvec = self._embedder.embed_query(claim)
        except Exception:
            return []
        keys: list[EntityKey] = []
        for emb in self._store.vector_search(slug, qvec, k=self._auto_anchor_k):
            key = self._provider.entity_key_of(slug, emb.node_id)
            if key and key not in keys:
                keys.append(key)
        if keys:
            noun = "entity" if len(keys) == 1 else "entities"
            warnings.append(f"auto-anchored to {len(keys)} code {noun}")
        return keys

    def _resolve_anchors(
        self, slug: str, keys: list[EntityKey], warnings: list[str]
    ) -> list[EntityKey]:
        """Drop anchors that don't resolve to a live code node."""
        if not keys:
            return []
        resolved = self._provider.resolve_many(slug, keys)
        out: list[EntityKey] = []
        for key in keys:
            if key in resolved:
                if key not in out:
                    out.append(key)
            else:
                warnings.append(f"dropped unresolved anchor: {key}")
        return out

    def _embed(self, node: MemoryNode, warnings: list[str]) -> list[float] | None:
        """Embed a node's content; non-fatal (None ⇒ BM25-only)."""
        try:
            rows = self._embedder.embed_nodes([(node.node_id, node.content)], slug=node.slug)
        except Exception:
            warnings.append("embedding unavailable; BM25 fallback")
            return None
        return list(rows[0].vector) if rows else None

    def _store_embedding(self, node: MemoryNode, vec: list[float] | None) -> None:
        if vec is None:
            return
        self._store.upsert_memory_embeddings(
            node.slug,
            [
                MemoryEmbedding(
                    slug=node.slug,
                    node_id=node.node_id,
                    vector=vec,
                    model=getattr(self._embedder, "model", ""),
                    dim=len(vec),
                )
            ],
        )

    def _apply_new(
        self,
        slug: str,
        node: MemoryNode,
        anchor_keys: list[EntityKey],
        link_ids: list[str],
        vec: list[float] | None,
        now: str,
        warnings: list[str],
        *,
        action: Literal["created", "linked"],
        relate_to: list[str] | None = None,
        tier: str | None = None,
    ) -> IngestedClaim:
        self._store.upsert_memory_nodes(slug, [node])
        self._store_embedding(node, vec)
        edges = self._anchor_edges(slug, node.node_id, anchor_keys, now)
        edges += self._relate_edges(slug, node.node_id, list(link_ids) + list(relate_to or []), now)
        if edges:
            self._store.upsert_memory_edges(slug, edges)
        return IngestedClaim(
            action=action, node_id=node.node_id, content=node.content,
            anchors=anchor_keys, tier=tier, warnings=warnings,
        )

    def _apply_merge(
        self,
        slug: str,
        candidate: MemoryNode,
        decision: DedupDecision,
        anchor_keys: list[EntityKey],
        link_ids: list[str],
        now: str,
        warnings: list[str],
    ) -> IngestedClaim:
        target = self._store.get_memory_node(slug, decision.target_node_id or "")
        if target is None:  # raced/absent — fall back to a fresh insert
            vec = self._embed(candidate, warnings)
            return self._apply_new(
                slug, candidate, anchor_keys, link_ids, vec, now, warnings, action="created"
            )

        existing = self._store.list_memory_edges(slug, node_id=target.node_id)
        union_anchor_keys = list(
            dict.fromkeys(
                [e.target for e in existing if e.type == "ANCHORS"] + anchor_keys
            )
        )
        carried_relates = [e.target for e in existing if e.type == "RELATES"] + list(link_ids)

        # Survivor keeps the crisper (shorter) text; ties keep the target.
        survivor_content = (
            candidate.content if len(candidate.content) < len(target.content) else target.content
        )
        survivor = MemoryNode(
            slug=slug,
            content=survivor_content,
            kind=target.kind,
            labels=list(dict.fromkeys(target.labels + candidate.labels)),
            corpus=target.corpus,
            provenance=MemoryProvenance(
                author_agent=target.provenance.author_agent,
                source=target.provenance.source,
                session_id=target.provenance.session_id,
                created_at=target.provenance.created_at,
                updated_at=now,
            ),
        )
        vec = self._embed(survivor, warnings)
        self._store.upsert_memory_nodes(slug, [survivor])
        self._store_embedding(survivor, vec)

        edges = self._anchor_edges(slug, survivor.node_id, union_anchor_keys, now)
        edges += self._relate_edges(slug, survivor.node_id, carried_relates, now)
        if edges:
            self._store.upsert_memory_edges(slug, edges)

        # Content changed identity ⇒ retire the old node: invalidate its edges
        # (history preserved) and drop the now-orphaned node + embedding so it
        # neither surfaces in retrieval nor pollutes the dedup ladder.
        if survivor.node_id != target.node_id:
            self._invalidate_edges(slug, existing, now)
            self._store.delete_memory_node(slug, target.node_id)

        return IngestedClaim(
            action="merged", node_id=survivor.node_id, content=survivor.content,
            anchors=union_anchor_keys, tier=decision.tier, warnings=warnings,
        )

    def _anchor_edges(
        self, slug: str, node_id: str, keys: list[EntityKey], now: str
    ) -> list[MemoryEdge]:
        return [
            MemoryEdge(slug=slug, source=node_id, target=key, type="ANCHORS", valid_at=now)
            for key in keys
        ]

    def _relate_edges(
        self, slug: str, node_id: str, target_ids: list[str], now: str
    ) -> list[MemoryEdge]:
        seen: set[str] = set()
        out: list[MemoryEdge] = []
        for tid in target_ids:
            if tid == node_id or tid in seen:
                continue
            seen.add(tid)
            out.append(
                MemoryEdge(slug=slug, source=node_id, target=tid, type="RELATES", valid_at=now)
            )
        return out

    def _invalidate_edges(self, slug: str, edges: list[MemoryEdge], now: str) -> None:
        retired = [e.model_copy(update={"invalid_at": now}) for e in edges if e.invalid_at is None]
        if retired:
            self._store.upsert_memory_edges(slug, retired)

__init__(*, store: WikiStoreBase, embedder: EmbedderProtocol, provider: AnchorResolver, deduper: InsightDeduper, condenser: InsightCondenser | None = None, clock: Any = None, max_anchors: int = 8, max_chars: int = MAX_INSIGHT_CHARS, auto_anchor_k: int = 3) -> None

Wire collaborators (all injected); condenser is optional.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
def __init__(
    self,
    *,
    store: WikiStoreBase,
    embedder: EmbedderProtocol,
    provider: AnchorResolver,
    deduper: InsightDeduper,
    condenser: InsightCondenser | None = None,
    clock: Any = None,
    max_anchors: int = 8,
    max_chars: int = MAX_INSIGHT_CHARS,
    auto_anchor_k: int = 3,
) -> None:
    """Wire collaborators (all injected); *condenser* is optional."""
    self._store = store
    self._embedder = embedder
    self._provider = provider
    self._deduper = deduper
    self._condenser = condenser
    self._clock = clock or utc_now_iso
    self._max_anchors = max_anchors
    self._max_chars = max_chars
    self._auto_anchor_k = auto_anchor_k

from_store(store: WikiStoreBase, *, embedder: EmbedderProtocol | None = None, slug: str | None = None, llm: Any = None, condenser: InsightCondenser | None = None, clock: Any = None, provider: AnchorResolver | None = None, deduper: InsightDeduper | None = None, max_anchors: int | None = None, max_chars: int | None = None) -> InsightIngestor classmethod

Build an ingestor with the standard collaborators (DRY across surfaces).

The single construction path shared by the SessionTool, the REST endpoint, and the MCP tool, so every surface dedups/anchors identically. embedder defaults to the wiki Embedder (BM25-only _NullEmbedder when none can be built); llm/condenser are opt-in (the in-session tool leaves them off — agents pre-atomize). provider overrides the default CodeStructureProvider so an alternate corpus (e.g. the SCG connector graph) can resolve its own anchors; None keeps the code-graph default (backward-compatible).

The wiki.memory.* knobs are read HERE, at the one composition root every production caller goes through, and passed down as arguments. The failure mode that made this necessary is invisible: this class and InsightDeduper take each knob as a keyword argument whose default EQUALS the config default, so an ingestor built with none of them behaves exactly like a correctly configured one on a default deployment — an operator who retunes dedup_k or max_anchors sees no error and no effect. Reading them here keeps both classes pure DI: a class that read config itself could not be constructed in a test without one, and a directly-constructed one would silently start obeying a deployment setting the test had no way to see.

Config decides the DEFAULT, never "whether" — an explicitly passed deduper/max_anchors/max_chars still wins.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
@classmethod
def from_store(
    cls,
    store: WikiStoreBase,
    *,
    embedder: EmbedderProtocol | None = None,
    slug: str | None = None,
    llm: Any = None,
    condenser: InsightCondenser | None = None,
    clock: Any = None,
    provider: AnchorResolver | None = None,
    deduper: InsightDeduper | None = None,
    max_anchors: int | None = None,
    max_chars: int | None = None,
) -> InsightIngestor:
    """Build an ingestor with the standard collaborators (DRY across surfaces).

    The single construction path shared by the SessionTool, the REST
    endpoint, and the MCP tool, so every surface dedups/anchors
    identically. ``embedder`` defaults to the wiki ``Embedder`` (BM25-only
    ``_NullEmbedder`` when none can be built); ``llm``/``condenser`` are
    opt-in (the in-session tool leaves them off — agents pre-atomize).
    ``provider`` overrides the default ``CodeStructureProvider`` so an
    alternate corpus (e.g. the SCG connector graph) can resolve its own
    anchors; ``None`` keeps the code-graph default (backward-compatible).

    The ``wiki.memory.*`` knobs are read HERE, at the one composition root
    every production caller goes through, and passed down as arguments. The
    failure mode that made this necessary is invisible: this class and
    ``InsightDeduper`` take each knob as a keyword argument whose default
    EQUALS the config default, so an ingestor built with none of them
    behaves exactly like a correctly configured one on a default
    deployment — an operator who retunes ``dedup_k`` or ``max_anchors``
    sees no error and no effect. Reading them here keeps both classes pure
    DI: a class that read config itself could not be constructed in a test
    without one, and a directly-constructed one would silently start
    obeying a deployment setting the test had no way to see.

    Config decides the DEFAULT, never "whether" — an explicitly passed
    ``deduper``/``max_anchors``/``max_chars`` still wins.
    """
    from .structure_provider import CodeStructureProvider

    if embedder is None:
        from .embedder import make_embedder_for_or_none

        embedder = make_embedder_for_or_none(store, slug or "") or _NullEmbedder()
    if deduper is None:
        deduper = InsightDeduper(
            store=store,
            llm=llm,
            fuzzy_jaccard=float(cls._setting("fuzzy_jaccard", 0.85)),
            dedup_k=int(cls._setting("dedup_k", 5)),
            dedup_cosine=float(cls._setting("dedup_cosine", 0.6)),
        )
    return cls(
        store=store,
        embedder=embedder,
        provider=provider or CodeStructureProvider(store),
        deduper=deduper,
        condenser=condenser,
        clock=clock,
        max_anchors=(
            max_anchors if max_anchors is not None else int(cls._setting("max_anchors", 8))
        ),
        max_chars=max_chars if max_chars is not None else cls._configured_max_chars(),
    )

ingest(slug: str, content: str | None = None, *, raw: str | None = None, anchors: list[EntityKey] | None = None, links: list[str] | None = None, kind: MemoryKind = 'propositional', labels: list[str] | None = None, corpus: str = 'code', condense: bool = False, source: MemorySource = 'indexer', author_agent: str = 'insight', session_id: str | None = None) -> IngestResult

Ingest a claim (or condensed raw blob) into the memory graph.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
def ingest(
    self,
    slug: str,
    content: str | None = None,
    *,
    raw: str | None = None,
    anchors: list[EntityKey] | None = None,
    links: list[str] | None = None,
    kind: MemoryKind = "propositional",
    labels: list[str] | None = None,
    corpus: str = "code",
    condense: bool = False,
    source: MemorySource = "indexer",
    author_agent: str = "insight",
    session_id: str | None = None,
) -> IngestResult:
    """Ingest a claim (or condensed raw blob) into the memory graph."""
    base_anchors = list(anchors or [])
    links = list(links or [])
    labels = list(labels or [])
    use_raw = raw is not None or condense
    claims, condense_warnings = self._claims_for(raw if raw is not None else content, use_raw)

    outcomes: list[IngestedClaim] = []
    if not claims:
        return IngestResult(
            claims=[
                IngestedClaim(
                    action="rejected",
                    content=(raw or content or ""),
                    warnings=condense_warnings,
                )
            ]
        )
    for claim in claims:
        outcomes.append(
            self._ingest_claim(
                slug,
                claim,
                base_anchors=base_anchors,
                links=links,
                kind=kind,
                labels=labels,
                corpus=corpus,
                source=source,
                author_agent=author_agent,
                session_id=session_id,
                auto_anchor=use_raw,
                seed_warnings=condense_warnings,
            )
        )
    return IngestResult(claims=outcomes)

llm_text(llm: Any, prompt: str) -> str

Invoke a chat model and coerce its reply to plain text.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory.py
93
94
95
96
97
def llm_text(llm: Any, prompt: str) -> str:
    """Invoke a chat model and coerce its reply to plain text."""
    resp = llm.invoke(prompt)
    content = getattr(resp, "content", resp)
    return content if isinstance(content, str) else str(content)

mewbo_graph.wiki.memory_types

Pydantic models for the multiplex memory layer.

These overlay an evolving memory + docs graph on the existing tree-sitter code graph (types.py). Three node families share one identity namespace (EntityKey):

  • Code entities — the existing GraphNode/GraphEdge (untouched).
  • Memory notes — MemoryNode (ultra-small atomic claims) reified as nodes that ANCHORS to code entities and RELATES to siblings.
  • Doc pages — DocPageNote (one generated wiki page = one node) anchored to the code it documents.

Conventions match types.py: model_config = ConfigDict(extra="forbid", populate_by_name=True) and snake_case attributes. MemoryNode.node_id is derived — sha1(slug | content.strip().lower())[:16] — so two notes with the same normalized claim collapse to the same id (the exact-dup dedup tier).

DocPageNote

Bases: BaseModel

A generated wiki page as a first-class multiplex node.

Anchored to code via anchor_keys (resolved from the page's frontmatter relevantSources). The incremental refresh propagates change impact onto these to decide keep/edit/regenerate/create.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory_types.py
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
class DocPageNote(BaseModel):
    """A generated wiki page as a first-class multiplex node.

    Anchored to code via ``anchor_keys`` (resolved from the page's
    frontmatter ``relevantSources``). The incremental refresh propagates
    change impact onto these to decide keep/edit/regenerate/create.
    """

    model_config = _CFG

    slug: str
    page_id: str
    title: str
    content_hash: str
    page_type: DocPageType
    anchor_keys: list[EntityKey] = Field(default_factory=list)
    staleness_score: float = 0.0
    staleness_reason: str = "clean"
    generation_policy: DocGenerationPolicy = "keep"
    last_indexed_commit: str | None = None
    # The EVIDENCE behind the verdict above, not a second copy of it.
    #
    # ``staleness_reason`` is one of four canned strings — a CATEGORY, which is
    # all a human skimming a preview needs but not enough for anything to act
    # on. A later pass cannot recover which anchors actually moved, because the
    # run's ``ChangeSet``/``GraphDelta`` are in-memory only and discarded the
    # moment the refresh returns; the page's own verdict was the only durable
    # trace, and it names nothing.
    #
    # These two are the intersections ``_assess`` already computes and would
    # otherwise discard, so they cost one store field rather than a second pass, and
    # they are bounded by the PAGE's own anchor count rather than by repository
    # size — a page with three anchors records at most three keys however large
    # the change was. The changed FILES are deliberately not stored beside them:
    # a doc note's anchor keys ARE bare file paths — ``DocStalenessPlanner``
    # takes them verbatim from the page's frontmatter ``relevantSources`` — so
    # the file list is already here and persisting it again would be one fact in
    # two places. (The ``EntityKey`` type also spells ``path#Symbol`` elsewhere,
    # which is why this reads as a prefix-split and is not one; a
    # ``split("#")`` over these keys does nothing.)
    stale_anchor_keys: list[EntityKey] = Field(default_factory=list)
    deleted_anchor_keys: list[EntityKey] = Field(default_factory=list)

FileManifest

Bases: BaseModel

Per-(slug, path) content-hash + entity index for scoped retraction.

The incremental refresh diffs the stored content_hash against the working tree and, for dirty files, retracts exactly the listed entity_keys instead of rebuilding the whole graph.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory_types.py
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
class FileManifest(BaseModel):
    """Per-``(slug, path)`` content-hash + entity index for scoped retraction.

    The incremental refresh diffs the stored ``content_hash`` against the
    working tree and, for dirty files, retracts exactly the listed
    ``entity_keys`` instead of rebuilding the whole graph.
    """

    model_config = _CFG

    slug: str
    path: str
    content_hash: str
    last_indexed_commit: str | None = None
    entity_keys: list[EntityKey] = Field(default_factory=list)

MemoryEdge

Bases: BaseModel

Directed multiplex edge.

ANCHORS: memory node_id → code EntityKey (the de-facto hyperedge fan-out). RELATES: memory node_id → memory node_id. Validity is a single nullable axis — invalid_at=None means live; setting it invalidates the edge (Graphiti invalidate-don't-delete).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory_types.py
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
class MemoryEdge(BaseModel):
    """Directed multiplex edge.

    ``ANCHORS``: memory ``node_id`` → code ``EntityKey`` (the de-facto
    hyperedge fan-out). ``RELATES``: memory ``node_id`` → memory ``node_id``.
    Validity is a single nullable axis — ``invalid_at=None`` means live;
    setting it invalidates the edge (Graphiti invalidate-don't-delete).
    """

    model_config = _CFG

    slug: str
    source: str
    target: str
    type: MemoryEdgeType
    weight: float = 1.0
    valid_at: str
    invalid_at: str | None = None

MemoryEmbedding

Bases: BaseModel

Dense embedding vector for a memory node (mirrors Embedding).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory_types.py
119
120
121
122
123
124
125
126
127
128
class MemoryEmbedding(BaseModel):
    """Dense embedding vector for a memory node (mirrors ``Embedding``)."""

    model_config = _CFG

    slug: str
    node_id: str
    vector: list[float]
    model: str
    dim: int

MemoryFilter

Bases: BaseModel

Optional facets applied to memory retrieval (all default to no-op).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory_types.py
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
class MemoryFilter(BaseModel):
    """Optional facets applied to memory retrieval (all default to no-op)."""

    model_config = _CFG

    corpus: str | None = None
    source: MemorySource | None = None
    kind: MemoryKind | None = None
    labels: list[str] | None = None
    valid_at: str | None = None
    exclude_invalidated: bool = True

    def matches_node(self, node: MemoryNode) -> bool:
        """Return True if *node* satisfies the node-level facets.

        Edge-level validity (``valid_at`` / ``exclude_invalidated``) is
        applied separately by the store/retriever against the edge set.
        """
        if self.corpus is not None and node.corpus != self.corpus:
            return False
        if self.source is not None and node.provenance.source != self.source:
            return False
        if self.kind is not None and node.kind != self.kind:
            return False
        if self.labels and not set(self.labels).issubset(set(node.labels)):
            return False
        return True

matches_node(node: MemoryNode) -> bool

Return True if node satisfies the node-level facets.

Edge-level validity (valid_at / exclude_invalidated) is applied separately by the store/retriever against the edge set.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory_types.py
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
def matches_node(self, node: MemoryNode) -> bool:
    """Return True if *node* satisfies the node-level facets.

    Edge-level validity (``valid_at`` / ``exclude_invalidated``) is
    applied separately by the store/retriever against the edge set.
    """
    if self.corpus is not None and node.corpus != self.corpus:
        return False
    if self.source is not None and node.provenance.source != self.source:
        return False
    if self.kind is not None and node.kind != self.kind:
        return False
    if self.labels and not set(self.labels).issubset(set(node.labels)):
        return False
    return True

MemoryNode

Bases: BaseModel

An atomic memory claim reified as a multiplex graph node.

node_id is always derived from (slug, content) — any supplied value is overwritten. This makes identical normalized claims share one id, which is exactly the exact-match tier of the dedup ladder.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory_types.py
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
class MemoryNode(BaseModel):
    """An atomic memory claim reified as a multiplex graph node.

    ``node_id`` is always derived from ``(slug, content)`` — any supplied
    value is overwritten. This makes identical normalized claims share one
    id, which is exactly the exact-match tier of the dedup ladder.
    """

    model_config = _CFG

    slug: str
    node_id: str = ""
    content: str = Field(max_length=MAX_INSIGHT_CHARS)
    kind: MemoryKind = "propositional"
    labels: list[str] = Field(default_factory=list)
    corpus: str = "code"
    provenance: MemoryProvenance
    # ISO timestamp of the last incremental-refresh anchor check (idempotency
    # guard in ``MemoryReconciler``); ``None`` until first reconciled.
    anchor_checked_at: str | None = None

    @staticmethod
    def compute_node_id(slug: str, content: str) -> str:
        """Deterministic id over ``(slug, normalized content)``."""
        h = hashlib.sha1(f"{slug}|{content.strip().lower()}".encode())
        return h.hexdigest()[:16]

    @model_validator(mode="after")
    def _derive_node_id(self) -> MemoryNode:
        """Force ``node_id`` to the derived value (single source of truth)."""
        derived = self.compute_node_id(self.slug, self.content)
        if self.node_id != derived:
            self.node_id = derived
        return self

compute_node_id(slug: str, content: str) -> str staticmethod

Deterministic id over (slug, normalized content).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory_types.py
84
85
86
87
88
@staticmethod
def compute_node_id(slug: str, content: str) -> str:
    """Deterministic id over ``(slug, normalized content)``."""
    h = hashlib.sha1(f"{slug}|{content.strip().lower()}".encode())
    return h.hexdigest()[:16]

MemoryProvenance

Bases: BaseModel

Who/when/how a memory note was created — citable audit trail.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/memory_types.py
48
49
50
51
52
53
54
55
56
57
class MemoryProvenance(BaseModel):
    """Who/when/how a memory note was created — citable audit trail."""

    model_config = _CFG

    author_agent: str
    source: MemorySource
    session_id: str | None = None
    created_at: str
    updated_at: str | None = None

mewbo_graph.wiki.store

Wiki persistence layer.

JSON-file backed implementation (default) + abstract base for the MongoDB impl that lands in Task 1.4. Layout under $MEWBO_HOME/wiki/:

projects/<slug>.json                    (Project model — DISPLAY snapshot)
settings/<slug>.json                    (ProjectSettings — the editable record)
pages/<slug>/_index.json                (page-id→title index for fast listing)
pages/<slug>/<page_id>.json             (full WikiPage including body)
jobs/<job_id>/job.json                  (IndexingJob model)
jobs/<job_id>/events.jsonl              (append-only event log with idx)
jobs/<job_id>/session.txt               (Mewbo session_id — one line)
qa/<answer_id>/answer.json              (QaAnswer model)
qa/<answer_id>/events.jsonl             (append-only event log with idx)

Slugs that contain slashes (e.g. "org/repo") are escaped as "org__repo" so they map safely to a single directory/filename segment.

JobPatch dataclass

The IndexingJob fields a caller NAMED, validated and nothing else.

The unit both drivers write, and what keeps two overlapping writers from losing each other's changes: "set these fields" and "rewrite the document that happens to hold them" are not the same operation. The second reverts every field a CONCURRENT writer changed between this writer's read and its write — a cancel landing between a progress writer's read and its write was silently undone, and the job carried on running with no record that a cancel had ever been asked for. Carrying only the named fields lets each backend narrow its write to what the caller actually asked for, so writers touching disjoint fields stop colliding at all.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
@dataclass(frozen=True)
class JobPatch:
    """The ``IndexingJob`` fields a caller NAMED, validated and nothing else.

    The unit both drivers write, and what keeps two overlapping writers from
    losing each other's changes: "set these fields" and "rewrite the document
    that happens to hold them" are not the same operation. The second reverts
    every field a CONCURRENT writer changed between this writer's read and its
    write — a cancel landing between a progress writer's read and its write was
    silently undone, and the job carried on running with no record that a cancel
    had ever been asked for. Carrying only the named fields lets each backend
    narrow its write to what the caller actually asked for, so writers touching
    disjoint fields stop colliding at all.
    """

    # ``progress`` follows the same named-field rule: its write is disjoint from
    # ``cancel_job`` writing ``status``, so neither can revert the other. Fresh
    # fan-out reporters can race on the ledger field itself and lose one step
    # update; that bounded loss heals on the next report and is acceptable beside
    # the unacceptable alternative of reverting an unrelated field.
    fields: dict[str, Any]

    @classmethod
    def build(cls, job: IndexingJob, fields: Mapping[str, Any]) -> JobPatch:
        """Validate *fields* against the WHOLE *job*, then keep only those keys.

        Validation stays whole-document — an unknown key still fails
        ``extra="forbid"`` and every value is still coerced by the field that
        owns it — because what needed narrowing is the WRITE, not the check.

        An explicit ``None`` is a VALUE here, never an omission: ``emit_phase``
        clears the three ``phase_progress_*`` fields by naming them, so dropping
        falsy values (as the ``PROJECT_UPDATABLE`` filter does, for a surface
        whose ``None`` genuinely means "not supplied") would silently discard
        the write the progress invariant depends on.
        """
        merged = job.model_dump(by_alias=False)
        merged.update(fields)
        coerced = IndexingJob.model_validate(merged).model_dump(by_alias=False)
        return cls(fields={name: coerced[name] for name in fields})

    def apply(self, job: IndexingJob) -> IndexingJob:
        """Return *job* with this patch's fields set, re-validated as a whole."""
        merged = job.model_dump(by_alias=False)
        merged.update(self.fields)
        return IndexingJob.model_validate(merged)

apply(job: IndexingJob) -> IndexingJob

Return job with this patch's fields set, re-validated as a whole.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
147
148
149
150
151
def apply(self, job: IndexingJob) -> IndexingJob:
    """Return *job* with this patch's fields set, re-validated as a whole."""
    merged = job.model_dump(by_alias=False)
    merged.update(self.fields)
    return IndexingJob.model_validate(merged)

build(job: IndexingJob, fields: Mapping[str, Any]) -> JobPatch classmethod

Validate fields against the WHOLE job, then keep only those keys.

Validation stays whole-document — an unknown key still fails extra="forbid" and every value is still coerced by the field that owns it — because what needed narrowing is the WRITE, not the check.

An explicit None is a VALUE here, never an omission: emit_phase clears the three phase_progress_* fields by naming them, so dropping falsy values (as the PROJECT_UPDATABLE filter does, for a surface whose None genuinely means "not supplied") would silently discard the write the progress invariant depends on.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
@classmethod
def build(cls, job: IndexingJob, fields: Mapping[str, Any]) -> JobPatch:
    """Validate *fields* against the WHOLE *job*, then keep only those keys.

    Validation stays whole-document — an unknown key still fails
    ``extra="forbid"`` and every value is still coerced by the field that
    owns it — because what needed narrowing is the WRITE, not the check.

    An explicit ``None`` is a VALUE here, never an omission: ``emit_phase``
    clears the three ``phase_progress_*`` fields by naming them, so dropping
    falsy values (as the ``PROJECT_UPDATABLE`` filter does, for a surface
    whose ``None`` genuinely means "not supplied") would silently discard
    the write the progress invariant depends on.
    """
    merged = job.model_dump(by_alias=False)
    merged.update(fields)
    coerced = IndexingJob.model_validate(merged).model_dump(by_alias=False)
    return cls(fields={name: coerced[name] for name in fields})

JsonWikiStore

Bases: WikiStoreBase

File-backed implementation under $MEWBO_HOME/wiki/ (or a custom root).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
2028
2029
2030
2031
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
2100
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
2115
2116
2117
2118
2119
2120
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
2143
2144
2145
2146
2147
2148
2149
2150
2151
2152
2153
2154
2155
2156
2157
2158
2159
2160
2161
2162
2163
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
2178
2179
2180
2181
2182
2183
2184
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
2204
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
2226
2227
2228
2229
2230
2231
2232
2233
2234
2235
2236
2237
2238
2239
2240
2241
2242
2243
2244
2245
2246
2247
2248
2249
2250
2251
2252
2253
2254
2255
2256
2257
2258
2259
2260
2261
2262
2263
2264
2265
2266
2267
2268
2269
2270
2271
2272
2273
2274
2275
2276
2277
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
2291
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
2308
2309
2310
2311
2312
2313
2314
2315
2316
2317
2318
2319
2320
2321
2322
2323
2324
2325
2326
2327
2328
2329
2330
2331
2332
2333
2334
2335
2336
2337
2338
2339
2340
2341
2342
2343
2344
2345
2346
2347
2348
2349
2350
2351
2352
2353
2354
2355
2356
2357
2358
2359
2360
2361
2362
2363
2364
2365
2366
2367
2368
2369
2370
2371
2372
2373
2374
2375
2376
2377
2378
2379
2380
2381
2382
2383
2384
2385
2386
2387
2388
2389
2390
2391
2392
2393
2394
2395
2396
2397
2398
2399
2400
2401
2402
2403
2404
2405
2406
2407
2408
2409
2410
2411
2412
2413
2414
2415
2416
2417
2418
2419
2420
2421
2422
2423
2424
2425
2426
2427
2428
2429
2430
2431
2432
2433
2434
2435
2436
2437
2438
2439
2440
2441
2442
2443
2444
2445
2446
2447
2448
2449
2450
2451
2452
2453
2454
2455
2456
2457
2458
2459
2460
2461
2462
2463
2464
2465
2466
2467
2468
2469
2470
2471
2472
2473
2474
2475
2476
2477
2478
2479
2480
2481
2482
2483
2484
2485
2486
2487
2488
2489
2490
2491
2492
2493
2494
2495
2496
2497
2498
2499
2500
2501
2502
2503
2504
2505
2506
2507
2508
2509
2510
2511
2512
2513
2514
2515
2516
2517
2518
2519
2520
2521
2522
2523
2524
2525
2526
2527
2528
2529
2530
2531
2532
2533
2534
2535
2536
2537
2538
2539
2540
2541
2542
2543
2544
2545
2546
2547
2548
2549
2550
2551
2552
2553
2554
2555
2556
2557
2558
2559
2560
2561
2562
2563
2564
2565
2566
2567
2568
2569
2570
2571
2572
2573
2574
2575
2576
2577
2578
2579
2580
2581
2582
2583
2584
2585
2586
2587
2588
2589
2590
2591
2592
2593
2594
2595
2596
2597
2598
2599
2600
2601
2602
2603
2604
2605
2606
2607
2608
2609
2610
2611
2612
2613
2614
2615
2616
2617
2618
2619
2620
2621
2622
2623
2624
2625
2626
2627
2628
2629
2630
2631
2632
2633
2634
2635
2636
2637
2638
2639
2640
2641
2642
2643
2644
2645
2646
2647
2648
2649
2650
2651
2652
2653
2654
2655
2656
2657
2658
2659
2660
2661
2662
class JsonWikiStore(WikiStoreBase):
    """File-backed implementation under ``$MEWBO_HOME/wiki/`` (or a custom root)."""

    def __init__(self, root_dir: str | Path | None = None) -> None:
        """Initialise and create the directory tree."""
        if root_dir is None:
            home = get_config_value("runtime", "cache_dir", default="") or ".mewbo"
            root_dir = Path(home) / "wiki"
        self.root_dir = Path(root_dir)
        self.root_dir.mkdir(parents=True, exist_ok=True)
        for sub in ("projects", "pages", "jobs", "qa"):
            (self.root_dir / sub).mkdir(parents=True, exist_ok=True)
        self._lock = threading.Lock()

    # -- Private helpers -----------------------------------------------------

    def _save_json(self, path: Path, model: Any) -> None:
        """Persist a Pydantic model as JSON (by_alias, mode=json)."""
        path.parent.mkdir(parents=True, exist_ok=True)
        path.write_text(
            model.model_dump_json(by_alias=True, indent=2), encoding="utf-8"
        )

    def _load_json(self, path: Path, model_cls: type[_M]) -> _M | None:
        """Load a Pydantic model from JSON, returning None if missing."""
        if not path.exists():
            return None
        try:
            return model_cls.model_validate_json(path.read_text(encoding="utf-8"))
        except Exception:
            logging.warning("Skipping malformed JSON at {}", path)
            return None

    def _event_path(self, scope: str, owner_id: str) -> Path:
        """Return path to the JSONL event log for jobs or qa."""
        return self.root_dir / scope / owner_id / "events.jsonl"

    def _append_event(
        self, scope: str, owner_id: str, event_dict: dict[str, Any]
    ) -> int:
        """Append event_dict to the JSONL log; return its monotonic idx."""
        with self._lock:
            path = self._event_path(scope, owner_id)
            path.parent.mkdir(parents=True, exist_ok=True)
            idx = 0
            if path.exists():
                lines = [ln for ln in path.read_text(encoding="utf-8").splitlines() if ln.strip()]
                idx = len(lines)
            payload = {**event_dict, "idx": idx}
            with path.open("a", encoding="utf-8") as fh:
                fh.write(json.dumps(payload) + "\n")
            return idx

    def _load_events(
        self, scope: str, owner_id: str, after_idx: int = -1
    ) -> list[dict[str, Any]]:
        """Load events from JSONL; filter to idx > after_idx."""
        path = self._event_path(scope, owner_id)
        if not path.exists():
            return []
        results: list[dict[str, Any]] = []
        for line in path.read_text(encoding="utf-8").splitlines():
            line = line.strip()
            if not line:
                continue
            try:
                rec: dict[str, Any] = json.loads(line)
            except json.JSONDecodeError:
                logging.warning("Skipping malformed event line in {}", path)
                continue
            if rec.get("idx", -1) > after_idx:
                results.append(rec)
        return results

    # -- Projects ------------------------------------------------------------

    def _project_path(self, slug: str) -> Path:
        """Filesystem path for a project JSON file."""
        return self.root_dir / "projects" / f"{_slug_to_path(slug)}.json"

    def create_project(self, project: Project) -> None:
        """Persist a new project record."""
        self._save_json(self._project_path(project.slug), project)

    def get_project(self, slug: str) -> Project | None:
        """Return the project for *slug*, or None if absent."""
        return self._load_json(self._project_path(slug), Project)

    def list_projects(self) -> list[Project]:
        """Return all projects sorted by indexed_at descending."""
        projects: list[Project] = []
        for p in (self.root_dir / "projects").glob("*.json"):
            proj = self._load_json(p, Project)
            if proj is not None:
                projects.append(proj)
        return sorted(projects, key=lambda pr: pr.indexed_at, reverse=True)

    def delete_project(self, slug: str) -> bool:
        """Delete project *slug*; return True if deleted, False if absent."""
        path = self._project_path(slug)
        if not path.exists():
            return False
        path.unlink()
        return True

    # -- Project settings (slug-keyed sidecar) -------------------------------

    def _settings_path(self, slug: str) -> Path:
        """Filesystem path for a slug's editable settings record."""
        d = self.root_dir / "settings"
        d.mkdir(parents=True, exist_ok=True)
        return d / f"{_slug_to_path(slug)}.json"

    def save_project_settings(self, slug: str, settings: ProjectSettings) -> None:
        """Persist (upsert) the editable settings record for *slug*."""
        self._save_json(self._settings_path(slug), settings)

    def get_project_settings(self, slug: str) -> ProjectSettings | None:
        """Return *slug*'s settings record, or None when never written."""
        return self._load_json(self._settings_path(slug), ProjectSettings)

    def delete_project_settings(self, slug: str) -> bool:
        """Delete *slug*'s settings file; return True if one existed."""
        path = self._settings_path(slug)
        if not path.exists():
            return False
        path.unlink()
        return True

    # -- Whole-slug reap (see WikiStoreBase.reap_slug for the family list) ---

    def reap_slug(self, slug: str) -> dict[str, int]:
        """See ``WikiStoreBase.reap_slug``.

        Counts are taken from each JSONL/file BEFORE its owning directory is
        removed, so a family's number here means the same thing as the Mongo
        driver's ``delete_many().deleted_count`` — one count per row, not "the
        directory existed". Graph/entity/memory each live under ONE directory
        (``_graph_dir``/``_memory_dir``), and pages under another
        (``_pages_dir``), so counting + removing those collapses into directory
        operations rather than per-file bookkeeping. Jobs and QA are id-keyed,
        not slug-keyed: their ids are gathered FIRST, before anything is
        deleted, then each owning directory is removed by id.
        """
        with self._lock:
            job_ids = [j.job_id for j in self.list_jobs(slug)]
            answer_ids = self._qa_answer_ids_for_slug(slug)

            graph_dir = self._graph_dir(slug)
            counts: dict[str, int] = {
                "graph_nodes": self._count_jsonl_lines(graph_dir / "nodes.jsonl"),
                "graph_edges": self._count_jsonl_lines(graph_dir / "edges.jsonl"),
                "embeddings": self._count_jsonl_lines(graph_dir / "embeddings.jsonl"),
            }
            self._rmdir_if_exists(graph_dir)

            mem_dir = self._memory_dir(slug)
            counts.update({
                "entities": self._count_jsonl_lines(mem_dir / "entities.jsonl"),
                "entity_edges": self._count_jsonl_lines(mem_dir / "entity_edges.jsonl"),
                "entity_embeddings": self._count_jsonl_lines(
                    mem_dir / "entity_embeddings.jsonl"
                ),
                "entity_recommendations": self._count_jsonl_lines(
                    mem_dir / "entity_recommendations.jsonl"
                ),
                "memory_nodes": self._count_jsonl_lines(mem_dir / "nodes.jsonl"),
                "memory_edges": self._count_jsonl_lines(mem_dir / "edges.jsonl"),
                "memory_embeddings": self._count_jsonl_lines(mem_dir / "embeddings.jsonl"),
                "doc_notes": self._count_jsonl_lines(mem_dir / "docs.jsonl"),
                "file_manifest": self._count_jsonl_lines(mem_dir / "manifest.jsonl"),
            })
            self._rmdir_if_exists(mem_dir)

            pages_dir = self._pages_dir(slug)
            counts["pages"] = (
                sum(
                    1
                    for p in pages_dir.glob("*.json")
                    if p.name not in ("_index.json", "_attribution.json")
                )
                if pages_dir.exists()
                else 0
            )
            self._rmdir_if_exists(pages_dir)

            recovery_path = self._recovery_path(slug)
            counts["recovery"] = 1 if recovery_path.exists() else 0
            if counts["recovery"]:
                recovery_path.unlink()

            counts["jobs"] = len(job_ids)
            counts["job_events"] = sum(
                self._count_jsonl_lines(self._event_path("jobs", jid)) for jid in job_ids
            )
            for job_id in job_ids:
                self._rmdir_if_exists(self._job_dir(job_id))

            counts["qa"] = len(answer_ids)
            counts["qa_events"] = sum(
                self._count_jsonl_lines(self._event_path("qa", aid)) for aid in answer_ids
            )
            for answer_id in answer_ids:
                self._rmdir_if_exists(self._qa_dir(answer_id))

        return counts

    def _qa_answer_ids_for_slug(self, slug: str) -> list[str]:
        """Scan ``qa/<answer_id>/answer.json`` for the ones belonging to *slug*.

        QA answers are id-keyed, not slug-keyed, so there is no single directory
        to remove wholesale the way graph/memory/pages allow — every answer_id
        for this slug has to be found via its own record first.
        """
        qa_root = self.root_dir / "qa"
        if not qa_root.exists():
            return []
        ids: list[str] = []
        for qa_dir in qa_root.iterdir():
            if not qa_dir.is_dir():
                continue
            answer = self._load_json(qa_dir / "answer.json", QaAnswer)
            if answer is not None and answer.slug == slug:
                ids.append(qa_dir.name)
        return ids

    @staticmethod
    def _count_jsonl_lines(path: Path) -> int:
        """Count non-empty lines in a JSONL file; 0 if the file is absent."""
        if not path.exists():
            return 0
        return sum(1 for ln in path.read_text(encoding="utf-8").splitlines() if ln.strip())

    @staticmethod
    def _rmdir_if_exists(path: Path) -> None:
        """Remove a directory tree if present; no-op if it was already gone."""
        if path.exists():
            shutil.rmtree(path)

    # -- Pages ---------------------------------------------------------------

    def _pages_dir(self, slug: str) -> Path:
        """Directory containing all pages for *slug*."""
        return self.root_dir / "pages" / _slug_to_path(slug)

    def _page_path(self, slug: str, page_id: str) -> Path:
        """Filesystem path for a page JSON file."""
        return self._pages_dir(slug) / f"{_slug_to_path(page_id)}.json"

    def _index_path(self, slug: str) -> Path:
        """Filesystem path for the page-id→title index."""
        return self._pages_dir(slug) / "_index.json"

    def _attribution_path(self, slug: str) -> Path:
        """Filesystem path for the page-id→{commit_sha,job_id} attribution sidecar.

        Page attribution rides a sidecar rather than the page JSON so the
        persisted ``WikiPage`` (an ``extra="forbid"`` console wire type) stays
        byte-identical — the same reason the Mongo driver keeps it a store column.
        """
        return self._pages_dir(slug) / "_attribution.json"

    def _load_index(self, slug: str) -> dict[str, str]:
        """Load the page-id→title index; returns {} if absent."""
        idx_path = self._index_path(slug)
        if not idx_path.exists():
            return {}
        try:
            return json.loads(idx_path.read_text(encoding="utf-8"))
        except Exception:
            return {}

    def _load_attribution(self, slug: str) -> dict[str, dict[str, Any]]:
        """Load the page attribution sidecar; returns {} if absent/unreadable."""
        path = self._attribution_path(slug)
        if not path.exists():
            return {}
        try:
            return json.loads(path.read_text(encoding="utf-8"))
        except Exception:
            return {}

    def save_page(
        self,
        slug: str,
        page: WikiPage,
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Persist *page* for the project *slug*; overwrites if same page_id."""
        pages_dir = self._pages_dir(slug)
        pages_dir.mkdir(parents=True, exist_ok=True)
        self._save_json(self._page_path(slug, page.id), page)
        index = self._load_index(slug)
        index[page.id] = page.title
        self._index_path(slug).write_text(json.dumps(index, indent=2), encoding="utf-8")
        attribution = self._load_attribution(slug)
        attribution[page.id] = {"commit_sha": commit_sha, "job_id": job_id}
        self._attribution_path(slug).write_text(
            json.dumps(attribution, indent=2), encoding="utf-8"
        )

    def _get_page_raw(self, slug: str, page_id: str) -> WikiPage | None:
        """Return a single wiki page, or None if absent (no doc-guard)."""
        return self._load_json(self._page_path(slug, page_id), WikiPage)

    def list_pages(self, slug: str) -> list[WikiPage]:
        """Return all pages for project *slug*."""
        pages_dir = self._pages_dir(slug)
        if not pages_dir.exists():
            return []
        pages: list[WikiPage] = []
        for p in pages_dir.glob("*.json"):
            if p.name in ("_index.json", "_attribution.json"):
                continue
            page = self._load_json(p, WikiPage)
            if page is not None:
                pages.append(page)
        return pages

    def delete_page(self, slug: str, page_id: str) -> bool:
        """Delete a single page on disk + drop it from the index."""
        path = self._page_path(slug, page_id)
        removed = path.exists()
        if removed:
            path.unlink()
        index = self._load_index(slug)
        if index.pop(page_id, None) is not None:
            self._index_path(slug).write_text(
                json.dumps(index, indent=2), encoding="utf-8"
            )
            removed = True
        return removed

    # -- Indexing jobs -------------------------------------------------------

    def _job_dir(self, job_id: str) -> Path:
        """Directory for a job's artefacts."""
        return self.root_dir / "jobs" / job_id

    def _job_path(self, job_id: str) -> Path:
        """Filesystem path for a job JSON file."""
        return self._job_dir(job_id) / "job.json"

    def _session_path(self, job_id: str) -> Path:
        """Filesystem path for the session-id text file."""
        return self._job_dir(job_id) / "session.txt"

    def create_job(self, job: IndexingJob) -> None:
        """Persist a new indexing job."""
        self._job_dir(job.job_id).mkdir(parents=True, exist_ok=True)
        self._save_json(self._job_path(job.job_id), job)

    def get_job(self, job_id: str) -> IndexingJob | None:
        """Return the indexing job, or None if absent."""
        return self._load_json(self._job_path(job_id), IndexingJob)

    def _write_job_patch(self, job_id: str, patch: JobPatch) -> IndexingJob:
        """Re-read, apply and save the job file under the store lock.

        The whole file is the unit of write here, so a field-scoped write is
        only real if the read the merge is built on and the write that replaces
        it cannot be interleaved — the lock has to span BOTH. That is why this
        re-reads rather than applying the patch to the snapshot the caller
        already validated against: the caller's read happened outside the lock
        and may already be stale.

        Every other read-modify-write in this class takes this lock; this path
        was the omission, not the convention.
        """
        with self._lock:
            job = self._load_json(self._job_path(job_id), IndexingJob)
            if job is None:
                raise KeyError(f"Job not found: {job_id}")
            updated = patch.apply(job)
            self._save_json(self._job_path(job_id), updated)
        return updated

    def list_jobs(self, slug: str | None = None) -> list[IndexingJob]:
        """Return all jobs, newest first by ``phase_started_at``, filtered to *slug*.

        ``iterdir()`` yields entries in arbitrary, platform-dependent filesystem
        order, so the list MUST be sorted before return, and NOT by ``job_id`` —
        a ``uuid4`` hex sorts RANDOMLY with respect to when a job actually ran,
        which is not merely unhelpful but ACTIVELY MISLEADING: at least one
        caller (``resolve_qa_clone_dir``) assumes this method returns
        most-recent-first and picks the FIRST ``complete`` hit, which under such
        a sort could be an arbitrary old checkout on any slug with 2+ completed
        jobs. ``phase_started_at`` is ISO-8601, so lexicographic ==
        chronological; a job that never emitted a phase (no timestamp) sorts
        last, never winning over one that has. :meth:`latest_job` is a thin
        convenience over this order (narrow by status, take the first).
        """
        jobs_root = self.root_dir / "jobs"
        if not jobs_root.exists():
            return []
        jobs: list[IndexingJob] = []
        for job_dir in jobs_root.iterdir():
            if not job_dir.is_dir():
                continue
            job = self._load_json(job_dir / "job.json", IndexingJob)
            if job is None:
                continue
            if slug is None or job.slug == slug:
                jobs.append(job)
        return sorted(jobs, key=lambda j: j.phase_started_at or "", reverse=True)

    def _append_job_event(self, job_id: str, event: dict[str, Any]) -> int:
        """Append one validated event to this JSON job timeline. Cost: ``O(1)``."""
        return self._append_event("jobs", job_id, event)

    def load_job_events(
        self, job_id: str, after_idx: int = -1
    ) -> list[dict[str, Any]]:
        """Return job events with idx > *after_idx* (-1 returns all)."""
        return self._load_events("jobs", job_id, after_idx)

    def cancel_job(self, job_id: str) -> bool:
        """Cancel *job_id*; return True on first cancel, False if already cancelled."""
        job = self.get_job(job_id)
        if job is None:
            return False
        if job.status == "cancelled":
            return False
        self.update_job(job_id, status="cancelled")
        self.append_job_event(job_id, {"type": "cancelled"})
        return True

    def attach_job_session(self, job_id: str, session_id: str) -> None:
        """Associate a Mewbo session_id with an indexing job."""
        path = self._session_path(job_id)
        path.parent.mkdir(parents=True, exist_ok=True)
        path.write_text(session_id, encoding="utf-8")

    def get_job_session(self, job_id: str) -> str | None:
        """Return the session_id attached to *job_id*, or None."""
        path = self._session_path(job_id)
        if not path.exists():
            return None
        return path.read_text(encoding="utf-8").strip() or None

    def find_job_by_session(self, session_id: str) -> str | None:
        """Reverse lookup: scan job dirs for the session.txt that matches *session_id*."""
        jobs_root = self.root_dir / "jobs"
        if not jobs_root.exists():
            return None
        for job_dir in jobs_root.iterdir():
            if not job_dir.is_dir():
                continue
            sess_file = job_dir / "session.txt"
            if sess_file.exists() and sess_file.read_text(encoding="utf-8").strip() == session_id:
                return job_dir.name
        return None

    def _job_plan_path(self, job_id: str) -> Path:
        """Filesystem path for the page-plan sidecar file."""
        return self._job_dir(job_id) / "plan.json"

    def _job_meta_path(self, job_id: str) -> Path:
        """Filesystem path for the job extra-metadata sidecar file."""
        return self._job_dir(job_id) / "meta.json"

    def _load_job_meta(self, job_id: str) -> dict[str, Any]:
        """Load job metadata sidecar; returns {} if absent."""
        path = self._job_meta_path(job_id)
        if not path.exists():
            return {}
        try:
            return json.loads(path.read_text(encoding="utf-8"))
        except Exception:
            return {}

    def save_job_plan(self, job_id: str, plan: list[dict[str, Any]]) -> None:
        """Persist the page-plan list for *job_id*; overwrites any previous plan."""
        path = self._job_plan_path(job_id)
        path.parent.mkdir(parents=True, exist_ok=True)
        path.write_text(json.dumps(plan, indent=2), encoding="utf-8")

    def get_job_plan(self, job_id: str) -> list[dict[str, Any]] | None:
        """Return the page-plan list, or None if no plan has been committed yet."""
        path = self._job_plan_path(job_id)
        if not path.exists():
            return None
        try:
            data = json.loads(path.read_text(encoding="utf-8"))
            return data if isinstance(data, list) else None
        except Exception:
            return None

    def _job_resume_path(self, job_id: str) -> Path:
        """Filesystem path for the resume-plan sidecar file."""
        return self._job_dir(job_id) / "resume.json"

    def save_resume_plan(self, job_id: str, plan: dict[str, Any]) -> None:
        """Persist the resume-plan dict for *job_id*; overwrites any previous one."""
        path = self._job_resume_path(job_id)
        path.parent.mkdir(parents=True, exist_ok=True)
        path.write_text(json.dumps(plan, indent=2), encoding="utf-8")

    def get_resume_plan(self, job_id: str) -> dict[str, Any] | None:
        """Return the persisted resume-plan dict, or None if the job isn't resuming."""
        path = self._job_resume_path(job_id)
        if not path.exists():
            return None
        try:
            data = json.loads(path.read_text(encoding="utf-8"))
            return data if isinstance(data, dict) else None
        except Exception:
            return None

    def _job_act_path(self, job_id: str) -> Path:
        """Filesystem path for the act-stage sidecar file."""
        return self._job_dir(job_id) / "act.json"

    def save_act_plan(self, job_id: str, plan: dict[str, Any]) -> None:
        """Persist the act-stage record for *job_id*; overwrites any previous one."""
        path = self._job_act_path(job_id)
        path.parent.mkdir(parents=True, exist_ok=True)
        path.write_text(json.dumps(plan, indent=2), encoding="utf-8")

    def get_act_plan(self, job_id: str) -> dict[str, Any] | None:
        """Return the persisted act-stage record, or None if stage 1 hasn't finished."""
        path = self._job_act_path(job_id)
        if not path.exists():
            return None
        try:
            data = json.loads(path.read_text(encoding="utf-8"))
            return data if isinstance(data, dict) else None
        except Exception:
            return None

    def get_job_submitted_count(self, job_id: str) -> int:
        """Return the number of pages submitted so far for *job_id*."""
        meta = self._load_job_meta(job_id)
        return int(meta.get("submitted_pages", 0))

    def claim_job_page(self, slug: str, job_id: str, page_id: str) -> PageClaim:
        """Record *page_id* under *job_id*, under the lock; count = set size."""
        if self.get_job(job_id) is None:
            raise KeyError(f"Job not found: {job_id}")
        with self._lock:
            meta = self._load_job_meta(job_id)
            claimed = self._claimed_ids(meta, slug, job_id)
            if page_id in claimed:
                return PageClaim(count=len(claimed), is_new=False)
            claimed.append(page_id)
            meta["submitted_page_ids"] = claimed
            # Kept in step purely so ``get_job_submitted_count`` stays a cheap
            # read; the SET is the authority, so the two can never disagree.
            meta["submitted_pages"] = len(claimed)
            path = self._job_meta_path(job_id)
            path.parent.mkdir(parents=True, exist_ok=True)
            path.write_text(json.dumps(meta, indent=2), encoding="utf-8")
            return PageClaim(count=len(claimed), is_new=True)

    def _claimed_ids(self, meta: dict[str, Any], slug: str, job_id: str) -> list[str]:
        """The job's claimed ids, seeded from attribution when it has no record.

        A MISSING key and an empty list are different answers: the first is a
        job with no claim record (recover what it wrote from attribution), the
        second is a job that has genuinely written nothing yet (and must not
        adopt some other index's pages).
        """
        stored = meta.get("submitted_page_ids")
        if stored is None:
            return sorted(self.page_ids_for_job(slug, job_id))
        return [str(p) for p in stored]

    def get_job_page_ids(self, slug: str, job_id: str) -> frozenset[str]:
        """Return the page ids *job_id* wrote (claim record, else attribution)."""
        return frozenset(self._claimed_ids(self._load_job_meta(job_id), slug, job_id))

    def page_ids_for_job(self, slug: str, job_id: str) -> frozenset[str]:
        """Page ids whose attribution sidecar entry names *job_id*."""
        return frozenset(
            pid
            for pid, meta in self._load_attribution(slug).items()
            if isinstance(meta, dict) and meta.get("job_id") == job_id
        )

    def _job_submission_path(self, job_id: str) -> Path:
        """Filesystem path for the submission sidecar file."""
        return self._job_dir(job_id) / "submission.json"

    def save_job_submission(self, job_id: str, submission: dict[str, Any]) -> None:
        """Persist the wizard submission dict for *job_id* (token must be absent)."""
        path = self._job_submission_path(job_id)
        path.parent.mkdir(parents=True, exist_ok=True)
        path.write_text(json.dumps(submission, indent=2), encoding="utf-8")

    def get_job_submission(self, job_id: str) -> dict[str, Any] | None:
        """Return the persisted submission dict, or None if not yet saved."""
        path = self._job_submission_path(job_id)
        if not path.exists():
            return None
        try:
            data = json.loads(path.read_text(encoding="utf-8"))
            return data if isinstance(data, dict) else None
        except Exception:
            return None

    # -- Repository credentials (isolated subdir, mode 0600) -----------------

    def _credentials_dir(self) -> Path:
        """Directory holding per-slug credential files (mode 0700)."""
        d = self.root_dir / "credentials"
        d.mkdir(parents=True, exist_ok=True)
        try:
            d.chmod(0o700)
        except OSError:  # pragma: no cover — best-effort on exotic filesystems
            pass
        return d

    def _credential_path(self, slug: str) -> Path:
        """Filesystem path for a slug's credential file."""
        return self._credentials_dir() / f"{_slug_to_path(slug)}.json"

    def save_credentials(self, slug: str, blob: dict[str, Any]) -> None:
        """Persist the encoded credential *blob* for *slug* at mode 0600."""
        path = self._credential_path(slug)
        path.write_text(json.dumps(blob, indent=2), encoding="utf-8")
        try:
            path.chmod(0o600)
        except OSError:  # pragma: no cover
            pass

    def get_credentials(self, slug: str) -> dict[str, Any] | None:
        """Return the encoded credential blob for *slug*, or None."""
        path = self._credential_path(slug)
        if not path.exists():
            return None
        try:
            data = json.loads(path.read_text(encoding="utf-8"))
            return data if isinstance(data, dict) else None
        except Exception:
            return None

    def delete_credentials(self, slug: str) -> bool:
        """Delete *slug*'s credential file; return True if one existed."""
        path = self._credential_path(slug)
        if not path.exists():
            return False
        path.unlink()
        return True

    def list_credentials(self) -> dict[str, dict[str, Any]]:
        """Return every stored credential blob keyed by scope.

        The scope is read from the blob's ``scope`` field (stamped by
        ``CredentialStore.save``) — authoritative and lossless. Only a blob
        missing that field falls back to inverting :func:`_slug_to_path`
        (``__`` → ``/``), which corrupts a scope containing a literal ``__``;
        malformed files are skipped. Read-only: never creates the dir.
        """
        out: dict[str, dict[str, Any]] = {}
        cred_dir = self.root_dir / "credentials"
        if not cred_dir.exists():
            return out
        for path in cred_dir.glob("*.json"):
            try:
                data = json.loads(path.read_text(encoding="utf-8"))
            except Exception:
                logging.warning("Skipping malformed credential file {}", path)
                continue
            if isinstance(data, dict):
                blob_scope = data.get("scope")
                scope = blob_scope if isinstance(blob_scope, str) else path.stem.replace("__", "/")
                out[scope] = data
        return out

    # -- Restart-recovery counter (slug-keyed sidecar) -----------------------

    def _recovery_path(self, slug: str) -> Path:
        """Filesystem path for a slug's recovery-attempt counter file."""
        d = self.root_dir / "recovery"
        d.mkdir(parents=True, exist_ok=True)
        return d / f"{_slug_to_path(slug)}.json"

    def get_recovery_attempts(self, slug: str) -> int:
        """Return the recovery-attempt count for *slug* (0 if never recovered)."""
        path = self._recovery_path(slug)
        if not path.exists():
            return 0
        try:
            data = json.loads(path.read_text(encoding="utf-8"))
            return int(data.get("attempts", 0)) if isinstance(data, dict) else 0
        except Exception:
            return 0

    def bump_recovery_attempts(self, slug: str) -> int:
        """Atomically increment *slug*'s recovery counter; return the new value."""
        with self._lock:
            count = self.get_recovery_attempts(slug) + 1
            self._recovery_path(slug).write_text(
                json.dumps({"attempts": count}, indent=2), encoding="utf-8"
            )
            return count

    def reset_recovery_attempts(self, slug: str) -> None:
        """Clear *slug*'s recovery counter file (user-initiated resume fresh budget)."""
        with self._lock:
            path = self._recovery_path(slug)
            if path.exists():
                path.unlink()

    # -- QA ------------------------------------------------------------------

    def _qa_dir(self, answer_id: str) -> Path:
        """Directory for a QA answer's artefacts."""
        return self.root_dir / "qa" / answer_id

    def _qa_path(self, answer_id: str) -> Path:
        """Filesystem path for a QA answer JSON file."""
        return self._qa_dir(answer_id) / "answer.json"

    def _qa_session_path(self, answer_id: str) -> Path:
        """Filesystem path for the QA session-id text file."""
        return self._qa_dir(answer_id) / "session.txt"

    def save_qa(self, answer: QaAnswer) -> None:
        """Persist a QA answer record (``slug`` round-trips through answer.json)."""
        self._qa_dir(answer.answer_id).mkdir(parents=True, exist_ok=True)
        self._save_json(self._qa_path(answer.answer_id), answer)

    def update_qa_fields(self, answer: QaAnswer) -> None:
        """Non-destructive field update.

        Session + events are separate files here, so a plain answer.json rewrite
        already preserves them.
        """
        self._save_json(self._qa_path(answer.answer_id), answer)

    def get_qa(self, answer_id: str) -> QaAnswer | None:
        """Return the QA answer, or None if absent."""
        return self._load_json(self._qa_path(answer_id), QaAnswer)

    def list_qa(self, status: str | None = None) -> list[QaAnswer]:
        """Return all QA answers, optionally filtered to *status*.

        ``iterdir()`` order is arbitrary — this is a boot-time/offline scan
        (mirrors :meth:`list_jobs`), never an interactive listing, so no
        ordering guarantee is made or needed here.
        """
        qa_root = self.root_dir / "qa"
        if not qa_root.exists():
            return []
        answers: list[QaAnswer] = []
        for qa_dir in qa_root.iterdir():
            if not qa_dir.is_dir():
                continue
            answer = self._load_json(qa_dir / "answer.json", QaAnswer)
            if answer is None:
                continue
            if status is None or answer.status == status:
                answers.append(answer)
        return answers

    def attach_qa_session(self, answer_id: str, session_id: str) -> None:
        """Associate a Mewbo session_id with a QA answer."""
        path = self._qa_session_path(answer_id)
        path.parent.mkdir(parents=True, exist_ok=True)
        path.write_text(session_id, encoding="utf-8")

    def get_qa_session(self, answer_id: str) -> str | None:
        """Return the session_id attached to *answer_id*, or None."""
        path = self._qa_session_path(answer_id)
        if not path.exists():
            return None
        return path.read_text(encoding="utf-8").strip() or None

    def find_qa_by_session(self, session_id: str) -> str | None:
        """Reverse lookup: scan qa dirs for the session.txt that matches *session_id*."""
        qa_root = self.root_dir / "qa"
        if not qa_root.exists():
            return None
        for qa_dir in qa_root.iterdir():
            if not qa_dir.is_dir():
                continue
            sess_file = qa_dir / "session.txt"
            if sess_file.exists() and sess_file.read_text(encoding="utf-8").strip() == session_id:
                return qa_dir.name
        return None

    def append_qa_event(self, answer_id: str, event: dict[str, Any]) -> int:
        """Append *event* to the QA event log; return the monotonic idx."""
        return self._append_event("qa", answer_id, event)

    def load_qa_events(
        self, answer_id: str, after_idx: int = -1
    ) -> list[dict[str, Any]]:
        """Return QA events with idx > *after_idx* (-1 returns all)."""
        return self._load_events("qa", answer_id, after_idx)

    # -- Graph + embeddings --------------------------------------------------

    def _graph_dir(self, slug: str) -> Path:
        """Directory for per-slug graph artefacts."""
        return self.root_dir / "graph" / _slug_to_path(slug)

    def _nodes_path(self, slug: str) -> Path:
        return self._graph_dir(slug) / "nodes.jsonl"

    def _edges_path(self, slug: str) -> Path:
        return self._graph_dir(slug) / "edges.jsonl"

    def _embeddings_path(self, slug: str) -> Path:
        return self._graph_dir(slug) / "embeddings.jsonl"

    def _load_jsonl(self, path: Path, model_cls: type[_M]) -> list[_M]:
        """Load a JSONL file; skip malformed lines. Returns [] if absent."""
        if not path.exists():
            return []
        out: list[_M] = []
        for line in path.read_text(encoding="utf-8").splitlines():
            line = line.strip()
            if not line:
                continue
            try:
                out.append(model_cls.model_validate_json(line))
            except Exception:
                logging.warning("Skipping malformed line in {}", path)
        return out

    def _load_graph_nodes(self, path: Path) -> list[GraphNode]:
        """Load a graph-node JSONL, dispatching each line to its per-kind class.

        ``GraphNode`` is a discriminated union (schema v2), so validation goes
        through :data:`GraphNodeAdapter` — the ``type`` discriminator picks
        ``FileNode``/``ClassNode``/… ; a line without ``subkind``/
        ``attributes`` validates to the defaults. Mirrors ``_load_jsonl`` but
        can't reuse it (the union is not a single ``BaseModel`` subclass).
        """
        if not path.exists():
            return []
        out: list[GraphNode] = []
        for line in path.read_text(encoding="utf-8").splitlines():
            line = line.strip()
            if not line:
                continue
            try:
                out.append(GraphNodeAdapter.validate_json(line))
            except Exception:
                logging.warning("Skipping malformed line in {}", path)
        return out

    def _write_jsonl(self, path: Path, items: list[Any]) -> None:
        """Atomically rewrite a JSONL file (tmp + rename)."""
        path.parent.mkdir(parents=True, exist_ok=True)
        tmp = path.with_suffix(".tmp")
        tmp.write_text(
            "\n".join(item.model_dump_json(by_alias=True) for item in items) + "\n",
            encoding="utf-8",
        )
        tmp.replace(path)

    def upsert_nodes(
        self,
        slug: str,
        nodes: Iterable[GraphNode],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
        on_progress: Callable[[int, int], None] | None = None,
    ) -> None:
        """Upsert graph nodes for *slug*; dedup by node_id, stamp attribution.

        Cost: ``O(nodes)`` — offline. The JSON driver writes one atomic file, so
        one non-empty input is one persistence batch.
        """
        items = list(nodes)
        with self._lock:
            existing = {n.node_id: n for n in self._load_graph_nodes(self._nodes_path(slug))}
            for node in items:
                existing[node.node_id] = self._stamp_attribution(node, commit_sha, job_id)
            self._write_jsonl(self._nodes_path(slug), list(existing.values()))
        if items and on_progress is not None:
            on_progress(1, 1)

    def upsert_edges(
        self,
        slug: str,
        edges: Iterable[GraphEdge],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
        on_progress: Callable[[int, int], None] | None = None,
    ) -> None:
        """Upsert graph edges for *slug*; dedup by (source, target, type).

        Cost: ``O(edges)`` — offline. The JSON driver writes one atomic file, so
        one non-empty input is one persistence batch.
        """
        items = list(edges)
        with self._lock:
            existing = {
                (e.source, e.target, e.type): e
                for e in self._load_jsonl(self._edges_path(slug), GraphEdge)
            }
            for edge in items:
                existing[(edge.source, edge.target, edge.type)] = self._stamp_attribution(
                    edge, commit_sha, job_id
                )
            self._write_jsonl(self._edges_path(slug), list(existing.values()))
        if items and on_progress is not None:
            on_progress(1, 1)

    def upsert_embeddings(
        self,
        slug: str,
        items: Iterable[Embedding],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert embedding vectors for *slug*; dedup by node_id."""
        with self._lock:
            existing = {
                e.node_id: e
                for e in self._load_jsonl(self._embeddings_path(slug), Embedding)
            }
            for item in items:
                existing[item.node_id] = self._stamp_attribution(item, commit_sha, job_id)
            self._write_jsonl(self._embeddings_path(slug), list(existing.values()))

    def query_graph(
        self,
        slug: str,
        *,
        scope: CommitScope,
        node_type: str | None = None,
        name_match: str | None = None,
        neighbors_of: str | None = None,
        node_ids: Collection[str] | None = None,
    ) -> list[GraphNode]:
        """See ``WikiStoreBase.query_graph``."""
        if neighbors_of is not None:
            # The seed's own generation bounds the walk: an edge is only a
            # neighbour relation if it belongs to the same generation as the
            # nodes we are about to return, or the neighbourhood spans commits.
            edges = [
                e
                for e in self._load_jsonl(self._edges_path(slug), GraphEdge)
                if scope.matches(e.commit_sha)
            ]
            related_ids: set[str] = set()
            for edge in edges:
                if edge.source == neighbors_of:
                    related_ids.add(edge.target)
                elif edge.target == neighbors_of:
                    related_ids.add(edge.source)
            all_nodes = self._load_graph_nodes(self._nodes_path(slug))
            return [
                n
                for n in all_nodes
                if n.node_id in related_ids and scope.matches(n.commit_sha)
            ]
        nodes = [
            n
            for n in self._load_graph_nodes(self._nodes_path(slug))
            if scope.matches(n.commit_sha)
        ]
        if node_ids is not None:
            wanted = set(node_ids)
            nodes = [n for n in nodes if n.node_id in wanted]
        if node_type is not None:
            nodes = [n for n in nodes if n.type == node_type]
        if name_match is not None:
            lower = name_match.lower()
            nodes = [n for n in nodes if lower in n.name.lower()]
        return nodes

    def list_edges(self, slug: str, *, scope: CommitScope) -> list[GraphEdge]:
        """See ``WikiStoreBase.list_edges``."""
        return [
            e
            for e in self._load_jsonl(self._edges_path(slug), GraphEdge)
            if scope.matches(e.commit_sha)
        ]

    def count_graph_nodes(self, slug: str, *, commit_sha: str | None) -> int:
        """Count *slug* nodes stamped exactly *commit_sha* (``None`` matches None)."""
        return sum(
            1
            for n in self._load_graph_nodes(self._nodes_path(slug))
            if n.commit_sha == commit_sha
        )

    def supersede_graph_artifacts(
        self, slug: str, *, keep_commit_sha: str
    ) -> dict[str, int]:
        """Drop prior-commit graph + entity artifacts, preserving ``None``-stamped rows.

        A row survives iff its ``commit_sha`` is ``None`` (a QA-minted or
        pre-isolation record) OR equals *keep_commit_sha*. Everything else — the
        artifacts of a superseded commit — is dropped.
        """
        counts: dict[str, int] = {}
        with self._lock:
            counts["nodes"] = self._retain_jsonl(
                self._nodes_path(slug), self._load_graph_nodes, keep_commit_sha
            )
            counts["edges"] = self._retain_jsonl(
                self._edges_path(slug),
                lambda p: self._load_jsonl(p, GraphEdge),
                keep_commit_sha,
            )
            counts["embeddings"] = self._retain_jsonl(
                self._embeddings_path(slug),
                lambda p: self._load_jsonl(p, Embedding),
                keep_commit_sha,
            )
            counts["entities"] = self._retain_jsonl(
                self._entities_path(slug),
                lambda p: self._load_jsonl(p, Entity),
                keep_commit_sha,
            )
            counts["entity_edges"] = self._retain_jsonl(
                self._entity_edges_path(slug),
                lambda p: self._load_jsonl(p, EntityRelation),
                keep_commit_sha,
            )
            counts["entity_embeddings"] = self._retain_jsonl(
                self._entity_embeddings_path(slug),
                lambda p: self._load_jsonl(p, EntityEmbedding),
                keep_commit_sha,
            )
        return counts

    def restamp_graph_artifacts(
        self, slug: str, *, from_commit: str, to_commit: str
    ) -> dict[str, int]:
        """See ``WikiStoreBase.restamp_graph_artifacts``."""
        counts: dict[str, int] = {}
        with self._lock:
            for key, path, loader in (
                ("nodes", self._nodes_path(slug), self._load_graph_nodes),
                ("edges", self._edges_path(slug),
                 lambda p: self._load_jsonl(p, GraphEdge)),
                ("embeddings", self._embeddings_path(slug),
                 lambda p: self._load_jsonl(p, Embedding)),
                ("entities", self._entities_path(slug),
                 lambda p: self._load_jsonl(p, Entity)),
                ("entity_edges", self._entity_edges_path(slug),
                 lambda p: self._load_jsonl(p, EntityRelation)),
                ("entity_embeddings", self._entity_embeddings_path(slug),
                 lambda p: self._load_jsonl(p, EntityEmbedding)),
            ):
                counts[key] = self._restamp_jsonl(path, loader, from_commit, to_commit)
        return counts

    def _restamp_jsonl(
        self, path: Path, loader: Any, from_commit: str, to_commit: str
    ) -> int:
        """Rewrite *path* moving ``from_commit`` rows to ``to_commit``; return #moved.

        Graph nodes are ``frozen=True``, so a row is REPLACED via ``model_copy``
        rather than mutated in place — the same reason hierarchy stamping happens
        on the wire dict rather than on the node.
        """
        items = loader(path)
        moved = 0
        out = []
        for it in items:
            if it.commit_sha == from_commit:
                out.append(it.model_copy(update={"commit_sha": to_commit}))
                moved += 1
            else:
                out.append(it)
        if moved:
            self._write_jsonl(path, out)
        return moved

    def _retain_jsonl(
        self, path: Path, loader: Any, keep_commit_sha: str
    ) -> int:
        """Rewrite *path* keeping only ``None``/``keep_commit_sha`` rows; return #dropped."""
        items = loader(path)
        kept = [
            it for it in items
            if it.commit_sha is None or it.commit_sha == keep_commit_sha
        ]
        dropped = len(items) - len(kept)
        if dropped:
            self._write_jsonl(path, kept)
        return dropped

    def vector_search(self, slug: str, qvec: list[float], k: int = 10) -> list[Embedding]:
        """Return top-k embeddings for *slug* by cosine similarity."""
        from .embedder import Embedder

        pool = self._load_jsonl(self._embeddings_path(slug), Embedding)
        if not pool:
            return []
        scored = [(emb, Embedder.cosine(qvec, emb.vector)) for emb in pool]
        scored.sort(key=lambda t: t[1], reverse=True)
        return [emb for emb, _ in scored[:k]]

    # -- Scoped graph deletes (incremental retract) --------------------------

    def delete_nodes_by_file(self, slug: str, file: str) -> int:
        """Delete *file*'s code nodes AND their vectors; return the NODE count."""
        with self._lock:
            nodes = self._load_graph_nodes(self._nodes_path(slug))
            keep = [n for n in nodes if n.file != file]
            removed = len(nodes) - len(keep)
            if not removed:
                return 0
            self._write_jsonl(self._nodes_path(slug), keep)
            doomed = {n.node_id for n in nodes if n.file == file}
            vectors = self._load_jsonl(self._embeddings_path(slug), Embedding)
            kept_vectors = [e for e in vectors if e.node_id not in doomed]
            if len(kept_vectors) != len(vectors):
                self._write_jsonl(self._embeddings_path(slug), kept_vectors)
            return removed

    def delete_edges_by_source_file(self, slug: str, file: str) -> int:
        """Delete edges whose ``source`` node belongs to *file*; return count."""
        with self._lock:
            file_ids = {
                n.node_id
                for n in self._load_graph_nodes(self._nodes_path(slug))
                if n.file == file
            }
            if not file_ids:
                return 0
            edges = self._load_jsonl(self._edges_path(slug), GraphEdge)
            keep = [e for e in edges if e.source not in file_ids]
            removed = len(edges) - len(keep)
            if removed:
                self._write_jsonl(self._edges_path(slug), keep)
            return removed

    # -- Memory layer (multiplex overlay) ------------------------------------

    def _memory_dir(self, slug: str) -> Path:
        """Directory for per-slug memory-layer artefacts."""
        return self.root_dir / "memory" / _slug_to_path(slug)

    def _memory_nodes_path(self, slug: str) -> Path:
        return self._memory_dir(slug) / "nodes.jsonl"

    def _memory_edges_path(self, slug: str) -> Path:
        return self._memory_dir(slug) / "edges.jsonl"

    def _memory_embeddings_path(self, slug: str) -> Path:
        return self._memory_dir(slug) / "embeddings.jsonl"

    def _doc_notes_path(self, slug: str) -> Path:
        return self._memory_dir(slug) / "docs.jsonl"

    def _manifest_path(self, slug: str) -> Path:
        return self._memory_dir(slug) / "manifest.jsonl"

    def upsert_memory_nodes(self, slug: str, nodes: Iterable[MemoryNode]) -> None:
        """Upsert memory nodes for *slug*; dedup by node_id."""
        with self._lock:
            existing = {
                n.node_id: n
                for n in self._load_jsonl(self._memory_nodes_path(slug), MemoryNode)
            }
            for node in nodes:
                existing[node.node_id] = node
            self._write_jsonl(self._memory_nodes_path(slug), list(existing.values()))

    def get_memory_node(self, slug: str, node_id: str) -> MemoryNode | None:
        """Return a single memory node, or None if absent."""
        for n in self._load_jsonl(self._memory_nodes_path(slug), MemoryNode):
            if n.node_id == node_id:
                return n
        return None

    def delete_memory_node(self, slug: str, node_id: str) -> bool:
        """Delete a memory node + its embedding; return True if one was removed."""
        with self._lock:
            nodes = self._load_jsonl(self._memory_nodes_path(slug), MemoryNode)
            keep = [n for n in nodes if n.node_id != node_id]
            removed = len(keep) != len(nodes)
            if removed:
                self._write_jsonl(self._memory_nodes_path(slug), keep)
                embs = self._load_jsonl(
                    self._memory_embeddings_path(slug), MemoryEmbedding
                )
                kept_embs = [e for e in embs if e.node_id != node_id]
                if len(kept_embs) != len(embs):
                    self._write_jsonl(self._memory_embeddings_path(slug), kept_embs)
            return removed

    def query_memory(
        self, slug: str, *, filt: MemoryFilter | None = None
    ) -> list[MemoryNode]:
        """Return memory nodes matching *filt*'s node-level facets."""
        nodes = self._load_jsonl(self._memory_nodes_path(slug), MemoryNode)
        if filt is None:
            return nodes
        return [n for n in nodes if filt.matches_node(n)]

    def upsert_memory_edges(self, slug: str, edges: Iterable[MemoryEdge]) -> None:
        """Upsert memory edges for *slug*; dedup by (source, target, type)."""
        with self._lock:
            existing = {
                (e.source, e.target, e.type): e
                for e in self._load_jsonl(self._memory_edges_path(slug), MemoryEdge)
            }
            for edge in edges:
                existing[(edge.source, edge.target, edge.type)] = edge
            self._write_jsonl(self._memory_edges_path(slug), list(existing.values()))

    def list_memory_edges(
        self,
        slug: str,
        *,
        node_id: str | None = None,
        include_invalidated: bool = False,
    ) -> list[MemoryEdge]:
        """Return memory edges, optionally scoped to ``source == node_id``."""
        out: list[MemoryEdge] = []
        for e in self._load_jsonl(self._memory_edges_path(slug), MemoryEdge):
            if node_id is not None and e.source != node_id:
                continue
            if e.invalid_at is not None and not include_invalidated:
                continue
            out.append(e)
        return out

    def memories_anchored_to(
        self,
        slug: str,
        entity_keys: Iterable[EntityKey],
        *,
        include_invalidated: bool = False,
    ) -> list[str]:
        """Reverse ANCHORS lookup: entity_keys → distinct memory node_ids."""
        keys = set(entity_keys)
        seen: list[str] = []
        seen_set: set[str] = set()
        for e in self._load_jsonl(self._memory_edges_path(slug), MemoryEdge):
            if e.type != "ANCHORS" or e.target not in keys:
                continue
            if e.invalid_at is not None and not include_invalidated:
                continue
            if e.source not in seen_set:
                seen_set.add(e.source)
                seen.append(e.source)
        return seen

    def _live_anchored_ids(self, slug: str) -> set[str]:
        """Memory node_ids with ≥1 live ANCHORS edge."""
        return {
            e.source
            for e in self._load_jsonl(self._memory_edges_path(slug), MemoryEdge)
            if e.type == "ANCHORS" and e.invalid_at is None
        }

    def upsert_memory_embeddings(
        self, slug: str, items: Iterable[MemoryEmbedding]
    ) -> None:
        """Upsert memory embedding vectors for *slug*; dedup by node_id."""
        with self._lock:
            existing = {
                e.node_id: e
                for e in self._load_jsonl(
                    self._memory_embeddings_path(slug), MemoryEmbedding
                )
            }
            for item in items:
                existing[item.node_id] = item
            self._write_jsonl(
                self._memory_embeddings_path(slug), list(existing.values())
            )

    def memory_vector_search(
        self,
        slug: str,
        qvec: list[float],
        k: int = 10,
        *,
        filt: MemoryFilter | None = None,
    ) -> list[MemoryEmbedding]:
        """Top-k memory embeddings by cosine, after applying *filt*."""
        pool = self._load_jsonl(self._memory_embeddings_path(slug), MemoryEmbedding)
        return self._rank_memory(slug, pool, qvec, k, filt)

    # -- Doc-page notes ------------------------------------------------------

    def upsert_doc_notes(self, slug: str, notes: Iterable[DocPageNote]) -> None:
        """Upsert doc-page notes for *slug*; dedup by page_id."""
        with self._lock:
            existing = {
                d.page_id: d
                for d in self._load_jsonl(self._doc_notes_path(slug), DocPageNote)
            }
            for note in notes:
                existing[note.page_id] = note
            self._write_jsonl(self._doc_notes_path(slug), list(existing.values()))

    def get_doc_note(self, slug: str, page_id: str) -> DocPageNote | None:
        """Return a single doc-page note, or None if absent."""
        for d in self._load_jsonl(self._doc_notes_path(slug), DocPageNote):
            if d.page_id == page_id:
                return d
        return None

    def list_doc_notes(self, slug: str) -> list[DocPageNote]:
        """Return every doc-page note for *slug*."""
        return self._load_jsonl(self._doc_notes_path(slug), DocPageNote)

    def delete_doc_note(self, slug: str, page_id: str) -> bool:
        """Delete a doc-page note; return True if one was removed."""
        with self._lock:
            notes = self._load_jsonl(self._doc_notes_path(slug), DocPageNote)
            keep = [d for d in notes if d.page_id != page_id]
            if len(keep) == len(notes):
                return False
            self._write_jsonl(self._doc_notes_path(slug), keep)
            return True

    # -- File manifest -------------------------------------------------------

    def upsert_file_manifest(
        self, slug: str, entries: Iterable[FileManifest]
    ) -> None:
        """Upsert file-manifest entries for *slug*; dedup by path."""
        with self._lock:
            existing = {
                m.path: m
                for m in self._load_jsonl(self._manifest_path(slug), FileManifest)
            }
            for entry in entries:
                existing[entry.path] = entry
            self._write_jsonl(self._manifest_path(slug), list(existing.values()))

    def get_file_manifest(self, slug: str, path: str) -> FileManifest | None:
        """Return a single file-manifest entry, or None if absent."""
        for m in self._load_jsonl(self._manifest_path(slug), FileManifest):
            if m.path == path:
                return m
        return None

    def list_file_manifest(self, slug: str) -> list[FileManifest]:
        """Return every file-manifest entry for *slug*."""
        return self._load_jsonl(self._manifest_path(slug), FileManifest)

    def delete_file_manifest(self, slug: str, path: str) -> bool:
        """Delete a file-manifest entry; return True if one was removed."""
        with self._lock:
            entries = self._load_jsonl(self._manifest_path(slug), FileManifest)
            keep = [m for m in entries if m.path != path]
            if len(keep) == len(entries):
                return False
            self._write_jsonl(self._manifest_path(slug), keep)
            return True

    # -- Abstract-entity layer (multiplex overlay) ---------------------------
    #
    # Persisted as JSONL under the same per-slug memory dir as memory nodes,
    # reusing the exact ``_load_jsonl`` / ``_write_jsonl`` upsert idiom so the
    # entity overlay can never desync from the memory overlay's conventions.

    def _entities_path(self, slug: str) -> Path:
        return self._memory_dir(slug) / "entities.jsonl"

    def _entity_embeddings_path(self, slug: str) -> Path:
        return self._memory_dir(slug) / "entity_embeddings.jsonl"

    def _entity_edges_path(self, slug: str) -> Path:
        return self._memory_dir(slug) / "entity_edges.jsonl"

    def _entity_recs_path(self, slug: str) -> Path:
        return self._memory_dir(slug) / "entity_recommendations.jsonl"

    def upsert_entities(
        self,
        slug: str,
        entities: Iterable[Entity],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert entities for *slug*; dedup by id, stamp attribution."""
        with self._lock:
            existing = {
                e.id: e for e in self._load_jsonl(self._entities_path(slug), Entity)
            }
            for entity in entities:
                existing[entity.id] = self._stamp_attribution(entity, commit_sha, job_id)
            self._write_jsonl(self._entities_path(slug), list(existing.values()))

    def get_entity(self, slug: str, entity_id: str) -> Entity | None:
        """Return a single entity, or None if absent."""
        for e in self._load_jsonl(self._entities_path(slug), Entity):
            if e.id == entity_id:
                return e
        return None

    def query_entities(
        self, slug: str, *, filt: EntityFilter | None = None
    ) -> list[Entity]:
        """Return entities matching *filt*'s facets."""
        entities = self._load_jsonl(self._entities_path(slug), Entity)
        if filt is None:
            return entities
        return [e for e in entities if filt.matches(e)]

    def count_entities(self, slug: str, *, commit_sha: str | None) -> int:
        """Count *slug* entities stamped exactly *commit_sha* (``None`` matches None)."""
        return sum(
            1
            for e in self._load_jsonl(self._entities_path(slug), Entity)
            if e.commit_sha == commit_sha
        )

    def upsert_entity_embeddings(
        self,
        slug: str,
        items: Iterable[EntityEmbedding],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert entity embedding vectors for *slug*; dedup by entity_id."""
        with self._lock:
            existing = {
                e.entity_id: e
                for e in self._load_jsonl(
                    self._entity_embeddings_path(slug), EntityEmbedding
                )
            }
            for item in items:
                existing[item.entity_id] = self._stamp_attribution(
                    item, commit_sha, job_id
                )
            self._write_jsonl(
                self._entity_embeddings_path(slug), list(existing.values())
            )

    def entity_vector_search(
        self, slug: str, qvec: list[float], k: int = 10
    ) -> list[EntityEmbedding]:
        """Return top-k entity embeddings for *slug* by cosine similarity."""
        pool = self._load_jsonl(self._entity_embeddings_path(slug), EntityEmbedding)
        return self._rank_embeddings(pool, qvec, k)

    def upsert_entity_edges(
        self,
        slug: str,
        edges: Iterable[EntityRelation],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert entity relations for *slug*; dedup by id."""
        with self._lock:
            existing = {
                e.id: e
                for e in self._load_jsonl(self._entity_edges_path(slug), EntityRelation)
            }
            for edge in edges:
                existing[edge.id] = self._stamp_attribution(edge, commit_sha, job_id)
            self._write_jsonl(self._entity_edges_path(slug), list(existing.values()))

    def list_entity_edges(
        self, slug: str, *, source_id: str | None = None
    ) -> list[EntityRelation]:
        """Return entity relations, optionally scoped to ``source_id``."""
        out = self._load_jsonl(self._entity_edges_path(slug), EntityRelation)
        if source_id is not None:
            out = [e for e in out if e.source_id == source_id]
        return out

    def save_entity_recommendation(
        self, slug: str, rec: EntityRecommendation
    ) -> None:
        """Upsert a recommendation for *slug*; dedup by id (a replay converges)."""
        with self._lock:
            existing = {
                r.id: r
                for r in self._load_jsonl(
                    self._entity_recs_path(slug), EntityRecommendation
                )
            }
            existing[rec.id] = rec
            self._write_jsonl(self._entity_recs_path(slug), list(existing.values()))

    def get_entity_recommendations(self, slug: str) -> list[EntityRecommendation]:
        """Return every persisted entity recommendation for *slug*."""
        return self._load_jsonl(self._entity_recs_path(slug), EntityRecommendation)

__init__(root_dir: str | Path | None = None) -> None

Initialise and create the directory tree.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
def __init__(self, root_dir: str | Path | None = None) -> None:
    """Initialise and create the directory tree."""
    if root_dir is None:
        home = get_config_value("runtime", "cache_dir", default="") or ".mewbo"
        root_dir = Path(home) / "wiki"
    self.root_dir = Path(root_dir)
    self.root_dir.mkdir(parents=True, exist_ok=True)
    for sub in ("projects", "pages", "jobs", "qa"):
        (self.root_dir / sub).mkdir(parents=True, exist_ok=True)
    self._lock = threading.Lock()

append_qa_event(answer_id: str, event: dict[str, Any]) -> int

Append event to the QA event log; return the monotonic idx.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1960
1961
1962
def append_qa_event(self, answer_id: str, event: dict[str, Any]) -> int:
    """Append *event* to the QA event log; return the monotonic idx."""
    return self._append_event("qa", answer_id, event)

attach_job_session(job_id: str, session_id: str) -> None

Associate a Mewbo session_id with an indexing job.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1605
1606
1607
1608
1609
def attach_job_session(self, job_id: str, session_id: str) -> None:
    """Associate a Mewbo session_id with an indexing job."""
    path = self._session_path(job_id)
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(session_id, encoding="utf-8")

attach_qa_session(answer_id: str, session_id: str) -> None

Associate a Mewbo session_id with a QA answer.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1934
1935
1936
1937
1938
def attach_qa_session(self, answer_id: str, session_id: str) -> None:
    """Associate a Mewbo session_id with a QA answer."""
    path = self._qa_session_path(answer_id)
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(session_id, encoding="utf-8")

bump_recovery_attempts(slug: str) -> int

Atomically increment slug's recovery counter; return the new value.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1866
1867
1868
1869
1870
1871
1872
1873
def bump_recovery_attempts(self, slug: str) -> int:
    """Atomically increment *slug*'s recovery counter; return the new value."""
    with self._lock:
        count = self.get_recovery_attempts(slug) + 1
        self._recovery_path(slug).write_text(
            json.dumps({"attempts": count}, indent=2), encoding="utf-8"
        )
        return count

cancel_job(job_id: str) -> bool

Cancel job_id; return True on first cancel, False if already cancelled.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
def cancel_job(self, job_id: str) -> bool:
    """Cancel *job_id*; return True on first cancel, False if already cancelled."""
    job = self.get_job(job_id)
    if job is None:
        return False
    if job.status == "cancelled":
        return False
    self.update_job(job_id, status="cancelled")
    self.append_job_event(job_id, {"type": "cancelled"})
    return True

claim_job_page(slug: str, job_id: str, page_id: str) -> PageClaim

Record page_id under job_id, under the lock; count = set size.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
def claim_job_page(self, slug: str, job_id: str, page_id: str) -> PageClaim:
    """Record *page_id* under *job_id*, under the lock; count = set size."""
    if self.get_job(job_id) is None:
        raise KeyError(f"Job not found: {job_id}")
    with self._lock:
        meta = self._load_job_meta(job_id)
        claimed = self._claimed_ids(meta, slug, job_id)
        if page_id in claimed:
            return PageClaim(count=len(claimed), is_new=False)
        claimed.append(page_id)
        meta["submitted_page_ids"] = claimed
        # Kept in step purely so ``get_job_submitted_count`` stays a cheap
        # read; the SET is the authority, so the two can never disagree.
        meta["submitted_pages"] = len(claimed)
        path = self._job_meta_path(job_id)
        path.parent.mkdir(parents=True, exist_ok=True)
        path.write_text(json.dumps(meta, indent=2), encoding="utf-8")
        return PageClaim(count=len(claimed), is_new=True)

count_entities(slug: str, *, commit_sha: str | None) -> int

Count slug entities stamped exactly commit_sha (None matches None).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2580
2581
2582
2583
2584
2585
2586
def count_entities(self, slug: str, *, commit_sha: str | None) -> int:
    """Count *slug* entities stamped exactly *commit_sha* (``None`` matches None)."""
    return sum(
        1
        for e in self._load_jsonl(self._entities_path(slug), Entity)
        if e.commit_sha == commit_sha
    )

count_graph_nodes(slug: str, *, commit_sha: str | None) -> int

Count slug nodes stamped exactly commit_sha (None matches None).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2156
2157
2158
2159
2160
2161
2162
def count_graph_nodes(self, slug: str, *, commit_sha: str | None) -> int:
    """Count *slug* nodes stamped exactly *commit_sha* (``None`` matches None)."""
    return sum(
        1
        for n in self._load_graph_nodes(self._nodes_path(slug))
        if n.commit_sha == commit_sha
    )

create_job(job: IndexingJob) -> None

Persist a new indexing job.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1525
1526
1527
1528
def create_job(self, job: IndexingJob) -> None:
    """Persist a new indexing job."""
    self._job_dir(job.job_id).mkdir(parents=True, exist_ok=True)
    self._save_json(self._job_path(job.job_id), job)

create_project(project: Project) -> None

Persist a new project record.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1256
1257
1258
def create_project(self, project: Project) -> None:
    """Persist a new project record."""
    self._save_json(self._project_path(project.slug), project)

delete_credentials(slug: str) -> bool

Delete slug's credential file; return True if one existed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1814
1815
1816
1817
1818
1819
1820
def delete_credentials(self, slug: str) -> bool:
    """Delete *slug*'s credential file; return True if one existed."""
    path = self._credential_path(slug)
    if not path.exists():
        return False
    path.unlink()
    return True

delete_doc_note(slug: str, page_id: str) -> bool

Delete a doc-page note; return True if one was removed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2483
2484
2485
2486
2487
2488
2489
2490
2491
def delete_doc_note(self, slug: str, page_id: str) -> bool:
    """Delete a doc-page note; return True if one was removed."""
    with self._lock:
        notes = self._load_jsonl(self._doc_notes_path(slug), DocPageNote)
        keep = [d for d in notes if d.page_id != page_id]
        if len(keep) == len(notes):
            return False
        self._write_jsonl(self._doc_notes_path(slug), keep)
        return True

delete_edges_by_source_file(slug: str, file: str) -> int

Delete edges whose source node belongs to file; return count.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
def delete_edges_by_source_file(self, slug: str, file: str) -> int:
    """Delete edges whose ``source`` node belongs to *file*; return count."""
    with self._lock:
        file_ids = {
            n.node_id
            for n in self._load_graph_nodes(self._nodes_path(slug))
            if n.file == file
        }
        if not file_ids:
            return 0
        edges = self._load_jsonl(self._edges_path(slug), GraphEdge)
        keep = [e for e in edges if e.source not in file_ids]
        removed = len(edges) - len(keep)
        if removed:
            self._write_jsonl(self._edges_path(slug), keep)
        return removed

delete_file_manifest(slug: str, path: str) -> bool

Delete a file-manifest entry; return True if one was removed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2519
2520
2521
2522
2523
2524
2525
2526
2527
def delete_file_manifest(self, slug: str, path: str) -> bool:
    """Delete a file-manifest entry; return True if one was removed."""
    with self._lock:
        entries = self._load_jsonl(self._manifest_path(slug), FileManifest)
        keep = [m for m in entries if m.path != path]
        if len(keep) == len(entries):
            return False
        self._write_jsonl(self._manifest_path(slug), keep)
        return True

delete_memory_node(slug: str, node_id: str) -> bool

Delete a memory node + its embedding; return True if one was removed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2348
2349
2350
2351
2352
2353
2354
2355
2356
2357
2358
2359
2360
2361
2362
def delete_memory_node(self, slug: str, node_id: str) -> bool:
    """Delete a memory node + its embedding; return True if one was removed."""
    with self._lock:
        nodes = self._load_jsonl(self._memory_nodes_path(slug), MemoryNode)
        keep = [n for n in nodes if n.node_id != node_id]
        removed = len(keep) != len(nodes)
        if removed:
            self._write_jsonl(self._memory_nodes_path(slug), keep)
            embs = self._load_jsonl(
                self._memory_embeddings_path(slug), MemoryEmbedding
            )
            kept_embs = [e for e in embs if e.node_id != node_id]
            if len(kept_embs) != len(embs):
                self._write_jsonl(self._memory_embeddings_path(slug), kept_embs)
        return removed

delete_nodes_by_file(slug: str, file: str) -> int

Delete file's code nodes AND their vectors; return the NODE count.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2276
2277
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
def delete_nodes_by_file(self, slug: str, file: str) -> int:
    """Delete *file*'s code nodes AND their vectors; return the NODE count."""
    with self._lock:
        nodes = self._load_graph_nodes(self._nodes_path(slug))
        keep = [n for n in nodes if n.file != file]
        removed = len(nodes) - len(keep)
        if not removed:
            return 0
        self._write_jsonl(self._nodes_path(slug), keep)
        doomed = {n.node_id for n in nodes if n.file == file}
        vectors = self._load_jsonl(self._embeddings_path(slug), Embedding)
        kept_vectors = [e for e in vectors if e.node_id not in doomed]
        if len(kept_vectors) != len(vectors):
            self._write_jsonl(self._embeddings_path(slug), kept_vectors)
        return removed

delete_page(slug: str, page_id: str) -> bool

Delete a single page on disk + drop it from the index.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
def delete_page(self, slug: str, page_id: str) -> bool:
    """Delete a single page on disk + drop it from the index."""
    path = self._page_path(slug, page_id)
    removed = path.exists()
    if removed:
        path.unlink()
    index = self._load_index(slug)
    if index.pop(page_id, None) is not None:
        self._index_path(slug).write_text(
            json.dumps(index, indent=2), encoding="utf-8"
        )
        removed = True
    return removed

delete_project(slug: str) -> bool

Delete project slug; return True if deleted, False if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1273
1274
1275
1276
1277
1278
1279
def delete_project(self, slug: str) -> bool:
    """Delete project *slug*; return True if deleted, False if absent."""
    path = self._project_path(slug)
    if not path.exists():
        return False
    path.unlink()
    return True

delete_project_settings(slug: str) -> bool

Delete slug's settings file; return True if one existed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1297
1298
1299
1300
1301
1302
1303
def delete_project_settings(self, slug: str) -> bool:
    """Delete *slug*'s settings file; return True if one existed."""
    path = self._settings_path(slug)
    if not path.exists():
        return False
    path.unlink()
    return True

Return top-k entity embeddings for slug by cosine similarity.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2612
2613
2614
2615
2616
2617
def entity_vector_search(
    self, slug: str, qvec: list[float], k: int = 10
) -> list[EntityEmbedding]:
    """Return top-k entity embeddings for *slug* by cosine similarity."""
    pool = self._load_jsonl(self._entity_embeddings_path(slug), EntityEmbedding)
    return self._rank_embeddings(pool, qvec, k)

find_job_by_session(session_id: str) -> str | None

Reverse lookup: scan job dirs for the session.txt that matches session_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
def find_job_by_session(self, session_id: str) -> str | None:
    """Reverse lookup: scan job dirs for the session.txt that matches *session_id*."""
    jobs_root = self.root_dir / "jobs"
    if not jobs_root.exists():
        return None
    for job_dir in jobs_root.iterdir():
        if not job_dir.is_dir():
            continue
        sess_file = job_dir / "session.txt"
        if sess_file.exists() and sess_file.read_text(encoding="utf-8").strip() == session_id:
            return job_dir.name
    return None

find_qa_by_session(session_id: str) -> str | None

Reverse lookup: scan qa dirs for the session.txt that matches session_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
def find_qa_by_session(self, session_id: str) -> str | None:
    """Reverse lookup: scan qa dirs for the session.txt that matches *session_id*."""
    qa_root = self.root_dir / "qa"
    if not qa_root.exists():
        return None
    for qa_dir in qa_root.iterdir():
        if not qa_dir.is_dir():
            continue
        sess_file = qa_dir / "session.txt"
        if sess_file.exists() and sess_file.read_text(encoding="utf-8").strip() == session_id:
            return qa_dir.name
    return None

get_act_plan(job_id: str) -> dict[str, Any] | None

Return the persisted act-stage record, or None if stage 1 hasn't finished.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
def get_act_plan(self, job_id: str) -> dict[str, Any] | None:
    """Return the persisted act-stage record, or None if stage 1 hasn't finished."""
    path = self._job_act_path(job_id)
    if not path.exists():
        return None
    try:
        data = json.loads(path.read_text(encoding="utf-8"))
        return data if isinstance(data, dict) else None
    except Exception:
        return None

get_credentials(slug: str) -> dict[str, Any] | None

Return the encoded credential blob for slug, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
def get_credentials(self, slug: str) -> dict[str, Any] | None:
    """Return the encoded credential blob for *slug*, or None."""
    path = self._credential_path(slug)
    if not path.exists():
        return None
    try:
        data = json.loads(path.read_text(encoding="utf-8"))
        return data if isinstance(data, dict) else None
    except Exception:
        return None

get_doc_note(slug: str, page_id: str) -> DocPageNote | None

Return a single doc-page note, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2472
2473
2474
2475
2476
2477
def get_doc_note(self, slug: str, page_id: str) -> DocPageNote | None:
    """Return a single doc-page note, or None if absent."""
    for d in self._load_jsonl(self._doc_notes_path(slug), DocPageNote):
        if d.page_id == page_id:
            return d
    return None

get_entity(slug: str, entity_id: str) -> Entity | None

Return a single entity, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2564
2565
2566
2567
2568
2569
def get_entity(self, slug: str, entity_id: str) -> Entity | None:
    """Return a single entity, or None if absent."""
    for e in self._load_jsonl(self._entities_path(slug), Entity):
        if e.id == entity_id:
            return e
    return None

get_entity_recommendations(slug: str) -> list[EntityRecommendation]

Return every persisted entity recommendation for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2660
2661
2662
def get_entity_recommendations(self, slug: str) -> list[EntityRecommendation]:
    """Return every persisted entity recommendation for *slug*."""
    return self._load_jsonl(self._entity_recs_path(slug), EntityRecommendation)

get_file_manifest(slug: str, path: str) -> FileManifest | None

Return a single file-manifest entry, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2508
2509
2510
2511
2512
2513
def get_file_manifest(self, slug: str, path: str) -> FileManifest | None:
    """Return a single file-manifest entry, or None if absent."""
    for m in self._load_jsonl(self._manifest_path(slug), FileManifest):
        if m.path == path:
            return m
    return None

get_job(job_id: str) -> IndexingJob | None

Return the indexing job, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1530
1531
1532
def get_job(self, job_id: str) -> IndexingJob | None:
    """Return the indexing job, or None if absent."""
    return self._load_json(self._job_path(job_id), IndexingJob)

get_job_page_ids(slug: str, job_id: str) -> frozenset[str]

Return the page ids job_id wrote (claim record, else attribution).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1745
1746
1747
def get_job_page_ids(self, slug: str, job_id: str) -> frozenset[str]:
    """Return the page ids *job_id* wrote (claim record, else attribution)."""
    return frozenset(self._claimed_ids(self._load_job_meta(job_id), slug, job_id))

get_job_plan(job_id: str) -> list[dict[str, Any]] | None

Return the page-plan list, or None if no plan has been committed yet.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
def get_job_plan(self, job_id: str) -> list[dict[str, Any]] | None:
    """Return the page-plan list, or None if no plan has been committed yet."""
    path = self._job_plan_path(job_id)
    if not path.exists():
        return None
    try:
        data = json.loads(path.read_text(encoding="utf-8"))
        return data if isinstance(data, list) else None
    except Exception:
        return None

get_job_session(job_id: str) -> str | None

Return the session_id attached to job_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1611
1612
1613
1614
1615
1616
def get_job_session(self, job_id: str) -> str | None:
    """Return the session_id attached to *job_id*, or None."""
    path = self._session_path(job_id)
    if not path.exists():
        return None
    return path.read_text(encoding="utf-8").strip() or None

get_job_submission(job_id: str) -> dict[str, Any] | None

Return the persisted submission dict, or None if not yet saved.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
def get_job_submission(self, job_id: str) -> dict[str, Any] | None:
    """Return the persisted submission dict, or None if not yet saved."""
    path = self._job_submission_path(job_id)
    if not path.exists():
        return None
    try:
        data = json.loads(path.read_text(encoding="utf-8"))
        return data if isinstance(data, dict) else None
    except Exception:
        return None

get_job_submitted_count(job_id: str) -> int

Return the number of pages submitted so far for job_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1708
1709
1710
1711
def get_job_submitted_count(self, job_id: str) -> int:
    """Return the number of pages submitted so far for *job_id*."""
    meta = self._load_job_meta(job_id)
    return int(meta.get("submitted_pages", 0))

get_memory_node(slug: str, node_id: str) -> MemoryNode | None

Return a single memory node, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2341
2342
2343
2344
2345
2346
def get_memory_node(self, slug: str, node_id: str) -> MemoryNode | None:
    """Return a single memory node, or None if absent."""
    for n in self._load_jsonl(self._memory_nodes_path(slug), MemoryNode):
        if n.node_id == node_id:
            return n
    return None

get_project(slug: str) -> Project | None

Return the project for slug, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1260
1261
1262
def get_project(self, slug: str) -> Project | None:
    """Return the project for *slug*, or None if absent."""
    return self._load_json(self._project_path(slug), Project)

get_project_settings(slug: str) -> ProjectSettings | None

Return slug's settings record, or None when never written.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1293
1294
1295
def get_project_settings(self, slug: str) -> ProjectSettings | None:
    """Return *slug*'s settings record, or None when never written."""
    return self._load_json(self._settings_path(slug), ProjectSettings)

get_qa(answer_id: str) -> QaAnswer | None

Return the QA answer, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1909
1910
1911
def get_qa(self, answer_id: str) -> QaAnswer | None:
    """Return the QA answer, or None if absent."""
    return self._load_json(self._qa_path(answer_id), QaAnswer)

get_qa_session(answer_id: str) -> str | None

Return the session_id attached to answer_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1940
1941
1942
1943
1944
1945
def get_qa_session(self, answer_id: str) -> str | None:
    """Return the session_id attached to *answer_id*, or None."""
    path = self._qa_session_path(answer_id)
    if not path.exists():
        return None
    return path.read_text(encoding="utf-8").strip() or None

get_recovery_attempts(slug: str) -> int

Return the recovery-attempt count for slug (0 if never recovered).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
def get_recovery_attempts(self, slug: str) -> int:
    """Return the recovery-attempt count for *slug* (0 if never recovered)."""
    path = self._recovery_path(slug)
    if not path.exists():
        return 0
    try:
        data = json.loads(path.read_text(encoding="utf-8"))
        return int(data.get("attempts", 0)) if isinstance(data, dict) else 0
    except Exception:
        return 0

get_resume_plan(job_id: str) -> dict[str, Any] | None

Return the persisted resume-plan dict, or None if the job isn't resuming.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
def get_resume_plan(self, job_id: str) -> dict[str, Any] | None:
    """Return the persisted resume-plan dict, or None if the job isn't resuming."""
    path = self._job_resume_path(job_id)
    if not path.exists():
        return None
    try:
        data = json.loads(path.read_text(encoding="utf-8"))
        return data if isinstance(data, dict) else None
    except Exception:
        return None

list_credentials() -> dict[str, dict[str, Any]]

Return every stored credential blob keyed by scope.

The scope is read from the blob's scope field (stamped by CredentialStore.save) — authoritative and lossless. Only a blob missing that field falls back to inverting :func:_slug_to_path (__ → /), which corrupts a scope containing a literal __; malformed files are skipped. Read-only: never creates the dir.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
def list_credentials(self) -> dict[str, dict[str, Any]]:
    """Return every stored credential blob keyed by scope.

    The scope is read from the blob's ``scope`` field (stamped by
    ``CredentialStore.save``) — authoritative and lossless. Only a blob
    missing that field falls back to inverting :func:`_slug_to_path`
    (``__`` → ``/``), which corrupts a scope containing a literal ``__``;
    malformed files are skipped. Read-only: never creates the dir.
    """
    out: dict[str, dict[str, Any]] = {}
    cred_dir = self.root_dir / "credentials"
    if not cred_dir.exists():
        return out
    for path in cred_dir.glob("*.json"):
        try:
            data = json.loads(path.read_text(encoding="utf-8"))
        except Exception:
            logging.warning("Skipping malformed credential file {}", path)
            continue
        if isinstance(data, dict):
            blob_scope = data.get("scope")
            scope = blob_scope if isinstance(blob_scope, str) else path.stem.replace("__", "/")
            out[scope] = data
    return out

list_doc_notes(slug: str) -> list[DocPageNote]

Return every doc-page note for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2479
2480
2481
def list_doc_notes(self, slug: str) -> list[DocPageNote]:
    """Return every doc-page note for *slug*."""
    return self._load_jsonl(self._doc_notes_path(slug), DocPageNote)

list_edges(slug: str, *, scope: CommitScope) -> list[GraphEdge]

See WikiStoreBase.list_edges.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2148
2149
2150
2151
2152
2153
2154
def list_edges(self, slug: str, *, scope: CommitScope) -> list[GraphEdge]:
    """See ``WikiStoreBase.list_edges``."""
    return [
        e
        for e in self._load_jsonl(self._edges_path(slug), GraphEdge)
        if scope.matches(e.commit_sha)
    ]

list_entity_edges(slug: str, *, source_id: str | None = None) -> list[EntityRelation]

Return entity relations, optionally scoped to source_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2637
2638
2639
2640
2641
2642
2643
2644
def list_entity_edges(
    self, slug: str, *, source_id: str | None = None
) -> list[EntityRelation]:
    """Return entity relations, optionally scoped to ``source_id``."""
    out = self._load_jsonl(self._entity_edges_path(slug), EntityRelation)
    if source_id is not None:
        out = [e for e in out if e.source_id == source_id]
    return out

list_file_manifest(slug: str) -> list[FileManifest]

Return every file-manifest entry for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2515
2516
2517
def list_file_manifest(self, slug: str) -> list[FileManifest]:
    """Return every file-manifest entry for *slug*."""
    return self._load_jsonl(self._manifest_path(slug), FileManifest)

list_jobs(slug: str | None = None) -> list[IndexingJob]

Return all jobs, newest first by phase_started_at, filtered to slug.

iterdir() yields entries in arbitrary, platform-dependent filesystem order, so the list MUST be sorted before return, and NOT by job_id — a uuid4 hex sorts RANDOMLY with respect to when a job actually ran, which is not merely unhelpful but ACTIVELY MISLEADING: at least one caller (resolve_qa_clone_dir) assumes this method returns most-recent-first and picks the FIRST complete hit, which under such a sort could be an arbitrary old checkout on any slug with 2+ completed jobs. phase_started_at is ISO-8601, so lexicographic == chronological; a job that never emitted a phase (no timestamp) sorts last, never winning over one that has. :meth:latest_job is a thin convenience over this order (narrow by status, take the first).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
def list_jobs(self, slug: str | None = None) -> list[IndexingJob]:
    """Return all jobs, newest first by ``phase_started_at``, filtered to *slug*.

    ``iterdir()`` yields entries in arbitrary, platform-dependent filesystem
    order, so the list MUST be sorted before return, and NOT by ``job_id`` —
    a ``uuid4`` hex sorts RANDOMLY with respect to when a job actually ran,
    which is not merely unhelpful but ACTIVELY MISLEADING: at least one
    caller (``resolve_qa_clone_dir``) assumes this method returns
    most-recent-first and picks the FIRST ``complete`` hit, which under such
    a sort could be an arbitrary old checkout on any slug with 2+ completed
    jobs. ``phase_started_at`` is ISO-8601, so lexicographic ==
    chronological; a job that never emitted a phase (no timestamp) sorts
    last, never winning over one that has. :meth:`latest_job` is a thin
    convenience over this order (narrow by status, take the first).
    """
    jobs_root = self.root_dir / "jobs"
    if not jobs_root.exists():
        return []
    jobs: list[IndexingJob] = []
    for job_dir in jobs_root.iterdir():
        if not job_dir.is_dir():
            continue
        job = self._load_json(job_dir / "job.json", IndexingJob)
        if job is None:
            continue
        if slug is None or job.slug == slug:
            jobs.append(job)
    return sorted(jobs, key=lambda j: j.phase_started_at or "", reverse=True)

list_memory_edges(slug: str, *, node_id: str | None = None, include_invalidated: bool = False) -> list[MemoryEdge]

Return memory edges, optionally scoped to source == node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2384
2385
2386
2387
2388
2389
2390
2391
2392
2393
2394
2395
2396
2397
2398
2399
def list_memory_edges(
    self,
    slug: str,
    *,
    node_id: str | None = None,
    include_invalidated: bool = False,
) -> list[MemoryEdge]:
    """Return memory edges, optionally scoped to ``source == node_id``."""
    out: list[MemoryEdge] = []
    for e in self._load_jsonl(self._memory_edges_path(slug), MemoryEdge):
        if node_id is not None and e.source != node_id:
            continue
        if e.invalid_at is not None and not include_invalidated:
            continue
        out.append(e)
    return out

list_pages(slug: str) -> list[WikiPage]

Return all pages for project slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
def list_pages(self, slug: str) -> list[WikiPage]:
    """Return all pages for project *slug*."""
    pages_dir = self._pages_dir(slug)
    if not pages_dir.exists():
        return []
    pages: list[WikiPage] = []
    for p in pages_dir.glob("*.json"):
        if p.name in ("_index.json", "_attribution.json"):
            continue
        page = self._load_json(p, WikiPage)
        if page is not None:
            pages.append(page)
    return pages

list_projects() -> list[Project]

Return all projects sorted by indexed_at descending.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1264
1265
1266
1267
1268
1269
1270
1271
def list_projects(self) -> list[Project]:
    """Return all projects sorted by indexed_at descending."""
    projects: list[Project] = []
    for p in (self.root_dir / "projects").glob("*.json"):
        proj = self._load_json(p, Project)
        if proj is not None:
            projects.append(proj)
    return sorted(projects, key=lambda pr: pr.indexed_at, reverse=True)

list_qa(status: str | None = None) -> list[QaAnswer]

Return all QA answers, optionally filtered to status.

iterdir() order is arbitrary — this is a boot-time/offline scan (mirrors :meth:list_jobs), never an interactive listing, so no ordering guarantee is made or needed here.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
def list_qa(self, status: str | None = None) -> list[QaAnswer]:
    """Return all QA answers, optionally filtered to *status*.

    ``iterdir()`` order is arbitrary — this is a boot-time/offline scan
    (mirrors :meth:`list_jobs`), never an interactive listing, so no
    ordering guarantee is made or needed here.
    """
    qa_root = self.root_dir / "qa"
    if not qa_root.exists():
        return []
    answers: list[QaAnswer] = []
    for qa_dir in qa_root.iterdir():
        if not qa_dir.is_dir():
            continue
        answer = self._load_json(qa_dir / "answer.json", QaAnswer)
        if answer is None:
            continue
        if status is None or answer.status == status:
            answers.append(answer)
    return answers

load_job_events(job_id: str, after_idx: int = -1) -> list[dict[str, Any]]

Return job events with idx > after_idx (-1 returns all).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1588
1589
1590
1591
1592
def load_job_events(
    self, job_id: str, after_idx: int = -1
) -> list[dict[str, Any]]:
    """Return job events with idx > *after_idx* (-1 returns all)."""
    return self._load_events("jobs", job_id, after_idx)

load_qa_events(answer_id: str, after_idx: int = -1) -> list[dict[str, Any]]

Return QA events with idx > after_idx (-1 returns all).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1964
1965
1966
1967
1968
def load_qa_events(
    self, answer_id: str, after_idx: int = -1
) -> list[dict[str, Any]]:
    """Return QA events with idx > *after_idx* (-1 returns all)."""
    return self._load_events("qa", answer_id, after_idx)

memories_anchored_to(slug: str, entity_keys: Iterable[EntityKey], *, include_invalidated: bool = False) -> list[str]

Reverse ANCHORS lookup: entity_keys → distinct memory node_ids.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2401
2402
2403
2404
2405
2406
2407
2408
2409
2410
2411
2412
2413
2414
2415
2416
2417
2418
2419
2420
def memories_anchored_to(
    self,
    slug: str,
    entity_keys: Iterable[EntityKey],
    *,
    include_invalidated: bool = False,
) -> list[str]:
    """Reverse ANCHORS lookup: entity_keys → distinct memory node_ids."""
    keys = set(entity_keys)
    seen: list[str] = []
    seen_set: set[str] = set()
    for e in self._load_jsonl(self._memory_edges_path(slug), MemoryEdge):
        if e.type != "ANCHORS" or e.target not in keys:
            continue
        if e.invalid_at is not None and not include_invalidated:
            continue
        if e.source not in seen_set:
            seen_set.add(e.source)
            seen.append(e.source)
    return seen

Top-k memory embeddings by cosine, after applying filt.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2447
2448
2449
2450
2451
2452
2453
2454
2455
2456
2457
def memory_vector_search(
    self,
    slug: str,
    qvec: list[float],
    k: int = 10,
    *,
    filt: MemoryFilter | None = None,
) -> list[MemoryEmbedding]:
    """Top-k memory embeddings by cosine, after applying *filt*."""
    pool = self._load_jsonl(self._memory_embeddings_path(slug), MemoryEmbedding)
    return self._rank_memory(slug, pool, qvec, k, filt)

page_ids_for_job(slug: str, job_id: str) -> frozenset[str]

Page ids whose attribution sidecar entry names job_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1749
1750
1751
1752
1753
1754
1755
def page_ids_for_job(self, slug: str, job_id: str) -> frozenset[str]:
    """Page ids whose attribution sidecar entry names *job_id*."""
    return frozenset(
        pid
        for pid, meta in self._load_attribution(slug).items()
        if isinstance(meta, dict) and meta.get("job_id") == job_id
    )

query_entities(slug: str, *, filt: EntityFilter | None = None) -> list[Entity]

Return entities matching filt's facets.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2571
2572
2573
2574
2575
2576
2577
2578
def query_entities(
    self, slug: str, *, filt: EntityFilter | None = None
) -> list[Entity]:
    """Return entities matching *filt*'s facets."""
    entities = self._load_jsonl(self._entities_path(slug), Entity)
    if filt is None:
        return entities
    return [e for e in entities if filt.matches(e)]

query_graph(slug: str, *, scope: CommitScope, node_type: str | None = None, name_match: str | None = None, neighbors_of: str | None = None, node_ids: Collection[str] | None = None) -> list[GraphNode]

See WikiStoreBase.query_graph.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
2115
2116
2117
2118
2119
2120
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
2143
2144
2145
2146
def query_graph(
    self,
    slug: str,
    *,
    scope: CommitScope,
    node_type: str | None = None,
    name_match: str | None = None,
    neighbors_of: str | None = None,
    node_ids: Collection[str] | None = None,
) -> list[GraphNode]:
    """See ``WikiStoreBase.query_graph``."""
    if neighbors_of is not None:
        # The seed's own generation bounds the walk: an edge is only a
        # neighbour relation if it belongs to the same generation as the
        # nodes we are about to return, or the neighbourhood spans commits.
        edges = [
            e
            for e in self._load_jsonl(self._edges_path(slug), GraphEdge)
            if scope.matches(e.commit_sha)
        ]
        related_ids: set[str] = set()
        for edge in edges:
            if edge.source == neighbors_of:
                related_ids.add(edge.target)
            elif edge.target == neighbors_of:
                related_ids.add(edge.source)
        all_nodes = self._load_graph_nodes(self._nodes_path(slug))
        return [
            n
            for n in all_nodes
            if n.node_id in related_ids and scope.matches(n.commit_sha)
        ]
    nodes = [
        n
        for n in self._load_graph_nodes(self._nodes_path(slug))
        if scope.matches(n.commit_sha)
    ]
    if node_ids is not None:
        wanted = set(node_ids)
        nodes = [n for n in nodes if n.node_id in wanted]
    if node_type is not None:
        nodes = [n for n in nodes if n.type == node_type]
    if name_match is not None:
        lower = name_match.lower()
        nodes = [n for n in nodes if lower in n.name.lower()]
    return nodes

query_memory(slug: str, *, filt: MemoryFilter | None = None) -> list[MemoryNode]

Return memory nodes matching filt's node-level facets.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2364
2365
2366
2367
2368
2369
2370
2371
def query_memory(
    self, slug: str, *, filt: MemoryFilter | None = None
) -> list[MemoryNode]:
    """Return memory nodes matching *filt*'s node-level facets."""
    nodes = self._load_jsonl(self._memory_nodes_path(slug), MemoryNode)
    if filt is None:
        return nodes
    return [n for n in nodes if filt.matches_node(n)]

reap_slug(slug: str) -> dict[str, int]

See WikiStoreBase.reap_slug.

Counts are taken from each JSONL/file BEFORE its owning directory is removed, so a family's number here means the same thing as the Mongo driver's delete_many().deleted_count — one count per row, not "the directory existed". Graph/entity/memory each live under ONE directory (_graph_dir/_memory_dir), and pages under another (_pages_dir), so counting + removing those collapses into directory operations rather than per-file bookkeeping. Jobs and QA are id-keyed, not slug-keyed: their ids are gathered FIRST, before anything is deleted, then each owning directory is removed by id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
def reap_slug(self, slug: str) -> dict[str, int]:
    """See ``WikiStoreBase.reap_slug``.

    Counts are taken from each JSONL/file BEFORE its owning directory is
    removed, so a family's number here means the same thing as the Mongo
    driver's ``delete_many().deleted_count`` — one count per row, not "the
    directory existed". Graph/entity/memory each live under ONE directory
    (``_graph_dir``/``_memory_dir``), and pages under another
    (``_pages_dir``), so counting + removing those collapses into directory
    operations rather than per-file bookkeeping. Jobs and QA are id-keyed,
    not slug-keyed: their ids are gathered FIRST, before anything is
    deleted, then each owning directory is removed by id.
    """
    with self._lock:
        job_ids = [j.job_id for j in self.list_jobs(slug)]
        answer_ids = self._qa_answer_ids_for_slug(slug)

        graph_dir = self._graph_dir(slug)
        counts: dict[str, int] = {
            "graph_nodes": self._count_jsonl_lines(graph_dir / "nodes.jsonl"),
            "graph_edges": self._count_jsonl_lines(graph_dir / "edges.jsonl"),
            "embeddings": self._count_jsonl_lines(graph_dir / "embeddings.jsonl"),
        }
        self._rmdir_if_exists(graph_dir)

        mem_dir = self._memory_dir(slug)
        counts.update({
            "entities": self._count_jsonl_lines(mem_dir / "entities.jsonl"),
            "entity_edges": self._count_jsonl_lines(mem_dir / "entity_edges.jsonl"),
            "entity_embeddings": self._count_jsonl_lines(
                mem_dir / "entity_embeddings.jsonl"
            ),
            "entity_recommendations": self._count_jsonl_lines(
                mem_dir / "entity_recommendations.jsonl"
            ),
            "memory_nodes": self._count_jsonl_lines(mem_dir / "nodes.jsonl"),
            "memory_edges": self._count_jsonl_lines(mem_dir / "edges.jsonl"),
            "memory_embeddings": self._count_jsonl_lines(mem_dir / "embeddings.jsonl"),
            "doc_notes": self._count_jsonl_lines(mem_dir / "docs.jsonl"),
            "file_manifest": self._count_jsonl_lines(mem_dir / "manifest.jsonl"),
        })
        self._rmdir_if_exists(mem_dir)

        pages_dir = self._pages_dir(slug)
        counts["pages"] = (
            sum(
                1
                for p in pages_dir.glob("*.json")
                if p.name not in ("_index.json", "_attribution.json")
            )
            if pages_dir.exists()
            else 0
        )
        self._rmdir_if_exists(pages_dir)

        recovery_path = self._recovery_path(slug)
        counts["recovery"] = 1 if recovery_path.exists() else 0
        if counts["recovery"]:
            recovery_path.unlink()

        counts["jobs"] = len(job_ids)
        counts["job_events"] = sum(
            self._count_jsonl_lines(self._event_path("jobs", jid)) for jid in job_ids
        )
        for job_id in job_ids:
            self._rmdir_if_exists(self._job_dir(job_id))

        counts["qa"] = len(answer_ids)
        counts["qa_events"] = sum(
            self._count_jsonl_lines(self._event_path("qa", aid)) for aid in answer_ids
        )
        for answer_id in answer_ids:
            self._rmdir_if_exists(self._qa_dir(answer_id))

    return counts

reset_recovery_attempts(slug: str) -> None

Clear slug's recovery counter file (user-initiated resume fresh budget).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1875
1876
1877
1878
1879
1880
def reset_recovery_attempts(self, slug: str) -> None:
    """Clear *slug*'s recovery counter file (user-initiated resume fresh budget)."""
    with self._lock:
        path = self._recovery_path(slug)
        if path.exists():
            path.unlink()

restamp_graph_artifacts(slug: str, *, from_commit: str, to_commit: str) -> dict[str, int]

See WikiStoreBase.restamp_graph_artifacts.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
def restamp_graph_artifacts(
    self, slug: str, *, from_commit: str, to_commit: str
) -> dict[str, int]:
    """See ``WikiStoreBase.restamp_graph_artifacts``."""
    counts: dict[str, int] = {}
    with self._lock:
        for key, path, loader in (
            ("nodes", self._nodes_path(slug), self._load_graph_nodes),
            ("edges", self._edges_path(slug),
             lambda p: self._load_jsonl(p, GraphEdge)),
            ("embeddings", self._embeddings_path(slug),
             lambda p: self._load_jsonl(p, Embedding)),
            ("entities", self._entities_path(slug),
             lambda p: self._load_jsonl(p, Entity)),
            ("entity_edges", self._entity_edges_path(slug),
             lambda p: self._load_jsonl(p, EntityRelation)),
            ("entity_embeddings", self._entity_embeddings_path(slug),
             lambda p: self._load_jsonl(p, EntityEmbedding)),
        ):
            counts[key] = self._restamp_jsonl(path, loader, from_commit, to_commit)
    return counts

save_act_plan(job_id: str, plan: dict[str, Any]) -> None

Persist the act-stage record for job_id; overwrites any previous one.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1691
1692
1693
1694
1695
def save_act_plan(self, job_id: str, plan: dict[str, Any]) -> None:
    """Persist the act-stage record for *job_id*; overwrites any previous one."""
    path = self._job_act_path(job_id)
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(json.dumps(plan, indent=2), encoding="utf-8")

save_credentials(slug: str, blob: dict[str, Any]) -> None

Persist the encoded credential blob for slug at mode 0600.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1794
1795
1796
1797
1798
1799
1800
1801
def save_credentials(self, slug: str, blob: dict[str, Any]) -> None:
    """Persist the encoded credential *blob* for *slug* at mode 0600."""
    path = self._credential_path(slug)
    path.write_text(json.dumps(blob, indent=2), encoding="utf-8")
    try:
        path.chmod(0o600)
    except OSError:  # pragma: no cover
        pass

save_entity_recommendation(slug: str, rec: EntityRecommendation) -> None

Upsert a recommendation for slug; dedup by id (a replay converges).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2646
2647
2648
2649
2650
2651
2652
2653
2654
2655
2656
2657
2658
def save_entity_recommendation(
    self, slug: str, rec: EntityRecommendation
) -> None:
    """Upsert a recommendation for *slug*; dedup by id (a replay converges)."""
    with self._lock:
        existing = {
            r.id: r
            for r in self._load_jsonl(
                self._entity_recs_path(slug), EntityRecommendation
            )
        }
        existing[rec.id] = rec
        self._write_jsonl(self._entity_recs_path(slug), list(existing.values()))

save_job_plan(job_id: str, plan: list[dict[str, Any]]) -> None

Persist the page-plan list for job_id; overwrites any previous plan.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1649
1650
1651
1652
1653
def save_job_plan(self, job_id: str, plan: list[dict[str, Any]]) -> None:
    """Persist the page-plan list for *job_id*; overwrites any previous plan."""
    path = self._job_plan_path(job_id)
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(json.dumps(plan, indent=2), encoding="utf-8")

save_job_submission(job_id: str, submission: dict[str, Any]) -> None

Persist the wizard submission dict for job_id (token must be absent).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1761
1762
1763
1764
1765
def save_job_submission(self, job_id: str, submission: dict[str, Any]) -> None:
    """Persist the wizard submission dict for *job_id* (token must be absent)."""
    path = self._job_submission_path(job_id)
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(json.dumps(submission, indent=2), encoding="utf-8")

save_page(slug: str, page: WikiPage, *, commit_sha: str | None = None, job_id: str | None = None) -> None

Persist page for the project slug; overwrites if same page_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
def save_page(
    self,
    slug: str,
    page: WikiPage,
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Persist *page* for the project *slug*; overwrites if same page_id."""
    pages_dir = self._pages_dir(slug)
    pages_dir.mkdir(parents=True, exist_ok=True)
    self._save_json(self._page_path(slug, page.id), page)
    index = self._load_index(slug)
    index[page.id] = page.title
    self._index_path(slug).write_text(json.dumps(index, indent=2), encoding="utf-8")
    attribution = self._load_attribution(slug)
    attribution[page.id] = {"commit_sha": commit_sha, "job_id": job_id}
    self._attribution_path(slug).write_text(
        json.dumps(attribution, indent=2), encoding="utf-8"
    )

save_project_settings(slug: str, settings: ProjectSettings) -> None

Persist (upsert) the editable settings record for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1289
1290
1291
def save_project_settings(self, slug: str, settings: ProjectSettings) -> None:
    """Persist (upsert) the editable settings record for *slug*."""
    self._save_json(self._settings_path(slug), settings)

save_qa(answer: QaAnswer) -> None

Persist a QA answer record (slug round-trips through answer.json).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1896
1897
1898
1899
def save_qa(self, answer: QaAnswer) -> None:
    """Persist a QA answer record (``slug`` round-trips through answer.json)."""
    self._qa_dir(answer.answer_id).mkdir(parents=True, exist_ok=True)
    self._save_json(self._qa_path(answer.answer_id), answer)

save_resume_plan(job_id: str, plan: dict[str, Any]) -> None

Persist the resume-plan dict for job_id; overwrites any previous one.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1670
1671
1672
1673
1674
def save_resume_plan(self, job_id: str, plan: dict[str, Any]) -> None:
    """Persist the resume-plan dict for *job_id*; overwrites any previous one."""
    path = self._job_resume_path(job_id)
    path.parent.mkdir(parents=True, exist_ok=True)
    path.write_text(json.dumps(plan, indent=2), encoding="utf-8")

supersede_graph_artifacts(slug: str, *, keep_commit_sha: str) -> dict[str, int]

Drop prior-commit graph + entity artifacts, preserving None-stamped rows.

A row survives iff its commit_sha is None (a QA-minted or pre-isolation record) OR equals keep_commit_sha. Everything else — the artifacts of a superseded commit — is dropped.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
2178
2179
2180
2181
2182
2183
2184
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
def supersede_graph_artifacts(
    self, slug: str, *, keep_commit_sha: str
) -> dict[str, int]:
    """Drop prior-commit graph + entity artifacts, preserving ``None``-stamped rows.

    A row survives iff its ``commit_sha`` is ``None`` (a QA-minted or
    pre-isolation record) OR equals *keep_commit_sha*. Everything else — the
    artifacts of a superseded commit — is dropped.
    """
    counts: dict[str, int] = {}
    with self._lock:
        counts["nodes"] = self._retain_jsonl(
            self._nodes_path(slug), self._load_graph_nodes, keep_commit_sha
        )
        counts["edges"] = self._retain_jsonl(
            self._edges_path(slug),
            lambda p: self._load_jsonl(p, GraphEdge),
            keep_commit_sha,
        )
        counts["embeddings"] = self._retain_jsonl(
            self._embeddings_path(slug),
            lambda p: self._load_jsonl(p, Embedding),
            keep_commit_sha,
        )
        counts["entities"] = self._retain_jsonl(
            self._entities_path(slug),
            lambda p: self._load_jsonl(p, Entity),
            keep_commit_sha,
        )
        counts["entity_edges"] = self._retain_jsonl(
            self._entity_edges_path(slug),
            lambda p: self._load_jsonl(p, EntityRelation),
            keep_commit_sha,
        )
        counts["entity_embeddings"] = self._retain_jsonl(
            self._entity_embeddings_path(slug),
            lambda p: self._load_jsonl(p, EntityEmbedding),
            keep_commit_sha,
        )
    return counts

update_qa_fields(answer: QaAnswer) -> None

Non-destructive field update.

Session + events are separate files here, so a plain answer.json rewrite already preserves them.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1901
1902
1903
1904
1905
1906
1907
def update_qa_fields(self, answer: QaAnswer) -> None:
    """Non-destructive field update.

    Session + events are separate files here, so a plain answer.json rewrite
    already preserves them.
    """
    self._save_json(self._qa_path(answer.answer_id), answer)

upsert_doc_notes(slug: str, notes: Iterable[DocPageNote]) -> None

Upsert doc-page notes for slug; dedup by page_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2461
2462
2463
2464
2465
2466
2467
2468
2469
2470
def upsert_doc_notes(self, slug: str, notes: Iterable[DocPageNote]) -> None:
    """Upsert doc-page notes for *slug*; dedup by page_id."""
    with self._lock:
        existing = {
            d.page_id: d
            for d in self._load_jsonl(self._doc_notes_path(slug), DocPageNote)
        }
        for note in notes:
            existing[note.page_id] = note
        self._write_jsonl(self._doc_notes_path(slug), list(existing.values()))

upsert_edges(slug: str, edges: Iterable[GraphEdge], *, commit_sha: str | None = None, job_id: str | None = None, on_progress: Callable[[int, int], None] | None = None) -> None

Upsert graph edges for slug; dedup by (source, target, type).

Cost: O(edges) — offline. The JSON driver writes one atomic file, so one non-empty input is one persistence batch.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
def upsert_edges(
    self,
    slug: str,
    edges: Iterable[GraphEdge],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
    on_progress: Callable[[int, int], None] | None = None,
) -> None:
    """Upsert graph edges for *slug*; dedup by (source, target, type).

    Cost: ``O(edges)`` — offline. The JSON driver writes one atomic file, so
    one non-empty input is one persistence batch.
    """
    items = list(edges)
    with self._lock:
        existing = {
            (e.source, e.target, e.type): e
            for e in self._load_jsonl(self._edges_path(slug), GraphEdge)
        }
        for edge in items:
            existing[(edge.source, edge.target, edge.type)] = self._stamp_attribution(
                edge, commit_sha, job_id
            )
        self._write_jsonl(self._edges_path(slug), list(existing.values()))
    if items and on_progress is not None:
        on_progress(1, 1)

upsert_embeddings(slug: str, items: Iterable[Embedding], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert embedding vectors for slug; dedup by node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
def upsert_embeddings(
    self,
    slug: str,
    items: Iterable[Embedding],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert embedding vectors for *slug*; dedup by node_id."""
    with self._lock:
        existing = {
            e.node_id: e
            for e in self._load_jsonl(self._embeddings_path(slug), Embedding)
        }
        for item in items:
            existing[item.node_id] = self._stamp_attribution(item, commit_sha, job_id)
        self._write_jsonl(self._embeddings_path(slug), list(existing.values()))

upsert_entities(slug: str, entities: Iterable[Entity], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert entities for slug; dedup by id, stamp attribution.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2547
2548
2549
2550
2551
2552
2553
2554
2555
2556
2557
2558
2559
2560
2561
2562
def upsert_entities(
    self,
    slug: str,
    entities: Iterable[Entity],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert entities for *slug*; dedup by id, stamp attribution."""
    with self._lock:
        existing = {
            e.id: e for e in self._load_jsonl(self._entities_path(slug), Entity)
        }
        for entity in entities:
            existing[entity.id] = self._stamp_attribution(entity, commit_sha, job_id)
        self._write_jsonl(self._entities_path(slug), list(existing.values()))

upsert_entity_edges(slug: str, edges: Iterable[EntityRelation], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert entity relations for slug; dedup by id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2619
2620
2621
2622
2623
2624
2625
2626
2627
2628
2629
2630
2631
2632
2633
2634
2635
def upsert_entity_edges(
    self,
    slug: str,
    edges: Iterable[EntityRelation],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert entity relations for *slug*; dedup by id."""
    with self._lock:
        existing = {
            e.id: e
            for e in self._load_jsonl(self._entity_edges_path(slug), EntityRelation)
        }
        for edge in edges:
            existing[edge.id] = self._stamp_attribution(edge, commit_sha, job_id)
        self._write_jsonl(self._entity_edges_path(slug), list(existing.values()))

upsert_entity_embeddings(slug: str, items: Iterable[EntityEmbedding], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert entity embedding vectors for slug; dedup by entity_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2588
2589
2590
2591
2592
2593
2594
2595
2596
2597
2598
2599
2600
2601
2602
2603
2604
2605
2606
2607
2608
2609
2610
def upsert_entity_embeddings(
    self,
    slug: str,
    items: Iterable[EntityEmbedding],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert entity embedding vectors for *slug*; dedup by entity_id."""
    with self._lock:
        existing = {
            e.entity_id: e
            for e in self._load_jsonl(
                self._entity_embeddings_path(slug), EntityEmbedding
            )
        }
        for item in items:
            existing[item.entity_id] = self._stamp_attribution(
                item, commit_sha, job_id
            )
        self._write_jsonl(
            self._entity_embeddings_path(slug), list(existing.values())
        )

upsert_file_manifest(slug: str, entries: Iterable[FileManifest]) -> None

Upsert file-manifest entries for slug; dedup by path.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2495
2496
2497
2498
2499
2500
2501
2502
2503
2504
2505
2506
def upsert_file_manifest(
    self, slug: str, entries: Iterable[FileManifest]
) -> None:
    """Upsert file-manifest entries for *slug*; dedup by path."""
    with self._lock:
        existing = {
            m.path: m
            for m in self._load_jsonl(self._manifest_path(slug), FileManifest)
        }
        for entry in entries:
            existing[entry.path] = entry
        self._write_jsonl(self._manifest_path(slug), list(existing.values()))

upsert_memory_edges(slug: str, edges: Iterable[MemoryEdge]) -> None

Upsert memory edges for slug; dedup by (source, target, type).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2373
2374
2375
2376
2377
2378
2379
2380
2381
2382
def upsert_memory_edges(self, slug: str, edges: Iterable[MemoryEdge]) -> None:
    """Upsert memory edges for *slug*; dedup by (source, target, type)."""
    with self._lock:
        existing = {
            (e.source, e.target, e.type): e
            for e in self._load_jsonl(self._memory_edges_path(slug), MemoryEdge)
        }
        for edge in edges:
            existing[(edge.source, edge.target, edge.type)] = edge
        self._write_jsonl(self._memory_edges_path(slug), list(existing.values()))

upsert_memory_embeddings(slug: str, items: Iterable[MemoryEmbedding]) -> None

Upsert memory embedding vectors for slug; dedup by node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2430
2431
2432
2433
2434
2435
2436
2437
2438
2439
2440
2441
2442
2443
2444
2445
def upsert_memory_embeddings(
    self, slug: str, items: Iterable[MemoryEmbedding]
) -> None:
    """Upsert memory embedding vectors for *slug*; dedup by node_id."""
    with self._lock:
        existing = {
            e.node_id: e
            for e in self._load_jsonl(
                self._memory_embeddings_path(slug), MemoryEmbedding
            )
        }
        for item in items:
            existing[item.node_id] = item
        self._write_jsonl(
            self._memory_embeddings_path(slug), list(existing.values())
        )

upsert_memory_nodes(slug: str, nodes: Iterable[MemoryNode]) -> None

Upsert memory nodes for slug; dedup by node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2330
2331
2332
2333
2334
2335
2336
2337
2338
2339
def upsert_memory_nodes(self, slug: str, nodes: Iterable[MemoryNode]) -> None:
    """Upsert memory nodes for *slug*; dedup by node_id."""
    with self._lock:
        existing = {
            n.node_id: n
            for n in self._load_jsonl(self._memory_nodes_path(slug), MemoryNode)
        }
        for node in nodes:
            existing[node.node_id] = node
        self._write_jsonl(self._memory_nodes_path(slug), list(existing.values()))

upsert_nodes(slug: str, nodes: Iterable[GraphNode], *, commit_sha: str | None = None, job_id: str | None = None, on_progress: Callable[[int, int], None] | None = None) -> None

Upsert graph nodes for slug; dedup by node_id, stamp attribution.

Cost: O(nodes) — offline. The JSON driver writes one atomic file, so one non-empty input is one persistence batch.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
def upsert_nodes(
    self,
    slug: str,
    nodes: Iterable[GraphNode],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
    on_progress: Callable[[int, int], None] | None = None,
) -> None:
    """Upsert graph nodes for *slug*; dedup by node_id, stamp attribution.

    Cost: ``O(nodes)`` — offline. The JSON driver writes one atomic file, so
    one non-empty input is one persistence batch.
    """
    items = list(nodes)
    with self._lock:
        existing = {n.node_id: n for n in self._load_graph_nodes(self._nodes_path(slug))}
        for node in items:
            existing[node.node_id] = self._stamp_attribution(node, commit_sha, job_id)
        self._write_jsonl(self._nodes_path(slug), list(existing.values()))
    if items and on_progress is not None:
        on_progress(1, 1)

Return top-k embeddings for slug by cosine similarity.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2263
2264
2265
2266
2267
2268
2269
2270
2271
2272
def vector_search(self, slug: str, qvec: list[float], k: int = 10) -> list[Embedding]:
    """Return top-k embeddings for *slug* by cosine similarity."""
    from .embedder import Embedder

    pool = self._load_jsonl(self._embeddings_path(slug), Embedding)
    if not pool:
        return []
    scored = [(emb, Embedder.cosine(qvec, emb.vector)) for emb in pool]
    scored.sort(key=lambda t: t[1], reverse=True)
    return [emb for emb, _ in scored[:k]]

MongoWikiStore

Bases: WikiStoreBase

MongoDB-backed wiki persistence.

Collections:

  • wiki_projects (slug PK)
  • wiki_pages ((slug, page_id) compound PK)
  • wiki_jobs (job_id PK; includes event_count for atomic $inc)
  • wiki_job_events ((job_id, idx) compound; append-only)
  • wiki_qa (answer_id PK; includes event_count)
  • wiki_qa_events ((answer_id, idx) compound; append-only)

Graph/embeddings collections are not created here — the methods raise NotImplementedError inherited from WikiStoreBase.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2720
2721
2722
2723
2724
2725
2726
2727
2728
2729
2730
2731
2732
2733
2734
2735
2736
2737
2738
2739
2740
2741
2742
2743
2744
2745
2746
2747
2748
2749
2750
2751
2752
2753
2754
2755
2756
2757
2758
2759
2760
2761
2762
2763
2764
2765
2766
2767
2768
2769
2770
2771
2772
2773
2774
2775
2776
2777
2778
2779
2780
2781
2782
2783
2784
2785
2786
2787
2788
2789
2790
2791
2792
2793
2794
2795
2796
2797
2798
2799
2800
2801
2802
2803
2804
2805
2806
2807
2808
2809
2810
2811
2812
2813
2814
2815
2816
2817
2818
2819
2820
2821
2822
2823
2824
2825
2826
2827
2828
2829
2830
2831
2832
2833
2834
2835
2836
2837
2838
2839
2840
2841
2842
2843
2844
2845
2846
2847
2848
2849
2850
2851
2852
2853
2854
2855
2856
2857
2858
2859
2860
2861
2862
2863
2864
2865
2866
2867
2868
2869
2870
2871
2872
2873
2874
2875
2876
2877
2878
2879
2880
2881
2882
2883
2884
2885
2886
2887
2888
2889
2890
2891
2892
2893
2894
2895
2896
2897
2898
2899
2900
2901
2902
2903
2904
2905
2906
2907
2908
2909
2910
2911
2912
2913
2914
2915
2916
2917
2918
2919
2920
2921
2922
2923
2924
2925
2926
2927
2928
2929
2930
2931
2932
2933
2934
2935
2936
2937
2938
2939
2940
2941
2942
2943
2944
2945
2946
2947
2948
2949
2950
2951
2952
2953
2954
2955
2956
2957
2958
2959
2960
2961
2962
2963
2964
2965
2966
2967
2968
2969
2970
2971
2972
2973
2974
2975
2976
2977
2978
2979
2980
2981
2982
2983
2984
2985
2986
2987
2988
2989
2990
2991
2992
2993
2994
2995
2996
2997
2998
2999
3000
3001
3002
3003
3004
3005
3006
3007
3008
3009
3010
3011
3012
3013
3014
3015
3016
3017
3018
3019
3020
3021
3022
3023
3024
3025
3026
3027
3028
3029
3030
3031
3032
3033
3034
3035
3036
3037
3038
3039
3040
3041
3042
3043
3044
3045
3046
3047
3048
3049
3050
3051
3052
3053
3054
3055
3056
3057
3058
3059
3060
3061
3062
3063
3064
3065
3066
3067
3068
3069
3070
3071
3072
3073
3074
3075
3076
3077
3078
3079
3080
3081
3082
3083
3084
3085
3086
3087
3088
3089
3090
3091
3092
3093
3094
3095
3096
3097
3098
3099
3100
3101
3102
3103
3104
3105
3106
3107
3108
3109
3110
3111
3112
3113
3114
3115
3116
3117
3118
3119
3120
3121
3122
3123
3124
3125
3126
3127
3128
3129
3130
3131
3132
3133
3134
3135
3136
3137
3138
3139
3140
3141
3142
3143
3144
3145
3146
3147
3148
3149
3150
3151
3152
3153
3154
3155
3156
3157
3158
3159
3160
3161
3162
3163
3164
3165
3166
3167
3168
3169
3170
3171
3172
3173
3174
3175
3176
3177
3178
3179
3180
3181
3182
3183
3184
3185
3186
3187
3188
3189
3190
3191
3192
3193
3194
3195
3196
3197
3198
3199
3200
3201
3202
3203
3204
3205
3206
3207
3208
3209
3210
3211
3212
3213
3214
3215
3216
3217
3218
3219
3220
3221
3222
3223
3224
3225
3226
3227
3228
3229
3230
3231
3232
3233
3234
3235
3236
3237
3238
3239
3240
3241
3242
3243
3244
3245
3246
3247
3248
3249
3250
3251
3252
3253
3254
3255
3256
3257
3258
3259
3260
3261
3262
3263
3264
3265
3266
3267
3268
3269
3270
3271
3272
3273
3274
3275
3276
3277
3278
3279
3280
3281
3282
3283
3284
3285
3286
3287
3288
3289
3290
3291
3292
3293
3294
3295
3296
3297
3298
3299
3300
3301
3302
3303
3304
3305
3306
3307
3308
3309
3310
3311
3312
3313
3314
3315
3316
3317
3318
3319
3320
3321
3322
3323
3324
3325
3326
3327
3328
3329
3330
3331
3332
3333
3334
3335
3336
3337
3338
3339
3340
3341
3342
3343
3344
3345
3346
3347
3348
3349
3350
3351
3352
3353
3354
3355
3356
3357
3358
3359
3360
3361
3362
3363
3364
3365
3366
3367
3368
3369
3370
3371
3372
3373
3374
3375
3376
3377
3378
3379
3380
3381
3382
3383
3384
3385
3386
3387
3388
3389
3390
3391
3392
3393
3394
3395
3396
3397
3398
3399
3400
3401
3402
3403
3404
3405
3406
3407
3408
3409
3410
3411
3412
3413
3414
3415
3416
3417
3418
3419
3420
3421
3422
3423
3424
3425
3426
3427
3428
3429
3430
3431
3432
3433
3434
3435
3436
3437
3438
3439
3440
3441
3442
3443
3444
3445
3446
3447
3448
3449
3450
3451
3452
3453
3454
3455
3456
3457
3458
3459
3460
3461
3462
3463
3464
3465
3466
3467
3468
3469
3470
3471
3472
3473
3474
3475
3476
3477
3478
3479
3480
3481
3482
3483
3484
3485
3486
3487
3488
3489
3490
3491
3492
3493
3494
3495
3496
3497
3498
3499
3500
3501
3502
3503
3504
3505
3506
3507
3508
3509
3510
3511
3512
3513
3514
3515
3516
3517
3518
3519
3520
3521
3522
3523
3524
3525
3526
3527
3528
3529
3530
3531
3532
3533
3534
3535
3536
3537
3538
3539
3540
3541
3542
3543
3544
3545
3546
3547
3548
3549
3550
3551
3552
3553
3554
3555
3556
3557
3558
3559
3560
3561
3562
3563
3564
3565
3566
3567
3568
3569
3570
3571
3572
3573
3574
3575
3576
3577
3578
3579
3580
3581
3582
3583
3584
3585
3586
3587
3588
3589
3590
3591
3592
3593
3594
3595
3596
3597
3598
3599
3600
3601
3602
3603
3604
3605
3606
3607
3608
3609
3610
3611
3612
3613
3614
3615
3616
3617
3618
3619
3620
3621
3622
3623
3624
3625
3626
3627
3628
3629
3630
3631
3632
3633
3634
3635
3636
3637
3638
3639
3640
3641
3642
3643
3644
3645
3646
3647
3648
3649
3650
3651
3652
3653
3654
3655
3656
3657
3658
3659
3660
3661
3662
3663
3664
3665
3666
3667
3668
3669
3670
3671
3672
3673
3674
3675
3676
3677
3678
3679
3680
3681
3682
3683
3684
3685
3686
3687
3688
3689
3690
3691
3692
3693
3694
3695
3696
3697
3698
3699
3700
3701
3702
3703
3704
3705
3706
3707
3708
3709
3710
3711
3712
3713
3714
3715
3716
3717
3718
3719
3720
3721
3722
3723
3724
3725
3726
3727
3728
3729
3730
3731
3732
3733
3734
3735
3736
3737
3738
3739
3740
3741
3742
3743
3744
3745
3746
3747
3748
3749
3750
3751
3752
3753
3754
3755
3756
3757
3758
3759
3760
3761
3762
3763
3764
3765
3766
3767
3768
3769
3770
3771
3772
3773
3774
3775
3776
3777
3778
3779
3780
3781
3782
3783
3784
3785
3786
3787
3788
3789
3790
3791
3792
3793
3794
3795
3796
3797
3798
3799
3800
3801
3802
3803
3804
3805
3806
3807
3808
3809
3810
3811
3812
3813
3814
3815
3816
3817
3818
3819
3820
3821
3822
3823
3824
3825
3826
3827
3828
3829
3830
3831
3832
3833
3834
3835
3836
3837
3838
3839
3840
3841
3842
3843
3844
3845
3846
3847
3848
3849
3850
3851
3852
3853
3854
3855
3856
3857
3858
3859
3860
3861
3862
3863
3864
3865
3866
3867
3868
3869
3870
3871
3872
3873
3874
3875
3876
3877
3878
3879
3880
3881
3882
3883
3884
3885
3886
3887
3888
3889
3890
3891
3892
3893
3894
3895
3896
3897
3898
3899
3900
3901
3902
3903
3904
3905
3906
3907
3908
3909
3910
3911
3912
3913
3914
3915
3916
3917
3918
3919
3920
3921
3922
3923
3924
3925
3926
3927
3928
3929
3930
3931
3932
3933
3934
3935
3936
3937
3938
3939
3940
3941
3942
3943
3944
3945
3946
3947
3948
3949
3950
3951
3952
3953
3954
3955
3956
3957
3958
3959
3960
3961
3962
3963
3964
3965
3966
3967
3968
3969
3970
3971
3972
3973
3974
3975
3976
3977
3978
3979
3980
3981
3982
3983
3984
3985
3986
3987
3988
3989
3990
3991
3992
3993
3994
3995
3996
3997
3998
3999
4000
4001
4002
4003
4004
4005
4006
4007
4008
4009
4010
4011
4012
4013
4014
4015
4016
4017
4018
4019
4020
4021
4022
4023
4024
4025
4026
4027
4028
4029
4030
4031
4032
4033
4034
4035
4036
4037
4038
4039
4040
4041
4042
4043
4044
4045
4046
4047
4048
4049
4050
4051
4052
4053
4054
4055
4056
4057
4058
4059
4060
4061
4062
4063
4064
4065
4066
4067
4068
4069
4070
4071
4072
4073
4074
4075
4076
4077
4078
4079
4080
4081
4082
4083
4084
4085
4086
4087
4088
4089
4090
4091
4092
4093
4094
4095
4096
4097
4098
4099
4100
4101
4102
4103
4104
4105
4106
4107
4108
4109
4110
4111
4112
4113
4114
4115
4116
4117
4118
4119
4120
4121
4122
4123
4124
4125
4126
4127
4128
4129
4130
4131
4132
4133
4134
4135
4136
4137
4138
4139
4140
4141
4142
4143
4144
4145
4146
4147
4148
4149
4150
4151
4152
4153
4154
4155
4156
4157
4158
4159
4160
4161
4162
4163
4164
4165
4166
4167
4168
4169
4170
4171
4172
4173
4174
4175
4176
4177
4178
4179
4180
4181
4182
4183
4184
4185
4186
4187
4188
4189
4190
4191
4192
4193
4194
4195
4196
4197
4198
4199
4200
4201
4202
4203
4204
4205
4206
4207
4208
4209
4210
4211
4212
4213
4214
4215
4216
4217
4218
4219
4220
class MongoWikiStore(WikiStoreBase):
    """MongoDB-backed wiki persistence.

    Collections:

    - ``wiki_projects``     (slug PK)
    - ``wiki_pages``        ((slug, page_id) compound PK)
    - ``wiki_jobs``         (job_id PK; includes ``event_count`` for atomic ``$inc``)
    - ``wiki_job_events``   ((job_id, idx) compound; append-only)
    - ``wiki_qa``           (answer_id PK; includes ``event_count``)
    - ``wiki_qa_events``    ((answer_id, idx) compound; append-only)

    Graph/embeddings collections are not created here — the methods raise
    ``NotImplementedError`` inherited from ``WikiStoreBase``.
    """

    #: Documents per ``bulk_write`` for the graph upserts. Bounded so a write of
    #: a whole repository's edges holds one batch of pending operations in
    #: memory rather than all of them; large enough that the per-round-trip cost
    #: amortises to nothing.
    _BULK_BATCH_SIZE = 1000

    def __init__(
        self,
        *,
        client: Any = None,
        uri: str | None = None,
        database: str | None = None,
    ) -> None:
        """Initialize MongoDB connection and ensure indexes exist."""
        if client is None:
            from pymongo import MongoClient

            _uri = uri or get_config_value(
                "storage", "mongodb", "uri", default="mongodb://localhost:27017"
            )
            client = MongoClient(_uri, serverSelectionTimeoutMS=5000)
            # Fail fast — mirrors MongoSessionStore.
            client.admin.command("ping")
        if database is None:
            database = get_config_value(
                "storage", "mongodb", "database", default="mewbo"
            )
        self._client = client
        self._db = client[database]
        self._ensure_indexes()

    # -- helpers -------------------------------------------------------------

    def _col(self, name: str) -> Any:
        """Return a MongoDB collection by name."""
        return self._db[name]

    def _ensure_indexes(self) -> None:
        """Create indexes idempotently on first connection."""
        from pymongo import ASCENDING

        def _idx(col: str, keys: list[tuple[str, Any]], name: str) -> None:
            self._col(col).create_index(keys, name=name, unique=True, background=True)

        _idx("wiki_projects", [("slug", ASCENDING)], "ix_projects_slug")
        _idx(
            "wiki_pages",
            [("slug", ASCENDING), ("page_id", ASCENDING)],
            "ix_pages_slug_pageid",
        )
        _idx("wiki_jobs", [("job_id", ASCENDING)], "ix_jobs_job_id")
        _idx(
            "wiki_job_events",
            [("job_id", ASCENDING), ("idx", ASCENDING)],
            "ix_job_events_job_idx",
        )
        _idx("wiki_qa", [("answer_id", ASCENDING)], "ix_qa_answer_id")
        _idx(
            "wiki_qa_events",
            [("answer_id", ASCENDING), ("idx", ASCENDING)],
            "ix_qa_events_answer_idx",
        )
        _idx("wiki_credentials", [("slug", ASCENDING)], "ix_credentials_slug")
        _idx("wiki_recovery", [("slug", ASCENDING)], "ix_recovery_slug")
        _idx("wiki_settings", [("slug", ASCENDING)], "ix_settings_slug")

    def _atomic_next_idx(self, col: str, owner_field: str, owner_id: str) -> int:
        """Atomically increment event_count on the owner document and return the next idx (0-based).

        Uses ``$inc`` on ``event_count`` and returns ``new_value - 1`` as the
        event's monotonic idx so that the first event gets idx=0.
        """
        from pymongo import ReturnDocument

        doc = self._col(col).find_one_and_update(
            {owner_field: owner_id},
            {"$inc": {"event_count": 1}},
            return_document=ReturnDocument.AFTER,
        )
        if doc is None:
            raise KeyError(f"No document in '{col}' with {owner_field}={owner_id!r}")
        return int(doc["event_count"]) - 1

    # -- Projects ------------------------------------------------------------

    def create_project(self, project: Project) -> None:
        """Persist a new project record."""
        doc = project.model_dump(by_alias=False)
        self._col("wiki_projects").replace_one(
            {"slug": project.slug}, doc, upsert=True
        )

    def get_project(self, slug: str) -> Project | None:
        """Return the project for *slug*, or None if absent."""
        doc = self._col("wiki_projects").find_one({"slug": slug})
        if doc is None:
            return None
        return Project.model_validate(_strip_mongo_meta(doc))

    def list_projects(self) -> list[Project]:
        """Return all projects sorted by indexed_at descending."""
        cursor = self._col("wiki_projects").find().sort("indexed_at", -1)
        return [Project.model_validate(_strip_mongo_meta(d)) for d in cursor]

    def delete_project(self, slug: str) -> bool:
        """Delete project *slug*; return True if deleted, False if absent."""
        result = self._col("wiki_projects").delete_one({"slug": slug})
        return result.deleted_count > 0

    # -- Project settings (slug-keyed collection) ----------------------------

    def save_project_settings(self, slug: str, settings: ProjectSettings) -> None:
        """Persist (upsert) the editable settings record for *slug*."""
        doc = settings.model_dump(by_alias=False)
        doc["slug"] = slug
        self._col("wiki_settings").replace_one({"slug": slug}, doc, upsert=True)

    def get_project_settings(self, slug: str) -> ProjectSettings | None:
        """Return *slug*'s settings record, or None when never written."""
        doc = self._col("wiki_settings").find_one({"slug": slug})
        if doc is None:
            return None
        try:
            return ProjectSettings.model_validate(_strip_mongo_meta(doc))
        except Exception:
            # A hand-edited / off-schema document must not break the read path —
            # the caller then falls back to the per-job submission scan,
            # exactly as it does for a project that has no record at all.
            logging.warning("Skipping malformed wiki_settings document for {}", slug)
            return None

    def delete_project_settings(self, slug: str) -> bool:
        """Delete *slug*'s settings document; return True if one existed."""
        return self._col("wiki_settings").delete_one({"slug": slug}).deleted_count > 0

    # -- Whole-slug reap (see WikiStoreBase.reap_slug for the family list) ---

    def reap_slug(self, slug: str) -> dict[str, int]:
        """See ``WikiStoreBase.reap_slug``.

        ``wiki_job_events``/``wiki_qa_events`` carry only their owning id
        (``job_id``/``answer_id``), never ``slug`` — so those ids are read from
        ``wiki_jobs``/``wiki_qa`` FIRST, before either collection is touched,
        and the two event collections are then swept by id.
        """
        job_ids = [
            str(d["job_id"])
            for d in self._col("wiki_jobs").find({"slug": slug}, {"job_id": 1})
        ]
        answer_ids = [
            str(d["answer_id"])
            for d in self._col("wiki_qa").find({"slug": slug}, {"answer_id": 1})
        ]

        counts: dict[str, int] = {}
        for key, coll in (
            ("graph_nodes", "wiki_graph_nodes"),
            ("graph_edges", "wiki_graph_edges"),
            ("embeddings", "wiki_embeddings"),
            ("entities", "wiki_entities"),
            ("entity_edges", "wiki_entity_edges"),
            ("entity_embeddings", "wiki_entity_embeddings"),
            ("entity_recommendations", "wiki_entity_recommendations"),
            ("memory_nodes", "wiki_memory_nodes"),
            ("memory_edges", "wiki_memory_edges"),
            ("memory_embeddings", "wiki_memory_embeddings"),
            ("doc_notes", "wiki_doc_notes"),
            ("file_manifest", "wiki_file_manifest"),
            ("pages", "wiki_pages"),
            ("recovery", "wiki_recovery"),
            ("jobs", "wiki_jobs"),
            ("qa", "wiki_qa"),
        ):
            counts[key] = int(self._col(coll).delete_many({"slug": slug}).deleted_count)

        counts["job_events"] = (
            int(
                self._col("wiki_job_events")
                .delete_many({"job_id": {"$in": job_ids}})
                .deleted_count
            )
            if job_ids
            else 0
        )
        counts["qa_events"] = (
            int(
                self._col("wiki_qa_events")
                .delete_many({"answer_id": {"$in": answer_ids}})
                .deleted_count
            )
            if answer_ids
            else 0
        )
        return counts

    # -- Pages ---------------------------------------------------------------

    # Page attribution rides two store columns, never ``WikiPage`` fields, so the
    # console wire type stays byte-identical; both are stripped before validation.
    _PAGE_STORE_COLS = ("slug", "page_id", "commit_sha", "job_id")

    def save_page(
        self,
        slug: str,
        page: WikiPage,
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Persist *page* for the project *slug*; overwrites if same page_id."""
        doc = {
            "slug": slug,
            "page_id": page.id,
            "commit_sha": commit_sha,
            "job_id": job_id,
            **page.model_dump(by_alias=False),
        }
        self._col("wiki_pages").replace_one(
            {"slug": slug, "page_id": page.id}, doc, upsert=True
        )

    def _get_page_raw(self, slug: str, page_id: str) -> WikiPage | None:
        """Return a single wiki page, or None if absent (no doc-guard)."""
        doc = self._col("wiki_pages").find_one({"slug": slug, "page_id": page_id})
        if doc is None:
            return None
        clean = _strip_mongo_meta(doc)
        for col in self._PAGE_STORE_COLS:
            clean.pop(col, None)
        return WikiPage.model_validate(clean)

    def list_pages(self, slug: str) -> list[WikiPage]:
        """Return all pages for project *slug*."""
        pages: list[WikiPage] = []
        for doc in self._col("wiki_pages").find({"slug": slug}):
            clean = _strip_mongo_meta(doc)
            for col in self._PAGE_STORE_COLS:
                clean.pop(col, None)
            page = WikiPage.model_validate(clean)
            pages.append(page)
        return pages

    def delete_page(self, slug: str, page_id: str) -> bool:
        """Delete a single wiki page document. Returns True on a hit."""
        result = self._col("wiki_pages").delete_one(
            {"slug": slug, "page_id": page_id}
        )
        return result.deleted_count > 0

    def prune_pages(self, slug: str, keep: Iterable[str]) -> int:
        """Bulk-drop pages not in *keep* in a single Mongo round-trip."""
        keep_list = list(keep)
        result = self._col("wiki_pages").delete_many(
            {"slug": slug, "page_id": {"$nin": keep_list}}
        )
        return int(result.deleted_count)

    # -- Indexing jobs -------------------------------------------------------

    def create_job(self, job: IndexingJob) -> None:
        """Persist a new indexing job."""
        doc = {"event_count": 0, **job.model_dump(by_alias=False)}
        self._col("wiki_jobs").replace_one({"job_id": job.job_id}, doc, upsert=True)

    def get_job(self, job_id: str) -> IndexingJob | None:
        """Return the indexing job, or None if absent."""
        doc = self._col("wiki_jobs").find_one({"job_id": job_id})
        if doc is None:
            return None
        return IndexingJob.model_validate(_clean_for_model(doc, IndexingJob))

    def _write_job_patch(self, job_id: str, patch: JobPatch) -> IndexingJob:
        """``$set`` exactly the patch's fields in ONE server-side update.

        No lock, and none would help: the API runs several worker processes, so
        a process-local lock proves nothing about the writer next door. Mongo's
        per-document atomicity is the mechanism instead — a ``$set`` naming only
        these fields leaves every field it does not name exactly as another
        writer left it, which removes the lost update without any locking at
        all. A ``$set`` of the whole dumped model instead is what makes two
        writers collide on fields neither has touched.

        Returning the AFTER document is part of the same property: the caller
        gets what is actually stored, a concurrent writer's fields included,
        rather than a locally merged guess that would report them reverted.
        """
        from pymongo import ReturnDocument

        doc = self._col("wiki_jobs").find_one_and_update(
            {"job_id": job_id},
            {"$set": patch.fields},
            return_document=ReturnDocument.AFTER,
        )
        if doc is None:
            raise KeyError(f"Job not found: {job_id}")
        return IndexingJob.model_validate(_clean_for_model(doc, IndexingJob))

    def list_jobs(self, slug: str | None = None) -> list[IndexingJob]:
        """Return all jobs, newest first by ``phase_started_at``, filtered to *slug*.

        Mongo ``find()`` has no inherent order, and sorting by ``job_id`` for
        reproducibility would not fix that: ``job_id`` is a ``uuid4`` hex, which
        sorts RANDOMLY with respect to when a job actually ran, and at least one
        caller (``resolve_qa_clone_dir``) assumes this returns most-recent-first
        and picks the FIRST ``complete`` hit — under such a sort an arbitrary
        old checkout. Sorted in Python
        rather than via a Mongo-side ``.sort()`` so both drivers apply the
        IDENTICAL rule (ISO-8601 ``phase_started_at``, so lexicographic ==
        chronological; a job with no timestamp sorts last) instead of relying
        on each backend's own null-ordering semantics to happen to agree.
        """
        query: dict[str, Any] = {}
        if slug is not None:
            query["slug"] = slug
        jobs: list[IndexingJob] = []
        for doc in self._col("wiki_jobs").find(query):
            jobs.append(IndexingJob.model_validate(_clean_for_model(doc, IndexingJob)))
        return sorted(jobs, key=lambda j: j.phase_started_at or "", reverse=True)

    def _append_job_event(self, job_id: str, event: dict[str, Any]) -> int:
        """Append one validated event to this Mongo job timeline. Cost: ``O(1)``."""
        idx = self._atomic_next_idx("wiki_jobs", "job_id", job_id)
        self._col("wiki_job_events").insert_one({"job_id": job_id, "idx": idx, **event})
        return idx

    def load_job_events(
        self, job_id: str, after_idx: int = -1
    ) -> list[dict[str, Any]]:
        """Return job events with idx > *after_idx* (-1 returns all)."""
        query: dict[str, Any] = {"job_id": job_id, "idx": {"$gt": after_idx}}
        results: list[dict[str, Any]] = []
        for doc in self._col("wiki_job_events").find(query).sort("idx", 1):
            clean = _strip_mongo_meta(doc)
            clean.pop("job_id", None)
            results.append(clean)
        return results

    def cancel_job(self, job_id: str) -> bool:
        """Cancel *job_id*; return True on first cancel, False if already cancelled."""
        job = self.get_job(job_id)
        if job is None:
            return False
        if job.status == "cancelled":
            return False
        self.update_job(job_id, status="cancelled")
        self.append_job_event(job_id, {"type": "cancelled"})
        return True

    def attach_job_session(self, job_id: str, session_id: str) -> None:
        """Associate a Mewbo session_id with an indexing job."""
        self._col("wiki_jobs").update_one(
            {"job_id": job_id},
            {"$set": {"session_id": session_id}},
        )

    def get_job_session(self, job_id: str) -> str | None:
        """Return the session_id attached to *job_id*, or None."""
        doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"session_id": 1})
        if doc is None:
            return None
        val = doc.get("session_id")
        return str(val) if val else None

    def find_job_by_session(self, session_id: str) -> str | None:
        """Reverse lookup: return the job_id for *session_id*, or None."""
        doc = self._col("wiki_jobs").find_one(
            {"session_id": session_id}, {"job_id": 1}
        )
        if doc is None:
            return None
        val = doc.get("job_id")
        return str(val) if val else None

    def save_job_plan(self, job_id: str, plan: list[dict[str, Any]]) -> None:
        """Persist the page-plan list for *job_id*; overwrites any previous plan."""
        self._col("wiki_jobs").update_one(
            {"job_id": job_id},
            {"$set": {"plan": plan}},
        )

    def get_job_plan(self, job_id: str) -> list[dict[str, Any]] | None:
        """Return the page-plan list, or None if no plan has been committed yet."""
        doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"plan": 1})
        if doc is None:
            return None
        plan = doc.get("plan")
        return plan if isinstance(plan, list) else None

    def save_resume_plan(self, job_id: str, plan: dict[str, Any]) -> None:
        """Persist the resume-plan dict on the job doc; overwrites any previous one."""
        self._col("wiki_jobs").update_one(
            {"job_id": job_id},
            {"$set": {"resume_plan": plan}},
        )

    def get_resume_plan(self, job_id: str) -> dict[str, Any] | None:
        """Return the persisted resume-plan dict, or None if the job isn't resuming."""
        doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"resume_plan": 1})
        if doc is None:
            return None
        val = doc.get("resume_plan")
        return val if isinstance(val, dict) else None

    def save_act_plan(self, job_id: str, plan: dict[str, Any]) -> None:
        """Persist the act-stage record on the job doc; overwrites any previous one."""
        self._col("wiki_jobs").update_one(
            {"job_id": job_id},
            {"$set": {"act_plan": plan}},
        )

    def get_act_plan(self, job_id: str) -> dict[str, Any] | None:
        """Return the persisted act-stage record, or None if stage 1 hasn't finished."""
        doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"act_plan": 1})
        if doc is None:
            return None
        val = doc.get("act_plan")
        return val if isinstance(val, dict) else None

    def get_job_submitted_count(self, job_id: str) -> int:
        """Return the number of pages submitted so far for *job_id*."""
        doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"submitted_pages": 1})
        if doc is None:
            return 0
        return int(doc.get("submitted_pages", 0))

    def claim_job_page(self, slug: str, job_id: str, page_id: str) -> PageClaim:
        """Claim + count in ONE conditional update, so racing writers can't double-count."""
        from pymongo import ReturnDocument

        self._seed_claim_record(slug, job_id)
        doc = self._col("wiki_jobs").find_one_and_update(
            {"job_id": job_id, "submitted_page_ids": {"$ne": page_id}},
            {"$addToSet": {"submitted_page_ids": page_id}},
            return_document=ReturnDocument.AFTER,
        )
        if doc is None:
            # No match means one of two things, and they are not
            # interchangeable: the id is already claimed (a re-submit — the
            # common case), or the job does not exist (a programming error every
            # other job-keyed write raises on). Read once more to tell them
            # apart rather than reporting a missing job as a quiet no-op claim.
            doc = self._col("wiki_jobs").find_one(
                {"job_id": job_id}, {"submitted_page_ids": 1}
            )
            if doc is None:
                raise KeyError(f"Job not found: {job_id}")
            return PageClaim(count=len(doc.get("submitted_page_ids") or []), is_new=False)
        count = len(doc.get("submitted_page_ids") or [])
        # Mirrored, never authoritative — the SET is the count, so the cheap
        # ``get_job_submitted_count`` read can never disagree with it.
        self._col("wiki_jobs").update_one(
            {"job_id": job_id}, {"$set": {"submitted_pages": count}}
        )
        return PageClaim(count=count, is_new=True)

    def _seed_claim_record(self, slug: str, job_id: str) -> None:
        """Give a pre-claim job its claim set from attribution, once.

        ``{"submitted_page_ids": {"$ne": <id>}}`` matches a document that lacks
        the field ENTIRELY, so without this a job written before claims existed
        would claim its way up from zero while its already-written pages sat
        unaccounted — reporting a fraction of its real progress and, on resume,
        regenerating pages it had already produced correctly.

        Idempotent and race-safe: the filter requires the field to be absent, so
        a second caller seeding concurrently either writes the same derived list
        or does nothing.
        """
        seed = self.page_ids_for_job(slug, job_id)
        if not seed:
            return
        self._col("wiki_jobs").update_one(
            {"job_id": job_id, "submitted_page_ids": {"$exists": False}},
            {"$set": {"submitted_page_ids": sorted(seed)}},
        )

    def get_job_page_ids(self, slug: str, job_id: str) -> frozenset[str]:
        """Return the page ids *job_id* wrote (claim record, else attribution)."""
        doc = self._col("wiki_jobs").find_one(
            {"job_id": job_id}, {"submitted_page_ids": 1}
        )
        if doc is None:
            return frozenset()
        ids = doc.get("submitted_page_ids")
        if ids is None:
            return self.page_ids_for_job(slug, job_id)
        return frozenset(str(p) for p in ids)

    def page_ids_for_job(self, slug: str, job_id: str) -> frozenset[str]:
        """Page ids whose stored attribution column names *job_id*."""
        return frozenset(
            str(d["page_id"])
            for d in self._col("wiki_pages").find(
                {"slug": slug, "job_id": job_id}, {"page_id": 1}
            )
            if d.get("page_id")
        )

    def save_job_submission(self, job_id: str, submission: dict[str, Any]) -> None:
        """Persist the wizard submission dict for *job_id* (token must be absent)."""
        self._col("wiki_jobs").update_one(
            {"job_id": job_id},
            {"$set": {"submission": submission}},
        )

    def get_job_submission(self, job_id: str) -> dict[str, Any] | None:
        """Return the persisted submission dict, or None if not yet saved."""
        doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"submission": 1})
        if doc is None:
            return None
        val = doc.get("submission")
        return val if isinstance(val, dict) else None

    # -- Repository credentials (isolated collection) ------------------------

    def save_credentials(self, slug: str, blob: dict[str, Any]) -> None:
        """Persist the encoded credential blob for *slug* (slug PK, upsert)."""
        self._col("wiki_credentials").replace_one(
            {"slug": slug}, {"slug": slug, "blob": blob}, upsert=True
        )

    def get_credentials(self, slug: str) -> dict[str, Any] | None:
        """Return the encoded credential blob for *slug*, or None."""
        doc = self._col("wiki_credentials").find_one({"slug": slug}, {"blob": 1})
        if doc is None:
            return None
        val = doc.get("blob")
        return val if isinstance(val, dict) else None

    def delete_credentials(self, slug: str) -> bool:
        """Delete *slug*'s credential document; return True if one existed."""
        result = self._col("wiki_credentials").delete_one({"slug": slug})
        return result.deleted_count > 0

    def list_credentials(self) -> dict[str, dict[str, Any]]:
        """Return every stored credential blob keyed by scope (full scan).

        Prefers the blob's own ``scope`` field (stamped by
        ``CredentialStore.save``, matching the JSON driver's precedence); falls
        back to the document's ``slug`` key for a blob missing that field.
        """
        out: dict[str, dict[str, Any]] = {}
        for doc in self._col("wiki_credentials").find({}, {"slug": 1, "blob": 1}):
            slug = doc.get("slug")
            blob = doc.get("blob")
            if isinstance(blob, dict):
                scope = blob["scope"] if isinstance(blob.get("scope"), str) else slug
                if isinstance(scope, str):
                    out[scope] = blob
        return out

    # -- Restart-recovery counter (slug-keyed collection) --------------------

    def get_recovery_attempts(self, slug: str) -> int:
        """Return the recovery-attempt count for *slug* (0 if never recovered)."""
        doc = self._col("wiki_recovery").find_one({"slug": slug}, {"attempts": 1})
        return int(doc.get("attempts", 0)) if doc else 0

    def bump_recovery_attempts(self, slug: str) -> int:
        """Atomically increment *slug*'s recovery counter; return the new value."""
        from pymongo import ReturnDocument

        doc = self._col("wiki_recovery").find_one_and_update(
            {"slug": slug},
            {"$inc": {"attempts": 1}},
            upsert=True,
            return_document=ReturnDocument.AFTER,
        )
        return int(doc["attempts"])

    def reset_recovery_attempts(self, slug: str) -> None:
        """Clear *slug*'s recovery counter document (user-initiated resume fresh budget)."""
        self._col("wiki_recovery").delete_one({"slug": slug})

    # -- QA ------------------------------------------------------------------

    def save_qa(self, answer: QaAnswer) -> None:
        """Persist a QA answer record (creation: resets event_count, no session yet)."""
        doc = {"event_count": 0, **answer.model_dump(by_alias=False)}
        self._col("wiki_qa").replace_one(
            {"answer_id": answer.answer_id}, doc, upsert=True
        )

    def update_qa_fields(self, answer: QaAnswer) -> None:
        """In-place ``$set`` of the QaAnswer fields only.

        Leaves ``event_count`` + ``session_id`` (this backend packs both into the
        same doc) intact, unlike ``save_qa``'s full replace.
        """
        self._col("wiki_qa").update_one(
            {"answer_id": answer.answer_id},
            {"$set": answer.model_dump(by_alias=False)},
        )

    def get_qa(self, answer_id: str) -> QaAnswer | None:
        """Return the QA answer, or None if absent."""
        doc = self._col("wiki_qa").find_one({"answer_id": answer_id})
        if doc is None:
            return None
        return QaAnswer.model_validate(_clean_for_model(doc, QaAnswer))

    def list_qa(self, status: str | None = None) -> list[QaAnswer]:
        """Return all QA answers, optionally filtered to *status* (server-side)."""
        query: dict[str, Any] = {} if status is None else {"status": status}
        return [
            QaAnswer.model_validate(_clean_for_model(doc, QaAnswer))
            for doc in self._col("wiki_qa").find(query)
        ]

    def attach_qa_session(self, answer_id: str, session_id: str) -> None:
        """Associate a Mewbo session_id with a QA answer."""
        self._col("wiki_qa").update_one(
            {"answer_id": answer_id},
            {"$set": {"session_id": session_id}},
        )

    def get_qa_session(self, answer_id: str) -> str | None:
        """Return the session_id attached to *answer_id*, or None."""
        doc = self._col("wiki_qa").find_one({"answer_id": answer_id}, {"session_id": 1})
        if doc is None:
            return None
        val = doc.get("session_id")
        return str(val) if val else None

    def find_qa_by_session(self, session_id: str) -> str | None:
        """Reverse lookup: return the answer_id for *session_id*, or None."""
        doc = self._col("wiki_qa").find_one(
            {"session_id": session_id}, {"answer_id": 1}
        )
        if doc is None:
            return None
        val = doc.get("answer_id")
        return str(val) if val else None

    def append_qa_event(self, answer_id: str, event: dict[str, Any]) -> int:
        """Append *event* to the QA event log; return the monotonic idx."""
        idx = self._atomic_next_idx("wiki_qa", "answer_id", answer_id)
        self._col("wiki_qa_events").insert_one(
            {"answer_id": answer_id, "idx": idx, **event}
        )
        return idx

    def load_qa_events(
        self, answer_id: str, after_idx: int = -1
    ) -> list[dict[str, Any]]:
        """Return QA events with idx > *after_idx* (-1 returns all)."""
        query: dict[str, Any] = {"answer_id": answer_id, "idx": {"$gt": after_idx}}
        results: list[dict[str, Any]] = []
        for doc in self._col("wiki_qa_events").find(query).sort("idx", 1):
            clean = _strip_mongo_meta(doc)
            clean.pop("answer_id", None)
            results.append(clean)
        return results

    # -- Graph + embeddings --------------------------------------------------

    def _ensure_graph_indexes(self) -> None:
        """Create graph collection indexes (called lazily on first upsert)."""
        if getattr(self, "_graph_idx_done", False):
            return
        from pymongo import ASCENDING

        self._col("wiki_graph_nodes").create_index(
            [("slug", ASCENDING), ("node_id", ASCENDING)],
            name="ix_graph_nodes_slug_nid",
            unique=True,
            background=True,
        )
        self._col("wiki_graph_edges").create_index(
            [
                ("slug", ASCENDING),
                ("source", ASCENDING),
                ("target", ASCENDING),
                ("type", ASCENDING),
            ],
            name="ix_graph_edges_slug_src_tgt_type",
            unique=True,
            background=True,
        )
        self._col("wiki_embeddings").create_index(
            [("slug", ASCENDING), ("node_id", ASCENDING)],
            name="ix_embeddings_slug_nid",
            unique=True,
            background=True,
        )
        # Non-unique (slug, commit_sha) indexes back the commit-scoped count the
        # resume predicate keys on AND the per-commit supersede sweep. The unique
        # keys above are unchanged — a node_id still identifies ONE row per slug,
        # so a re-index overwrites the shared symbol and supersede reaps only the
        # commit-only stragglers (deleted files).
        for coll in ("wiki_graph_nodes", "wiki_graph_edges", "wiki_embeddings"):
            self._col(coll).create_index(
                [("slug", ASCENDING), ("commit_sha", ASCENDING)],
                name="ix_" + coll + "_slug_commit",
                background=True,
            )
        self._graph_idx_done = True

    def _bulk_upsert(
        self,
        collection: Any,
        ops: Iterable[tuple[dict[str, Any], dict[str, Any]]],
        *,
        total_batches: int | None = None,
        on_progress: Callable[[int, int], None] | None = None,
    ) -> None:
        """Apply ``(filter, document)`` ``$set`` upserts in bounded unordered batches.

        Cost: ``O(documents written)`` in server work, but
        ``ceil(n / _BULK_BATCH_SIZE)`` round-trips rather than ``n`` — a graph
        phase writes ~130k documents, and one round-trip apiece dominated the
        phase's write leg on a local socket, worse still over a network. Each
        op is the same single-document ``$set`` upsert a per-document loop
        issues, so the unique indexes still do the dedup and re-writing a
        document is still idempotent.

        The batch bound is what keeps memory ``O(_BULK_BATCH_SIZE)`` instead of
        ``O(all documents)``: *ops* is consumed lazily, so a caller streaming
        79k edges never materialises 79k pending operations.

        **A batch is keyed by its filter, keeping the LAST write.** A
        per-document loop applied two updates to one key in input order, so the
        last one won; ``ordered=False`` explicitly permits out-of-order or
        parallel execution, and MongoDB does not document that two operations
        sharing a filter in one batch resolve in list order. Folding here keeps
        last-write-wins a property of THIS code rather than of a server
        version — the same guarantee the loop gave, at the cost of one dict.
        Across batches it needs no help: batches are issued sequentially, so a
        later batch's write lands after an earlier one's.
        """
        from pymongo import UpdateOne

        batch: dict[tuple[tuple[str, Any], ...], Any] = {}
        completed = 0
        for filt, doc in ops:
            batch[tuple(sorted(filt.items()))] = UpdateOne(
                filt, {"$set": doc}, upsert=True
            )
            if len(batch) >= self._BULK_BATCH_SIZE:
                collection.bulk_write(list(batch.values()), ordered=False)
                completed += 1
                if on_progress is not None and total_batches is not None:
                    on_progress(completed, total_batches)
                batch = {}
        if batch:
            collection.bulk_write(list(batch.values()), ordered=False)
            completed += 1
            if on_progress is not None and total_batches is not None:
                on_progress(completed, total_batches)

    def upsert_nodes(
        self,
        slug: str,
        nodes: Iterable[GraphNode],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
        on_progress: Callable[[int, int], None] | None = None,
    ) -> None:
        """Upsert graph nodes for *slug*; dedup by (slug, node_id), stamp attribution.

        Cost: ``O(nodes)`` — offline, the ``graph`` phase. Batched through
        ``_bulk_upsert``, so the round-trip count is ``O(nodes / batch)``.
        """
        self._ensure_graph_indexes()
        total_batches = (
            math.ceil(len(nodes) / self._BULK_BATCH_SIZE)
            if isinstance(nodes, Collection)
            else None
        )
        self._bulk_upsert(
            self._col("wiki_graph_nodes"),
            (
                (
                    {"slug": slug, "node_id": node.node_id},
                    self._stamp_attribution(node, commit_sha, job_id).model_dump(by_alias=False),
                )
                for node in nodes
            ),
            total_batches=total_batches,
            on_progress=on_progress,
        )

    def upsert_edges(
        self,
        slug: str,
        edges: Iterable[GraphEdge],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
        on_progress: Callable[[int, int], None] | None = None,
    ) -> None:
        """Upsert graph edges for *slug*; dedup by (slug, source, target, type).

        Cost: ``O(edges)`` — offline, the ``graph`` phase. Batched through
        ``_bulk_upsert``, so the round-trip count is ``O(edges / batch)``.
        """
        self._ensure_graph_indexes()
        total_batches = (
            math.ceil(len(edges) / self._BULK_BATCH_SIZE)
            if isinstance(edges, Collection)
            else None
        )
        self._bulk_upsert(
            self._col("wiki_graph_edges"),
            (
                (
                    {
                        "slug": slug,
                        "source": edge.source,
                        "target": edge.target,
                        "type": edge.type,
                    },
                    self._stamp_attribution(edge, commit_sha, job_id).model_dump(by_alias=False),
                )
                for edge in edges
            ),
            total_batches=total_batches,
            on_progress=on_progress,
        )

    def upsert_embeddings(
        self,
        slug: str,
        items: Iterable[Embedding],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert embedding vectors for *slug*; dedup by (slug, node_id).

        Cost: ``O(vectors)`` — offline, the ``graph`` phase. Batched through
        ``_bulk_upsert``, so the round-trip count is ``O(vectors / batch)``.
        """
        self._ensure_graph_indexes()
        self._bulk_upsert(
            self._col("wiki_embeddings"),
            (
                (
                    {"slug": slug, "node_id": item.node_id},
                    self._embedding_doc(item, commit_sha, job_id),
                )
                for item in items
            ),
        )

    def _embedding_doc(
        self, item: Embedding, commit_sha: str | None, job_id: str | None
    ) -> dict[str, Any]:
        """The stored embedding document: the model dump plus the packed vector."""
        stamped = self._stamp_attribution(item, commit_sha, job_id)
        doc = stamped.model_dump(by_alias=False)
        # Written alongside the canonical list so ``vector_search`` can skip
        # the BSON-array decode entirely — see ``_VEC_F32``.
        doc[_VEC_F32] = _pack_f32(stamped.vector)
        return doc

    def query_graph(
        self,
        slug: str,
        *,
        scope: CommitScope,
        node_type: str | None = None,
        name_match: str | None = None,
        neighbors_of: str | None = None,
        node_ids: Collection[str] | None = None,
    ) -> list[GraphNode]:
        """See ``WikiStoreBase.query_graph``."""
        import re

        commit = scope.filter_fields()
        if neighbors_of is not None:
            # The seed's own generation bounds the walk — see the JSON driver.
            edge_query = {
                "slug": slug,
                "$or": [{"source": neighbors_of}, {"target": neighbors_of}],
                **commit,
            }
            related_ids: set[str] = set()
            for edge_doc in self._col("wiki_graph_edges").find(edge_query):
                src = edge_doc.get("source")
                tgt = edge_doc.get("target")
                if src == neighbors_of:
                    related_ids.add(tgt)
                else:
                    related_ids.add(src)
            if not related_ids:
                return []
            cursor = self._col("wiki_graph_nodes").find(
                {"slug": slug, "node_id": {"$in": list(related_ids)}, **commit}
            )
            return [GraphNodeAdapter.validate_python(_strip_mongo_meta(d)) for d in cursor]
        query: dict[str, Any] = {"slug": slug, **commit}
        if node_ids is not None:
            query["node_id"] = {"$in": list(node_ids)}
        if node_type is not None:
            query["type"] = node_type
        if name_match is not None:
            query["name"] = {"$regex": re.escape(name_match), "$options": "i"}
        cursor = self._col("wiki_graph_nodes").find(query)
        return [GraphNodeAdapter.validate_python(_strip_mongo_meta(d)) for d in cursor]

    def list_edges(self, slug: str, *, scope: CommitScope) -> list[GraphEdge]:
        """See ``WikiStoreBase.list_edges``."""
        cursor = self._col("wiki_graph_edges").find({"slug": slug, **scope.filter_fields()})
        return [GraphEdge.model_validate(_strip_mongo_meta(d)) for d in cursor]

    def count_graph_nodes(self, slug: str, *, commit_sha: str | None) -> int:
        """Count *slug* nodes stamped exactly *commit_sha* (``None`` matches None)."""
        self._ensure_graph_indexes()
        return int(
            self._col("wiki_graph_nodes").count_documents(
                {"slug": slug, "commit_sha": commit_sha}
            )
        )

    def supersede_graph_artifacts(
        self, slug: str, *, keep_commit_sha: str
    ) -> dict[str, int]:
        """Drop prior-commit graph + entity artifacts, preserving ``None``-stamped rows.

        ``{"$nin": [None, keep]}`` matches a REAL commit other than *keep* while
        leaving both ``None``-valued and field-absent rows untouched — the
        QA-minted / pre-isolation records supersede must not reap.
        """
        self._ensure_graph_indexes()
        self._ensure_memory_indexes()
        stale = {"slug": slug, "commit_sha": {"$nin": [None, keep_commit_sha]}}
        counts: dict[str, int] = {}
        for key, coll in (
            ("nodes", "wiki_graph_nodes"),
            ("edges", "wiki_graph_edges"),
            ("embeddings", "wiki_embeddings"),
            ("entities", "wiki_entities"),
            ("entity_edges", "wiki_entity_edges"),
            ("entity_embeddings", "wiki_entity_embeddings"),
        ):
            counts[key] = int(self._col(coll).delete_many(stale).deleted_count)
        return counts

    def restamp_graph_artifacts(
        self, slug: str, *, from_commit: str, to_commit: str
    ) -> dict[str, int]:
        """See ``WikiStoreBase.restamp_graph_artifacts``.

        An EXACT ``commit_sha`` match, deliberately not the ``$nin`` shape
        ``supersede_graph_artifacts`` uses: supersede asks "everything except
        the keeper", which would here sweep up older generations and
        field-absent rows and stamp genuinely dead code as live.
        """
        self._ensure_graph_indexes()
        self._ensure_memory_indexes()
        prior = {"slug": slug, "commit_sha": from_commit}
        patch = {"$set": {"commit_sha": to_commit}}
        counts: dict[str, int] = {}
        for key, coll in (
            ("nodes", "wiki_graph_nodes"),
            ("edges", "wiki_graph_edges"),
            ("embeddings", "wiki_embeddings"),
            ("entities", "wiki_entities"),
            ("entity_edges", "wiki_entity_edges"),
            ("entity_embeddings", "wiki_entity_embeddings"),
        ):
            counts[key] = int(self._col(coll).update_many(prior, patch).modified_count)
        return counts

    def vector_search(self, slug: str, qvec: list[float], k: int = 10) -> list[Embedding]:
        """Return top-k embeddings for *slug* by cosine similarity.

        Cost: ``O(embeddings for the slug)`` — every stored vector is still
        scored, so this remains the documented scale seam. What the packed path
        removes is the per-element Python cost of getting there: it reads only
        ``node_id`` + the ``_VEC_F32`` buffer, scores the whole project as one
        NumPy matrix product, and materialises exactly *k* ``Embedding`` models.

        Falls back to the pure-Python scan when NumPy is absent or any row
        lacks ``_VEC_F32`` — never to a PARTIAL result.
        """
        from .embedder import Embedder

        packed = self._vector_search_packed(slug, qvec, k)
        if packed is not None:
            return packed
        pool = [
            Embedding.model_validate(_clean_for_model(d, Embedding))
            for d in self._col("wiki_embeddings").find({"slug": slug})
        ]
        if not pool:
            return []
        scored = [(emb, Embedder.cosine(qvec, emb.vector)) for emb in pool]
        scored.sort(key=lambda t: t[1], reverse=True)
        return [emb for emb, _ in scored[:k]]

    def _vector_search_packed(
        self, slug: str, qvec: list[float], k: int
    ) -> list[Embedding] | None:
        """Top-k over the packed ``_VEC_F32`` buffers, or ``None`` if unusable.

        ``None`` means "this path cannot answer" — the caller falls back. It is
        returned rather than a short result whenever ANY row is missing or
        mis-sized, because scoring the subset that happens to be packed would
        silently search part of a project and still look like a complete answer.
        """
        try:
            import numpy as np  # noqa: PLC0415
        except ImportError:  # pragma: no cover - numpy ships with the retrieval extra
            return None
        if not qvec:
            return None
        width = len(qvec) * 4
        col = self._col("wiki_embeddings")
        ids: list[str] = []
        bufs: list[bytes] = []
        for d in col.find({"slug": slug}, {"node_id": 1, _VEC_F32: 1}):
            buf = d.get(_VEC_F32)
            if not isinstance(buf, bytes | bytearray) or len(buf) != width:
                return None
            ids.append(str(d["node_id"]))
            bufs.append(bytes(buf))
        if not ids:
            return None
        mat = np.frombuffer(b"".join(bufs), dtype="<f4").reshape(len(ids), len(qvec))
        q = np.asarray(qvec, dtype=np.float32)
        norms = np.linalg.norm(mat, axis=1) * float(np.linalg.norm(q))
        sims = np.divide(
            mat @ q, norms, out=np.zeros(len(ids), dtype=np.float32), where=norms != 0
        )
        # STABLE sort, matching the pure-Python path's ``sorted(..., reverse=True)``.
        # This project's vectors do produce exact score ties, and an unstable sort
        # reorders them — the two paths would then disagree on identical data.
        order = np.argsort(-sims, kind="stable")[:k]
        top = [ids[int(i)] for i in order]
        rank = {node_id: r for r, node_id in enumerate(top)}
        found = [
            Embedding.model_validate(_clean_for_model(d, Embedding))
            for d in col.find({"slug": slug, "node_id": {"$in": top}})
        ]
        found.sort(key=lambda e: rank[e.node_id])
        return found

    # -- Scoped graph deletes (incremental retract) --------------------------

    def delete_nodes_by_file(self, slug: str, file: str) -> int:
        """Delete *file*'s code nodes AND their vectors; return the NODE count."""
        # Read the ids BEFORE the delete — the same node→file join
        # ``delete_edges_by_source_file`` does, for the same reason: after the
        # nodes are gone nothing relates a vector back to a file.
        doomed = [
            d["node_id"]
            for d in self._col("wiki_graph_nodes").find(
                {"slug": slug, "file": file}, {"node_id": 1}
            )
        ]
        if not doomed:
            return 0
        result = self._col("wiki_graph_nodes").delete_many({"slug": slug, "file": file})
        self._col("wiki_embeddings").delete_many(
            {"slug": slug, "node_id": {"$in": doomed}}
        )
        return int(result.deleted_count)

    def delete_edges_by_source_file(self, slug: str, file: str) -> int:
        """Delete edges whose ``source`` node belongs to *file*; return count."""
        file_ids = [
            d["node_id"]
            for d in self._col("wiki_graph_nodes").find(
                {"slug": slug, "file": file}, {"node_id": 1}
            )
        ]
        if not file_ids:
            return 0
        result = self._col("wiki_graph_edges").delete_many(
            {"slug": slug, "source": {"$in": file_ids}}
        )
        return int(result.deleted_count)

    # -- Memory layer (multiplex overlay) ------------------------------------

    def _ensure_memory_indexes(self) -> None:
        """Create memory-layer collection indexes (lazy, on first upsert)."""
        if getattr(self, "_mem_idx_done", False):
            return
        from pymongo import ASCENDING

        self._col("wiki_memory_nodes").create_index(
            [("slug", ASCENDING), ("node_id", ASCENDING)],
            name="ix_mem_nodes_slug_nid", unique=True, background=True,
        )
        self._col("wiki_memory_edges").create_index(
            [("slug", ASCENDING), ("source", ASCENDING), ("target", ASCENDING),
             ("type", ASCENDING)],
            name="ix_mem_edges_key", unique=True, background=True,
        )
        self._col("wiki_memory_edges").create_index(
            [("slug", ASCENDING), ("type", ASCENDING), ("target", ASCENDING)],
            name="ix_mem_edges_anchor", background=True,
        )
        self._col("wiki_memory_embeddings").create_index(
            [("slug", ASCENDING), ("node_id", ASCENDING)],
            name="ix_mem_emb_slug_nid", unique=True, background=True,
        )
        self._col("wiki_doc_notes").create_index(
            [("slug", ASCENDING), ("page_id", ASCENDING)],
            name="ix_doc_notes_slug_pid", unique=True, background=True,
        )
        self._col("wiki_file_manifest").create_index(
            [("slug", ASCENDING), ("path", ASCENDING)],
            name="ix_manifest_slug_path", unique=True, background=True,
        )
        # Abstract-entity overlay collections (same lazy-index pattern).
        self._col("wiki_entities").create_index(
            [("slug", ASCENDING), ("id", ASCENDING)],
            name="ix_entities_slug_id", unique=True, background=True,
        )
        self._col("wiki_entity_embeddings").create_index(
            [("slug", ASCENDING), ("entity_id", ASCENDING)],
            name="ix_entity_emb_slug_eid", unique=True, background=True,
        )
        self._col("wiki_entity_edges").create_index(
            [("slug", ASCENDING), ("id", ASCENDING)],
            name="ix_entity_edges_slug_id", unique=True, background=True,
        )
        self._col("wiki_entity_recommendations").create_index(
            [("slug", ASCENDING)],
            name="ix_entity_recs_slug", background=True,
        )
        # Backs the keyed recommendation upsert. Deliberately NOT unique, unlike
        # its entity/edge siblings: recommendations were appended unkeyed for
        # long enough that a live collection can already hold duplicate rows,
        # and a unique index build fails outright on those — taking every other
        # index in this method down with it. The upsert converges new writes; a
        # dedup of the historical rows is a migration, not an index.
        self._col("wiki_entity_recommendations").create_index(
            [("slug", ASCENDING), ("id", ASCENDING)],
            name="ix_entity_recs_slug_id", background=True,
        )
        # Non-unique (slug, commit_sha) indexes back the commit-scoped count the
        # resume predicate keys on AND the per-commit supersede sweep, mirroring
        # the graph collections in _ensure_graph_indexes. The unique keys above
        # are unchanged — an entity id still identifies ONE row per slug, so a
        # re-index overwrites the shared entity and supersede reaps only the
        # commit-only stragglers.
        for coll in ("wiki_entities", "wiki_entity_edges", "wiki_entity_embeddings"):
            self._col(coll).create_index(
                [("slug", ASCENDING), ("commit_sha", ASCENDING)],
                name="ix_" + coll + "_slug_commit",
                background=True,
            )
        self._mem_idx_done = True

    def upsert_memory_nodes(self, slug: str, nodes: Iterable[MemoryNode]) -> None:
        """Upsert memory nodes for *slug*; dedup by (slug, node_id)."""
        self._ensure_memory_indexes()
        col = self._col("wiki_memory_nodes")
        for node in nodes:
            col.update_one(
                {"slug": slug, "node_id": node.node_id},
                {"$set": node.model_dump(by_alias=False)},
                upsert=True,
            )

    def get_memory_node(self, slug: str, node_id: str) -> MemoryNode | None:
        """Return a single memory node, or None if absent."""
        doc = self._col("wiki_memory_nodes").find_one({"slug": slug, "node_id": node_id})
        if doc is None:
            return None
        return MemoryNode.model_validate(_strip_mongo_meta(doc))

    def delete_memory_node(self, slug: str, node_id: str) -> bool:
        """Delete a memory node + its embedding; return True if one was removed."""
        result = self._col("wiki_memory_nodes").delete_one(
            {"slug": slug, "node_id": node_id}
        )
        self._col("wiki_memory_embeddings").delete_one(
            {"slug": slug, "node_id": node_id}
        )
        return result.deleted_count > 0

    def query_memory(
        self, slug: str, *, filt: MemoryFilter | None = None
    ) -> list[MemoryNode]:
        """Return memory nodes matching *filt*'s node-level facets."""
        nodes = [
            MemoryNode.model_validate(_strip_mongo_meta(d))
            for d in self._col("wiki_memory_nodes").find({"slug": slug})
        ]
        if filt is None:
            return nodes
        return [n for n in nodes if filt.matches_node(n)]

    def upsert_memory_edges(self, slug: str, edges: Iterable[MemoryEdge]) -> None:
        """Upsert memory edges for *slug*; dedup by (slug, source, target, type)."""
        self._ensure_memory_indexes()
        col = self._col("wiki_memory_edges")
        for edge in edges:
            col.update_one(
                {"slug": slug, "source": edge.source, "target": edge.target,
                 "type": edge.type},
                {"$set": edge.model_dump(by_alias=False)},
                upsert=True,
            )

    def list_memory_edges(
        self,
        slug: str,
        *,
        node_id: str | None = None,
        include_invalidated: bool = False,
    ) -> list[MemoryEdge]:
        """Return memory edges, optionally scoped to ``source == node_id``."""
        query: dict[str, Any] = {"slug": slug}
        if node_id is not None:
            query["source"] = node_id
        if not include_invalidated:
            query["invalid_at"] = None
        return [
            MemoryEdge.model_validate(_strip_mongo_meta(d))
            for d in self._col("wiki_memory_edges").find(query)
        ]

    def memories_anchored_to(
        self,
        slug: str,
        entity_keys: Iterable[EntityKey],
        *,
        include_invalidated: bool = False,
    ) -> list[str]:
        """Reverse ANCHORS lookup: entity_keys → distinct memory node_ids."""
        query: dict[str, Any] = {
            "slug": slug, "type": "ANCHORS", "target": {"$in": list(entity_keys)}
        }
        if not include_invalidated:
            query["invalid_at"] = None
        seen: list[str] = []
        seen_set: set[str] = set()
        for d in self._col("wiki_memory_edges").find(query):
            src = d.get("source")
            if src not in seen_set:
                seen_set.add(src)
                seen.append(src)
        return seen

    def _live_anchored_ids(self, slug: str) -> set[str]:
        """Memory node_ids with ≥1 live ANCHORS edge."""
        return {
            d["source"]
            for d in self._col("wiki_memory_edges").find(
                {"slug": slug, "type": "ANCHORS", "invalid_at": None}, {"source": 1}
            )
        }

    def upsert_memory_embeddings(
        self, slug: str, items: Iterable[MemoryEmbedding]
    ) -> None:
        """Upsert memory embedding vectors for *slug*; dedup by (slug, node_id)."""
        self._ensure_memory_indexes()
        col = self._col("wiki_memory_embeddings")
        for item in items:
            col.update_one(
                {"slug": slug, "node_id": item.node_id},
                {"$set": item.model_dump(by_alias=False)},
                upsert=True,
            )

    def memory_vector_search(
        self,
        slug: str,
        qvec: list[float],
        k: int = 10,
        *,
        filt: MemoryFilter | None = None,
    ) -> list[MemoryEmbedding]:
        """Top-k memory embeddings by cosine, after applying *filt*."""
        pool = [
            MemoryEmbedding.model_validate(_strip_mongo_meta(d))
            for d in self._col("wiki_memory_embeddings").find({"slug": slug})
        ]
        return self._rank_memory(slug, pool, qvec, k, filt)

    # -- Doc-page notes ------------------------------------------------------

    def upsert_doc_notes(self, slug: str, notes: Iterable[DocPageNote]) -> None:
        """Upsert doc-page notes for *slug*; dedup by (slug, page_id)."""
        self._ensure_memory_indexes()
        col = self._col("wiki_doc_notes")
        for note in notes:
            col.update_one(
                {"slug": slug, "page_id": note.page_id},
                {"$set": note.model_dump(by_alias=False)},
                upsert=True,
            )

    def get_doc_note(self, slug: str, page_id: str) -> DocPageNote | None:
        """Return a single doc-page note, or None if absent."""
        doc = self._col("wiki_doc_notes").find_one({"slug": slug, "page_id": page_id})
        if doc is None:
            return None
        return DocPageNote.model_validate(_strip_mongo_meta(doc))

    def list_doc_notes(self, slug: str) -> list[DocPageNote]:
        """Return every doc-page note for *slug*."""
        return [
            DocPageNote.model_validate(_strip_mongo_meta(d))
            for d in self._col("wiki_doc_notes").find({"slug": slug})
        ]

    def delete_doc_note(self, slug: str, page_id: str) -> bool:
        """Delete a doc-page note; return True if one was removed."""
        result = self._col("wiki_doc_notes").delete_one(
            {"slug": slug, "page_id": page_id}
        )
        return result.deleted_count > 0

    # -- File manifest -------------------------------------------------------

    def upsert_file_manifest(
        self, slug: str, entries: Iterable[FileManifest]
    ) -> None:
        """Upsert file-manifest entries for *slug*; dedup by (slug, path)."""
        self._ensure_memory_indexes()
        col = self._col("wiki_file_manifest")
        for entry in entries:
            col.update_one(
                {"slug": slug, "path": entry.path},
                {"$set": entry.model_dump(by_alias=False)},
                upsert=True,
            )

    def get_file_manifest(self, slug: str, path: str) -> FileManifest | None:
        """Return a single file-manifest entry, or None if absent."""
        doc = self._col("wiki_file_manifest").find_one({"slug": slug, "path": path})
        if doc is None:
            return None
        return FileManifest.model_validate(_strip_mongo_meta(doc))

    def list_file_manifest(self, slug: str) -> list[FileManifest]:
        """Return every file-manifest entry for *slug*."""
        return [
            FileManifest.model_validate(_strip_mongo_meta(d))
            for d in self._col("wiki_file_manifest").find({"slug": slug})
        ]

    def delete_file_manifest(self, slug: str, path: str) -> bool:
        """Delete a file-manifest entry; return True if one was removed."""
        result = self._col("wiki_file_manifest").delete_one(
            {"slug": slug, "path": path}
        )
        return result.deleted_count > 0

    # -- Abstract-entity layer (multiplex overlay) ---------------------------
    #
    # Mirrors the memory-node Mongo block exactly: per-(slug, key) upsert,
    # ``slug`` carried as a store-internal field and stripped on read.

    def upsert_entities(
        self,
        slug: str,
        entities: Iterable[Entity],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert entities for *slug*; dedup by (slug, id), stamp attribution."""
        self._ensure_memory_indexes()
        col = self._col("wiki_entities")
        for entity in entities:
            stamped = self._stamp_attribution(entity, commit_sha, job_id)
            col.update_one(
                {"slug": slug, "id": entity.id},
                {"$set": {"slug": slug, **stamped.model_dump(by_alias=False)}},
                upsert=True,
            )

    def get_entity(self, slug: str, entity_id: str) -> Entity | None:
        """Return a single entity, or None if absent."""
        doc = self._col("wiki_entities").find_one({"slug": slug, "id": entity_id})
        if doc is None:
            return None
        clean = _strip_mongo_meta(doc)
        clean.pop("slug", None)
        return Entity.model_validate(clean)

    def query_entities(
        self, slug: str, *, filt: EntityFilter | None = None
    ) -> list[Entity]:
        """Return entities matching *filt*'s facets."""
        out: list[Entity] = []
        for doc in self._col("wiki_entities").find({"slug": slug}):
            clean = _strip_mongo_meta(doc)
            clean.pop("slug", None)
            out.append(Entity.model_validate(clean))
        if filt is None:
            return out
        return [e for e in out if filt.matches(e)]

    def count_entities(self, slug: str, *, commit_sha: str | None) -> int:
        """Count *slug* entities stamped exactly *commit_sha* (``None`` matches None)."""
        self._ensure_memory_indexes()
        return int(
            self._col("wiki_entities").count_documents(
                {"slug": slug, "commit_sha": commit_sha}
            )
        )

    def upsert_entity_embeddings(
        self,
        slug: str,
        items: Iterable[EntityEmbedding],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert entity embedding vectors for *slug*; dedup by (slug, entity_id)."""
        self._ensure_memory_indexes()
        col = self._col("wiki_entity_embeddings")
        for item in items:
            stamped = self._stamp_attribution(item, commit_sha, job_id)
            col.update_one(
                {"slug": slug, "entity_id": item.entity_id},
                {"$set": stamped.model_dump(by_alias=False)},
                upsert=True,
            )

    def entity_vector_search(
        self, slug: str, qvec: list[float], k: int = 10
    ) -> list[EntityEmbedding]:
        """Return top-k entity embeddings for *slug* by cosine (in-memory scoring)."""
        pool = [
            EntityEmbedding.model_validate(_strip_mongo_meta(d))
            for d in self._col("wiki_entity_embeddings").find({"slug": slug})
        ]
        return self._rank_embeddings(pool, qvec, k)

    def upsert_entity_edges(
        self,
        slug: str,
        edges: Iterable[EntityRelation],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert entity relations for *slug*; dedup by (slug, id)."""
        self._ensure_memory_indexes()
        col = self._col("wiki_entity_edges")
        for edge in edges:
            stamped = self._stamp_attribution(edge, commit_sha, job_id)
            col.update_one(
                {"slug": slug, "id": edge.id},
                {"$set": {"slug": slug, **stamped.model_dump(by_alias=False)}},
                upsert=True,
            )

    def list_entity_edges(
        self, slug: str, *, source_id: str | None = None
    ) -> list[EntityRelation]:
        """Return entity relations, optionally scoped to ``source_id``."""
        query: dict[str, Any] = {"slug": slug}
        if source_id is not None:
            query["source_id"] = source_id
        out: list[EntityRelation] = []
        for d in self._col("wiki_entity_edges").find(query):
            clean = _strip_mongo_meta(d)
            clean.pop("slug", None)
            out.append(EntityRelation.model_validate(clean))
        return out

    def save_entity_recommendation(
        self, slug: str, rec: EntityRecommendation
    ) -> None:
        """Upsert a recommendation for *slug*; dedup by (slug, id)."""
        self._ensure_memory_indexes()
        self._col("wiki_entity_recommendations").update_one(
            {"slug": slug, "id": rec.id},
            {"$set": {"slug": slug, **rec.model_dump(by_alias=False)}},
            upsert=True,
        )

    def get_entity_recommendations(self, slug: str) -> list[EntityRecommendation]:
        """Return every persisted entity recommendation for *slug*."""
        out: list[EntityRecommendation] = []
        for d in self._col("wiki_entity_recommendations").find({"slug": slug}):
            clean = _strip_mongo_meta(d)
            clean.pop("slug", None)
            out.append(EntityRecommendation.model_validate(clean))
        return out

__init__(*, client: Any = None, uri: str | None = None, database: str | None = None) -> None

Initialize MongoDB connection and ensure indexes exist.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2742
2743
2744
2745
2746
2747
2748
2749
2750
2751
2752
2753
2754
2755
2756
2757
2758
2759
2760
2761
2762
2763
2764
2765
def __init__(
    self,
    *,
    client: Any = None,
    uri: str | None = None,
    database: str | None = None,
) -> None:
    """Initialize MongoDB connection and ensure indexes exist."""
    if client is None:
        from pymongo import MongoClient

        _uri = uri or get_config_value(
            "storage", "mongodb", "uri", default="mongodb://localhost:27017"
        )
        client = MongoClient(_uri, serverSelectionTimeoutMS=5000)
        # Fail fast — mirrors MongoSessionStore.
        client.admin.command("ping")
    if database is None:
        database = get_config_value(
            "storage", "mongodb", "database", default="mewbo"
        )
    self._client = client
    self._db = client[database]
    self._ensure_indexes()

append_qa_event(answer_id: str, event: dict[str, Any]) -> int

Append event to the QA event log; return the monotonic idx.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3370
3371
3372
3373
3374
3375
3376
def append_qa_event(self, answer_id: str, event: dict[str, Any]) -> int:
    """Append *event* to the QA event log; return the monotonic idx."""
    idx = self._atomic_next_idx("wiki_qa", "answer_id", answer_id)
    self._col("wiki_qa_events").insert_one(
        {"answer_id": answer_id, "idx": idx, **event}
    )
    return idx

attach_job_session(job_id: str, session_id: str) -> None

Associate a Mewbo session_id with an indexing job.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3084
3085
3086
3087
3088
3089
def attach_job_session(self, job_id: str, session_id: str) -> None:
    """Associate a Mewbo session_id with an indexing job."""
    self._col("wiki_jobs").update_one(
        {"job_id": job_id},
        {"$set": {"session_id": session_id}},
    )

attach_qa_session(answer_id: str, session_id: str) -> None

Associate a Mewbo session_id with a QA answer.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3345
3346
3347
3348
3349
3350
def attach_qa_session(self, answer_id: str, session_id: str) -> None:
    """Associate a Mewbo session_id with a QA answer."""
    self._col("wiki_qa").update_one(
        {"answer_id": answer_id},
        {"$set": {"session_id": session_id}},
    )

bump_recovery_attempts(slug: str) -> int

Atomically increment slug's recovery counter; return the new value.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3294
3295
3296
3297
3298
3299
3300
3301
3302
3303
3304
def bump_recovery_attempts(self, slug: str) -> int:
    """Atomically increment *slug*'s recovery counter; return the new value."""
    from pymongo import ReturnDocument

    doc = self._col("wiki_recovery").find_one_and_update(
        {"slug": slug},
        {"$inc": {"attempts": 1}},
        upsert=True,
        return_document=ReturnDocument.AFTER,
    )
    return int(doc["attempts"])

cancel_job(job_id: str) -> bool

Cancel job_id; return True on first cancel, False if already cancelled.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3073
3074
3075
3076
3077
3078
3079
3080
3081
3082
def cancel_job(self, job_id: str) -> bool:
    """Cancel *job_id*; return True on first cancel, False if already cancelled."""
    job = self.get_job(job_id)
    if job is None:
        return False
    if job.status == "cancelled":
        return False
    self.update_job(job_id, status="cancelled")
    self.append_job_event(job_id, {"type": "cancelled"})
    return True

claim_job_page(slug: str, job_id: str, page_id: str) -> PageClaim

Claim + count in ONE conditional update, so racing writers can't double-count.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3161
3162
3163
3164
3165
3166
3167
3168
3169
3170
3171
3172
3173
3174
3175
3176
3177
3178
3179
3180
3181
3182
3183
3184
3185
3186
3187
3188
3189
def claim_job_page(self, slug: str, job_id: str, page_id: str) -> PageClaim:
    """Claim + count in ONE conditional update, so racing writers can't double-count."""
    from pymongo import ReturnDocument

    self._seed_claim_record(slug, job_id)
    doc = self._col("wiki_jobs").find_one_and_update(
        {"job_id": job_id, "submitted_page_ids": {"$ne": page_id}},
        {"$addToSet": {"submitted_page_ids": page_id}},
        return_document=ReturnDocument.AFTER,
    )
    if doc is None:
        # No match means one of two things, and they are not
        # interchangeable: the id is already claimed (a re-submit — the
        # common case), or the job does not exist (a programming error every
        # other job-keyed write raises on). Read once more to tell them
        # apart rather than reporting a missing job as a quiet no-op claim.
        doc = self._col("wiki_jobs").find_one(
            {"job_id": job_id}, {"submitted_page_ids": 1}
        )
        if doc is None:
            raise KeyError(f"Job not found: {job_id}")
        return PageClaim(count=len(doc.get("submitted_page_ids") or []), is_new=False)
    count = len(doc.get("submitted_page_ids") or [])
    # Mirrored, never authoritative — the SET is the count, so the cheap
    # ``get_job_submitted_count`` read can never disagree with it.
    self._col("wiki_jobs").update_one(
        {"job_id": job_id}, {"$set": {"submitted_pages": count}}
    )
    return PageClaim(count=count, is_new=True)

count_entities(slug: str, *, commit_sha: str | None) -> int

Count slug entities stamped exactly commit_sha (None matches None).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4131
4132
4133
4134
4135
4136
4137
4138
def count_entities(self, slug: str, *, commit_sha: str | None) -> int:
    """Count *slug* entities stamped exactly *commit_sha* (``None`` matches None)."""
    self._ensure_memory_indexes()
    return int(
        self._col("wiki_entities").count_documents(
            {"slug": slug, "commit_sha": commit_sha}
        )
    )

count_graph_nodes(slug: str, *, commit_sha: str | None) -> int

Count slug nodes stamped exactly commit_sha (None matches None).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3643
3644
3645
3646
3647
3648
3649
3650
def count_graph_nodes(self, slug: str, *, commit_sha: str | None) -> int:
    """Count *slug* nodes stamped exactly *commit_sha* (``None`` matches None)."""
    self._ensure_graph_indexes()
    return int(
        self._col("wiki_graph_nodes").count_documents(
            {"slug": slug, "commit_sha": commit_sha}
        )
    )

create_job(job: IndexingJob) -> None

Persist a new indexing job.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2995
2996
2997
2998
def create_job(self, job: IndexingJob) -> None:
    """Persist a new indexing job."""
    doc = {"event_count": 0, **job.model_dump(by_alias=False)}
    self._col("wiki_jobs").replace_one({"job_id": job.job_id}, doc, upsert=True)

create_project(project: Project) -> None

Persist a new project record.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2821
2822
2823
2824
2825
2826
def create_project(self, project: Project) -> None:
    """Persist a new project record."""
    doc = project.model_dump(by_alias=False)
    self._col("wiki_projects").replace_one(
        {"slug": project.slug}, doc, upsert=True
    )

delete_credentials(slug: str) -> bool

Delete slug's credential document; return True if one existed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3265
3266
3267
3268
def delete_credentials(self, slug: str) -> bool:
    """Delete *slug*'s credential document; return True if one existed."""
    result = self._col("wiki_credentials").delete_one({"slug": slug})
    return result.deleted_count > 0

delete_doc_note(slug: str, page_id: str) -> bool

Delete a doc-page note; return True if one was removed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4042
4043
4044
4045
4046
4047
def delete_doc_note(self, slug: str, page_id: str) -> bool:
    """Delete a doc-page note; return True if one was removed."""
    result = self._col("wiki_doc_notes").delete_one(
        {"slug": slug, "page_id": page_id}
    )
    return result.deleted_count > 0

delete_edges_by_source_file(slug: str, file: str) -> int

Delete edges whose source node belongs to file; return count.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3797
3798
3799
3800
3801
3802
3803
3804
3805
3806
3807
3808
3809
3810
def delete_edges_by_source_file(self, slug: str, file: str) -> int:
    """Delete edges whose ``source`` node belongs to *file*; return count."""
    file_ids = [
        d["node_id"]
        for d in self._col("wiki_graph_nodes").find(
            {"slug": slug, "file": file}, {"node_id": 1}
        )
    ]
    if not file_ids:
        return 0
    result = self._col("wiki_graph_edges").delete_many(
        {"slug": slug, "source": {"$in": file_ids}}
    )
    return int(result.deleted_count)

delete_file_manifest(slug: str, path: str) -> bool

Delete a file-manifest entry; return True if one was removed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4078
4079
4080
4081
4082
4083
def delete_file_manifest(self, slug: str, path: str) -> bool:
    """Delete a file-manifest entry; return True if one was removed."""
    result = self._col("wiki_file_manifest").delete_one(
        {"slug": slug, "path": path}
    )
    return result.deleted_count > 0

delete_memory_node(slug: str, node_id: str) -> bool

Delete a memory node + its embedding; return True if one was removed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3904
3905
3906
3907
3908
3909
3910
3911
3912
def delete_memory_node(self, slug: str, node_id: str) -> bool:
    """Delete a memory node + its embedding; return True if one was removed."""
    result = self._col("wiki_memory_nodes").delete_one(
        {"slug": slug, "node_id": node_id}
    )
    self._col("wiki_memory_embeddings").delete_one(
        {"slug": slug, "node_id": node_id}
    )
    return result.deleted_count > 0

delete_nodes_by_file(slug: str, file: str) -> int

Delete file's code nodes AND their vectors; return the NODE count.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3778
3779
3780
3781
3782
3783
3784
3785
3786
3787
3788
3789
3790
3791
3792
3793
3794
3795
def delete_nodes_by_file(self, slug: str, file: str) -> int:
    """Delete *file*'s code nodes AND their vectors; return the NODE count."""
    # Read the ids BEFORE the delete — the same node→file join
    # ``delete_edges_by_source_file`` does, for the same reason: after the
    # nodes are gone nothing relates a vector back to a file.
    doomed = [
        d["node_id"]
        for d in self._col("wiki_graph_nodes").find(
            {"slug": slug, "file": file}, {"node_id": 1}
        )
    ]
    if not doomed:
        return 0
    result = self._col("wiki_graph_nodes").delete_many({"slug": slug, "file": file})
    self._col("wiki_embeddings").delete_many(
        {"slug": slug, "node_id": {"$in": doomed}}
    )
    return int(result.deleted_count)

delete_page(slug: str, page_id: str) -> bool

Delete a single wiki page document. Returns True on a hit.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2978
2979
2980
2981
2982
2983
def delete_page(self, slug: str, page_id: str) -> bool:
    """Delete a single wiki page document. Returns True on a hit."""
    result = self._col("wiki_pages").delete_one(
        {"slug": slug, "page_id": page_id}
    )
    return result.deleted_count > 0

delete_project(slug: str) -> bool

Delete project slug; return True if deleted, False if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2840
2841
2842
2843
def delete_project(self, slug: str) -> bool:
    """Delete project *slug*; return True if deleted, False if absent."""
    result = self._col("wiki_projects").delete_one({"slug": slug})
    return result.deleted_count > 0

delete_project_settings(slug: str) -> bool

Delete slug's settings document; return True if one existed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2867
2868
2869
def delete_project_settings(self, slug: str) -> bool:
    """Delete *slug*'s settings document; return True if one existed."""
    return self._col("wiki_settings").delete_one({"slug": slug}).deleted_count > 0

Return top-k entity embeddings for slug by cosine (in-memory scoring).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4159
4160
4161
4162
4163
4164
4165
4166
4167
def entity_vector_search(
    self, slug: str, qvec: list[float], k: int = 10
) -> list[EntityEmbedding]:
    """Return top-k entity embeddings for *slug* by cosine (in-memory scoring)."""
    pool = [
        EntityEmbedding.model_validate(_strip_mongo_meta(d))
        for d in self._col("wiki_entity_embeddings").find({"slug": slug})
    ]
    return self._rank_embeddings(pool, qvec, k)

find_job_by_session(session_id: str) -> str | None

Reverse lookup: return the job_id for session_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3099
3100
3101
3102
3103
3104
3105
3106
3107
def find_job_by_session(self, session_id: str) -> str | None:
    """Reverse lookup: return the job_id for *session_id*, or None."""
    doc = self._col("wiki_jobs").find_one(
        {"session_id": session_id}, {"job_id": 1}
    )
    if doc is None:
        return None
    val = doc.get("job_id")
    return str(val) if val else None

find_qa_by_session(session_id: str) -> str | None

Reverse lookup: return the answer_id for session_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3360
3361
3362
3363
3364
3365
3366
3367
3368
def find_qa_by_session(self, session_id: str) -> str | None:
    """Reverse lookup: return the answer_id for *session_id*, or None."""
    doc = self._col("wiki_qa").find_one(
        {"session_id": session_id}, {"answer_id": 1}
    )
    if doc is None:
        return None
    val = doc.get("answer_id")
    return str(val) if val else None

get_act_plan(job_id: str) -> dict[str, Any] | None

Return the persisted act-stage record, or None if stage 1 hasn't finished.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3146
3147
3148
3149
3150
3151
3152
def get_act_plan(self, job_id: str) -> dict[str, Any] | None:
    """Return the persisted act-stage record, or None if stage 1 hasn't finished."""
    doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"act_plan": 1})
    if doc is None:
        return None
    val = doc.get("act_plan")
    return val if isinstance(val, dict) else None

get_credentials(slug: str) -> dict[str, Any] | None

Return the encoded credential blob for slug, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3257
3258
3259
3260
3261
3262
3263
def get_credentials(self, slug: str) -> dict[str, Any] | None:
    """Return the encoded credential blob for *slug*, or None."""
    doc = self._col("wiki_credentials").find_one({"slug": slug}, {"blob": 1})
    if doc is None:
        return None
    val = doc.get("blob")
    return val if isinstance(val, dict) else None

get_doc_note(slug: str, page_id: str) -> DocPageNote | None

Return a single doc-page note, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4028
4029
4030
4031
4032
4033
def get_doc_note(self, slug: str, page_id: str) -> DocPageNote | None:
    """Return a single doc-page note, or None if absent."""
    doc = self._col("wiki_doc_notes").find_one({"slug": slug, "page_id": page_id})
    if doc is None:
        return None
    return DocPageNote.model_validate(_strip_mongo_meta(doc))

get_entity(slug: str, entity_id: str) -> Entity | None

Return a single entity, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4109
4110
4111
4112
4113
4114
4115
4116
def get_entity(self, slug: str, entity_id: str) -> Entity | None:
    """Return a single entity, or None if absent."""
    doc = self._col("wiki_entities").find_one({"slug": slug, "id": entity_id})
    if doc is None:
        return None
    clean = _strip_mongo_meta(doc)
    clean.pop("slug", None)
    return Entity.model_validate(clean)

get_entity_recommendations(slug: str) -> list[EntityRecommendation]

Return every persisted entity recommendation for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4213
4214
4215
4216
4217
4218
4219
4220
def get_entity_recommendations(self, slug: str) -> list[EntityRecommendation]:
    """Return every persisted entity recommendation for *slug*."""
    out: list[EntityRecommendation] = []
    for d in self._col("wiki_entity_recommendations").find({"slug": slug}):
        clean = _strip_mongo_meta(d)
        clean.pop("slug", None)
        out.append(EntityRecommendation.model_validate(clean))
    return out

get_file_manifest(slug: str, path: str) -> FileManifest | None

Return a single file-manifest entry, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4064
4065
4066
4067
4068
4069
def get_file_manifest(self, slug: str, path: str) -> FileManifest | None:
    """Return a single file-manifest entry, or None if absent."""
    doc = self._col("wiki_file_manifest").find_one({"slug": slug, "path": path})
    if doc is None:
        return None
    return FileManifest.model_validate(_strip_mongo_meta(doc))

get_job(job_id: str) -> IndexingJob | None

Return the indexing job, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3000
3001
3002
3003
3004
3005
def get_job(self, job_id: str) -> IndexingJob | None:
    """Return the indexing job, or None if absent."""
    doc = self._col("wiki_jobs").find_one({"job_id": job_id})
    if doc is None:
        return None
    return IndexingJob.model_validate(_clean_for_model(doc, IndexingJob))

get_job_page_ids(slug: str, job_id: str) -> frozenset[str]

Return the page ids job_id wrote (claim record, else attribution).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3212
3213
3214
3215
3216
3217
3218
3219
3220
3221
3222
def get_job_page_ids(self, slug: str, job_id: str) -> frozenset[str]:
    """Return the page ids *job_id* wrote (claim record, else attribution)."""
    doc = self._col("wiki_jobs").find_one(
        {"job_id": job_id}, {"submitted_page_ids": 1}
    )
    if doc is None:
        return frozenset()
    ids = doc.get("submitted_page_ids")
    if ids is None:
        return self.page_ids_for_job(slug, job_id)
    return frozenset(str(p) for p in ids)

get_job_plan(job_id: str) -> list[dict[str, Any]] | None

Return the page-plan list, or None if no plan has been committed yet.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3116
3117
3118
3119
3120
3121
3122
def get_job_plan(self, job_id: str) -> list[dict[str, Any]] | None:
    """Return the page-plan list, or None if no plan has been committed yet."""
    doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"plan": 1})
    if doc is None:
        return None
    plan = doc.get("plan")
    return plan if isinstance(plan, list) else None

get_job_session(job_id: str) -> str | None

Return the session_id attached to job_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3091
3092
3093
3094
3095
3096
3097
def get_job_session(self, job_id: str) -> str | None:
    """Return the session_id attached to *job_id*, or None."""
    doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"session_id": 1})
    if doc is None:
        return None
    val = doc.get("session_id")
    return str(val) if val else None

get_job_submission(job_id: str) -> dict[str, Any] | None

Return the persisted submission dict, or None if not yet saved.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3241
3242
3243
3244
3245
3246
3247
def get_job_submission(self, job_id: str) -> dict[str, Any] | None:
    """Return the persisted submission dict, or None if not yet saved."""
    doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"submission": 1})
    if doc is None:
        return None
    val = doc.get("submission")
    return val if isinstance(val, dict) else None

get_job_submitted_count(job_id: str) -> int

Return the number of pages submitted so far for job_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3154
3155
3156
3157
3158
3159
def get_job_submitted_count(self, job_id: str) -> int:
    """Return the number of pages submitted so far for *job_id*."""
    doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"submitted_pages": 1})
    if doc is None:
        return 0
    return int(doc.get("submitted_pages", 0))

get_memory_node(slug: str, node_id: str) -> MemoryNode | None

Return a single memory node, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3897
3898
3899
3900
3901
3902
def get_memory_node(self, slug: str, node_id: str) -> MemoryNode | None:
    """Return a single memory node, or None if absent."""
    doc = self._col("wiki_memory_nodes").find_one({"slug": slug, "node_id": node_id})
    if doc is None:
        return None
    return MemoryNode.model_validate(_strip_mongo_meta(doc))

get_project(slug: str) -> Project | None

Return the project for slug, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2828
2829
2830
2831
2832
2833
def get_project(self, slug: str) -> Project | None:
    """Return the project for *slug*, or None if absent."""
    doc = self._col("wiki_projects").find_one({"slug": slug})
    if doc is None:
        return None
    return Project.model_validate(_strip_mongo_meta(doc))

get_project_settings(slug: str) -> ProjectSettings | None

Return slug's settings record, or None when never written.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2853
2854
2855
2856
2857
2858
2859
2860
2861
2862
2863
2864
2865
def get_project_settings(self, slug: str) -> ProjectSettings | None:
    """Return *slug*'s settings record, or None when never written."""
    doc = self._col("wiki_settings").find_one({"slug": slug})
    if doc is None:
        return None
    try:
        return ProjectSettings.model_validate(_strip_mongo_meta(doc))
    except Exception:
        # A hand-edited / off-schema document must not break the read path —
        # the caller then falls back to the per-job submission scan,
        # exactly as it does for a project that has no record at all.
        logging.warning("Skipping malformed wiki_settings document for {}", slug)
        return None

get_qa(answer_id: str) -> QaAnswer | None

Return the QA answer, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3330
3331
3332
3333
3334
3335
def get_qa(self, answer_id: str) -> QaAnswer | None:
    """Return the QA answer, or None if absent."""
    doc = self._col("wiki_qa").find_one({"answer_id": answer_id})
    if doc is None:
        return None
    return QaAnswer.model_validate(_clean_for_model(doc, QaAnswer))

get_qa_session(answer_id: str) -> str | None

Return the session_id attached to answer_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3352
3353
3354
3355
3356
3357
3358
def get_qa_session(self, answer_id: str) -> str | None:
    """Return the session_id attached to *answer_id*, or None."""
    doc = self._col("wiki_qa").find_one({"answer_id": answer_id}, {"session_id": 1})
    if doc is None:
        return None
    val = doc.get("session_id")
    return str(val) if val else None

get_recovery_attempts(slug: str) -> int

Return the recovery-attempt count for slug (0 if never recovered).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3289
3290
3291
3292
def get_recovery_attempts(self, slug: str) -> int:
    """Return the recovery-attempt count for *slug* (0 if never recovered)."""
    doc = self._col("wiki_recovery").find_one({"slug": slug}, {"attempts": 1})
    return int(doc.get("attempts", 0)) if doc else 0

get_resume_plan(job_id: str) -> dict[str, Any] | None

Return the persisted resume-plan dict, or None if the job isn't resuming.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3131
3132
3133
3134
3135
3136
3137
def get_resume_plan(self, job_id: str) -> dict[str, Any] | None:
    """Return the persisted resume-plan dict, or None if the job isn't resuming."""
    doc = self._col("wiki_jobs").find_one({"job_id": job_id}, {"resume_plan": 1})
    if doc is None:
        return None
    val = doc.get("resume_plan")
    return val if isinstance(val, dict) else None

list_credentials() -> dict[str, dict[str, Any]]

Return every stored credential blob keyed by scope (full scan).

Prefers the blob's own scope field (stamped by CredentialStore.save, matching the JSON driver's precedence); falls back to the document's slug key for a blob missing that field.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3270
3271
3272
3273
3274
3275
3276
3277
3278
3279
3280
3281
3282
3283
3284
3285
def list_credentials(self) -> dict[str, dict[str, Any]]:
    """Return every stored credential blob keyed by scope (full scan).

    Prefers the blob's own ``scope`` field (stamped by
    ``CredentialStore.save``, matching the JSON driver's precedence); falls
    back to the document's ``slug`` key for a blob missing that field.
    """
    out: dict[str, dict[str, Any]] = {}
    for doc in self._col("wiki_credentials").find({}, {"slug": 1, "blob": 1}):
        slug = doc.get("slug")
        blob = doc.get("blob")
        if isinstance(blob, dict):
            scope = blob["scope"] if isinstance(blob.get("scope"), str) else slug
            if isinstance(scope, str):
                out[scope] = blob
    return out

list_doc_notes(slug: str) -> list[DocPageNote]

Return every doc-page note for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4035
4036
4037
4038
4039
4040
def list_doc_notes(self, slug: str) -> list[DocPageNote]:
    """Return every doc-page note for *slug*."""
    return [
        DocPageNote.model_validate(_strip_mongo_meta(d))
        for d in self._col("wiki_doc_notes").find({"slug": slug})
    ]

list_edges(slug: str, *, scope: CommitScope) -> list[GraphEdge]

See WikiStoreBase.list_edges.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3638
3639
3640
3641
def list_edges(self, slug: str, *, scope: CommitScope) -> list[GraphEdge]:
    """See ``WikiStoreBase.list_edges``."""
    cursor = self._col("wiki_graph_edges").find({"slug": slug, **scope.filter_fields()})
    return [GraphEdge.model_validate(_strip_mongo_meta(d)) for d in cursor]

list_entity_edges(slug: str, *, source_id: str | None = None) -> list[EntityRelation]

Return entity relations, optionally scoped to source_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4188
4189
4190
4191
4192
4193
4194
4195
4196
4197
4198
4199
4200
def list_entity_edges(
    self, slug: str, *, source_id: str | None = None
) -> list[EntityRelation]:
    """Return entity relations, optionally scoped to ``source_id``."""
    query: dict[str, Any] = {"slug": slug}
    if source_id is not None:
        query["source_id"] = source_id
    out: list[EntityRelation] = []
    for d in self._col("wiki_entity_edges").find(query):
        clean = _strip_mongo_meta(d)
        clean.pop("slug", None)
        out.append(EntityRelation.model_validate(clean))
    return out

list_file_manifest(slug: str) -> list[FileManifest]

Return every file-manifest entry for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4071
4072
4073
4074
4075
4076
def list_file_manifest(self, slug: str) -> list[FileManifest]:
    """Return every file-manifest entry for *slug*."""
    return [
        FileManifest.model_validate(_strip_mongo_meta(d))
        for d in self._col("wiki_file_manifest").find({"slug": slug})
    ]

list_jobs(slug: str | None = None) -> list[IndexingJob]

Return all jobs, newest first by phase_started_at, filtered to slug.

Mongo find() has no inherent order, and sorting by job_id for reproducibility would not fix that: job_id is a uuid4 hex, which sorts RANDOMLY with respect to when a job actually ran, and at least one caller (resolve_qa_clone_dir) assumes this returns most-recent-first and picks the FIRST complete hit — under such a sort an arbitrary old checkout. Sorted in Python rather than via a Mongo-side .sort() so both drivers apply the IDENTICAL rule (ISO-8601 phase_started_at, so lexicographic == chronological; a job with no timestamp sorts last) instead of relying on each backend's own null-ordering semantics to happen to agree.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3033
3034
3035
3036
3037
3038
3039
3040
3041
3042
3043
3044
3045
3046
3047
3048
3049
3050
3051
3052
3053
def list_jobs(self, slug: str | None = None) -> list[IndexingJob]:
    """Return all jobs, newest first by ``phase_started_at``, filtered to *slug*.

    Mongo ``find()`` has no inherent order, and sorting by ``job_id`` for
    reproducibility would not fix that: ``job_id`` is a ``uuid4`` hex, which
    sorts RANDOMLY with respect to when a job actually ran, and at least one
    caller (``resolve_qa_clone_dir``) assumes this returns most-recent-first
    and picks the FIRST ``complete`` hit — under such a sort an arbitrary
    old checkout. Sorted in Python
    rather than via a Mongo-side ``.sort()`` so both drivers apply the
    IDENTICAL rule (ISO-8601 ``phase_started_at``, so lexicographic ==
    chronological; a job with no timestamp sorts last) instead of relying
    on each backend's own null-ordering semantics to happen to agree.
    """
    query: dict[str, Any] = {}
    if slug is not None:
        query["slug"] = slug
    jobs: list[IndexingJob] = []
    for doc in self._col("wiki_jobs").find(query):
        jobs.append(IndexingJob.model_validate(_clean_for_model(doc, IndexingJob)))
    return sorted(jobs, key=lambda j: j.phase_started_at or "", reverse=True)

list_memory_edges(slug: str, *, node_id: str | None = None, include_invalidated: bool = False) -> list[MemoryEdge]

Return memory edges, optionally scoped to source == node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3938
3939
3940
3941
3942
3943
3944
3945
3946
3947
3948
3949
3950
3951
3952
3953
3954
def list_memory_edges(
    self,
    slug: str,
    *,
    node_id: str | None = None,
    include_invalidated: bool = False,
) -> list[MemoryEdge]:
    """Return memory edges, optionally scoped to ``source == node_id``."""
    query: dict[str, Any] = {"slug": slug}
    if node_id is not None:
        query["source"] = node_id
    if not include_invalidated:
        query["invalid_at"] = None
    return [
        MemoryEdge.model_validate(_strip_mongo_meta(d))
        for d in self._col("wiki_memory_edges").find(query)
    ]

list_pages(slug: str) -> list[WikiPage]

Return all pages for project slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2967
2968
2969
2970
2971
2972
2973
2974
2975
2976
def list_pages(self, slug: str) -> list[WikiPage]:
    """Return all pages for project *slug*."""
    pages: list[WikiPage] = []
    for doc in self._col("wiki_pages").find({"slug": slug}):
        clean = _strip_mongo_meta(doc)
        for col in self._PAGE_STORE_COLS:
            clean.pop(col, None)
        page = WikiPage.model_validate(clean)
        pages.append(page)
    return pages

list_projects() -> list[Project]

Return all projects sorted by indexed_at descending.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2835
2836
2837
2838
def list_projects(self) -> list[Project]:
    """Return all projects sorted by indexed_at descending."""
    cursor = self._col("wiki_projects").find().sort("indexed_at", -1)
    return [Project.model_validate(_strip_mongo_meta(d)) for d in cursor]

list_qa(status: str | None = None) -> list[QaAnswer]

Return all QA answers, optionally filtered to status (server-side).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3337
3338
3339
3340
3341
3342
3343
def list_qa(self, status: str | None = None) -> list[QaAnswer]:
    """Return all QA answers, optionally filtered to *status* (server-side)."""
    query: dict[str, Any] = {} if status is None else {"status": status}
    return [
        QaAnswer.model_validate(_clean_for_model(doc, QaAnswer))
        for doc in self._col("wiki_qa").find(query)
    ]

load_job_events(job_id: str, after_idx: int = -1) -> list[dict[str, Any]]

Return job events with idx > after_idx (-1 returns all).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3061
3062
3063
3064
3065
3066
3067
3068
3069
3070
3071
def load_job_events(
    self, job_id: str, after_idx: int = -1
) -> list[dict[str, Any]]:
    """Return job events with idx > *after_idx* (-1 returns all)."""
    query: dict[str, Any] = {"job_id": job_id, "idx": {"$gt": after_idx}}
    results: list[dict[str, Any]] = []
    for doc in self._col("wiki_job_events").find(query).sort("idx", 1):
        clean = _strip_mongo_meta(doc)
        clean.pop("job_id", None)
        results.append(clean)
    return results

load_qa_events(answer_id: str, after_idx: int = -1) -> list[dict[str, Any]]

Return QA events with idx > after_idx (-1 returns all).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3378
3379
3380
3381
3382
3383
3384
3385
3386
3387
3388
def load_qa_events(
    self, answer_id: str, after_idx: int = -1
) -> list[dict[str, Any]]:
    """Return QA events with idx > *after_idx* (-1 returns all)."""
    query: dict[str, Any] = {"answer_id": answer_id, "idx": {"$gt": after_idx}}
    results: list[dict[str, Any]] = []
    for doc in self._col("wiki_qa_events").find(query).sort("idx", 1):
        clean = _strip_mongo_meta(doc)
        clean.pop("answer_id", None)
        results.append(clean)
    return results

memories_anchored_to(slug: str, entity_keys: Iterable[EntityKey], *, include_invalidated: bool = False) -> list[str]

Reverse ANCHORS lookup: entity_keys → distinct memory node_ids.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3956
3957
3958
3959
3960
3961
3962
3963
3964
3965
3966
3967
3968
3969
3970
3971
3972
3973
3974
3975
3976
def memories_anchored_to(
    self,
    slug: str,
    entity_keys: Iterable[EntityKey],
    *,
    include_invalidated: bool = False,
) -> list[str]:
    """Reverse ANCHORS lookup: entity_keys → distinct memory node_ids."""
    query: dict[str, Any] = {
        "slug": slug, "type": "ANCHORS", "target": {"$in": list(entity_keys)}
    }
    if not include_invalidated:
        query["invalid_at"] = None
    seen: list[str] = []
    seen_set: set[str] = set()
    for d in self._col("wiki_memory_edges").find(query):
        src = d.get("source")
        if src not in seen_set:
            seen_set.add(src)
            seen.append(src)
    return seen

Top-k memory embeddings by cosine, after applying filt.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4000
4001
4002
4003
4004
4005
4006
4007
4008
4009
4010
4011
4012
4013
def memory_vector_search(
    self,
    slug: str,
    qvec: list[float],
    k: int = 10,
    *,
    filt: MemoryFilter | None = None,
) -> list[MemoryEmbedding]:
    """Top-k memory embeddings by cosine, after applying *filt*."""
    pool = [
        MemoryEmbedding.model_validate(_strip_mongo_meta(d))
        for d in self._col("wiki_memory_embeddings").find({"slug": slug})
    ]
    return self._rank_memory(slug, pool, qvec, k, filt)

page_ids_for_job(slug: str, job_id: str) -> frozenset[str]

Page ids whose stored attribution column names job_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3224
3225
3226
3227
3228
3229
3230
3231
3232
def page_ids_for_job(self, slug: str, job_id: str) -> frozenset[str]:
    """Page ids whose stored attribution column names *job_id*."""
    return frozenset(
        str(d["page_id"])
        for d in self._col("wiki_pages").find(
            {"slug": slug, "job_id": job_id}, {"page_id": 1}
        )
        if d.get("page_id")
    )

prune_pages(slug: str, keep: Iterable[str]) -> int

Bulk-drop pages not in keep in a single Mongo round-trip.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2985
2986
2987
2988
2989
2990
2991
def prune_pages(self, slug: str, keep: Iterable[str]) -> int:
    """Bulk-drop pages not in *keep* in a single Mongo round-trip."""
    keep_list = list(keep)
    result = self._col("wiki_pages").delete_many(
        {"slug": slug, "page_id": {"$nin": keep_list}}
    )
    return int(result.deleted_count)

query_entities(slug: str, *, filt: EntityFilter | None = None) -> list[Entity]

Return entities matching filt's facets.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4118
4119
4120
4121
4122
4123
4124
4125
4126
4127
4128
4129
def query_entities(
    self, slug: str, *, filt: EntityFilter | None = None
) -> list[Entity]:
    """Return entities matching *filt*'s facets."""
    out: list[Entity] = []
    for doc in self._col("wiki_entities").find({"slug": slug}):
        clean = _strip_mongo_meta(doc)
        clean.pop("slug", None)
        out.append(Entity.model_validate(clean))
    if filt is None:
        return out
    return [e for e in out if filt.matches(e)]

query_graph(slug: str, *, scope: CommitScope, node_type: str | None = None, name_match: str | None = None, neighbors_of: str | None = None, node_ids: Collection[str] | None = None) -> list[GraphNode]

See WikiStoreBase.query_graph.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3593
3594
3595
3596
3597
3598
3599
3600
3601
3602
3603
3604
3605
3606
3607
3608
3609
3610
3611
3612
3613
3614
3615
3616
3617
3618
3619
3620
3621
3622
3623
3624
3625
3626
3627
3628
3629
3630
3631
3632
3633
3634
3635
3636
def query_graph(
    self,
    slug: str,
    *,
    scope: CommitScope,
    node_type: str | None = None,
    name_match: str | None = None,
    neighbors_of: str | None = None,
    node_ids: Collection[str] | None = None,
) -> list[GraphNode]:
    """See ``WikiStoreBase.query_graph``."""
    import re

    commit = scope.filter_fields()
    if neighbors_of is not None:
        # The seed's own generation bounds the walk — see the JSON driver.
        edge_query = {
            "slug": slug,
            "$or": [{"source": neighbors_of}, {"target": neighbors_of}],
            **commit,
        }
        related_ids: set[str] = set()
        for edge_doc in self._col("wiki_graph_edges").find(edge_query):
            src = edge_doc.get("source")
            tgt = edge_doc.get("target")
            if src == neighbors_of:
                related_ids.add(tgt)
            else:
                related_ids.add(src)
        if not related_ids:
            return []
        cursor = self._col("wiki_graph_nodes").find(
            {"slug": slug, "node_id": {"$in": list(related_ids)}, **commit}
        )
        return [GraphNodeAdapter.validate_python(_strip_mongo_meta(d)) for d in cursor]
    query: dict[str, Any] = {"slug": slug, **commit}
    if node_ids is not None:
        query["node_id"] = {"$in": list(node_ids)}
    if node_type is not None:
        query["type"] = node_type
    if name_match is not None:
        query["name"] = {"$regex": re.escape(name_match), "$options": "i"}
    cursor = self._col("wiki_graph_nodes").find(query)
    return [GraphNodeAdapter.validate_python(_strip_mongo_meta(d)) for d in cursor]

query_memory(slug: str, *, filt: MemoryFilter | None = None) -> list[MemoryNode]

Return memory nodes matching filt's node-level facets.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3914
3915
3916
3917
3918
3919
3920
3921
3922
3923
3924
def query_memory(
    self, slug: str, *, filt: MemoryFilter | None = None
) -> list[MemoryNode]:
    """Return memory nodes matching *filt*'s node-level facets."""
    nodes = [
        MemoryNode.model_validate(_strip_mongo_meta(d))
        for d in self._col("wiki_memory_nodes").find({"slug": slug})
    ]
    if filt is None:
        return nodes
    return [n for n in nodes if filt.matches_node(n)]

reap_slug(slug: str) -> dict[str, int]

See WikiStoreBase.reap_slug.

wiki_job_events/wiki_qa_events carry only their owning id (job_id/answer_id), never slug — so those ids are read from wiki_jobs/wiki_qa FIRST, before either collection is touched, and the two event collections are then swept by id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2873
2874
2875
2876
2877
2878
2879
2880
2881
2882
2883
2884
2885
2886
2887
2888
2889
2890
2891
2892
2893
2894
2895
2896
2897
2898
2899
2900
2901
2902
2903
2904
2905
2906
2907
2908
2909
2910
2911
2912
2913
2914
2915
2916
2917
2918
2919
2920
2921
2922
2923
2924
2925
2926
2927
2928
2929
def reap_slug(self, slug: str) -> dict[str, int]:
    """See ``WikiStoreBase.reap_slug``.

    ``wiki_job_events``/``wiki_qa_events`` carry only their owning id
    (``job_id``/``answer_id``), never ``slug`` — so those ids are read from
    ``wiki_jobs``/``wiki_qa`` FIRST, before either collection is touched,
    and the two event collections are then swept by id.
    """
    job_ids = [
        str(d["job_id"])
        for d in self._col("wiki_jobs").find({"slug": slug}, {"job_id": 1})
    ]
    answer_ids = [
        str(d["answer_id"])
        for d in self._col("wiki_qa").find({"slug": slug}, {"answer_id": 1})
    ]

    counts: dict[str, int] = {}
    for key, coll in (
        ("graph_nodes", "wiki_graph_nodes"),
        ("graph_edges", "wiki_graph_edges"),
        ("embeddings", "wiki_embeddings"),
        ("entities", "wiki_entities"),
        ("entity_edges", "wiki_entity_edges"),
        ("entity_embeddings", "wiki_entity_embeddings"),
        ("entity_recommendations", "wiki_entity_recommendations"),
        ("memory_nodes", "wiki_memory_nodes"),
        ("memory_edges", "wiki_memory_edges"),
        ("memory_embeddings", "wiki_memory_embeddings"),
        ("doc_notes", "wiki_doc_notes"),
        ("file_manifest", "wiki_file_manifest"),
        ("pages", "wiki_pages"),
        ("recovery", "wiki_recovery"),
        ("jobs", "wiki_jobs"),
        ("qa", "wiki_qa"),
    ):
        counts[key] = int(self._col(coll).delete_many({"slug": slug}).deleted_count)

    counts["job_events"] = (
        int(
            self._col("wiki_job_events")
            .delete_many({"job_id": {"$in": job_ids}})
            .deleted_count
        )
        if job_ids
        else 0
    )
    counts["qa_events"] = (
        int(
            self._col("wiki_qa_events")
            .delete_many({"answer_id": {"$in": answer_ids}})
            .deleted_count
        )
        if answer_ids
        else 0
    )
    return counts

reset_recovery_attempts(slug: str) -> None

Clear slug's recovery counter document (user-initiated resume fresh budget).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3306
3307
3308
def reset_recovery_attempts(self, slug: str) -> None:
    """Clear *slug*'s recovery counter document (user-initiated resume fresh budget)."""
    self._col("wiki_recovery").delete_one({"slug": slug})

restamp_graph_artifacts(slug: str, *, from_commit: str, to_commit: str) -> dict[str, int]

See WikiStoreBase.restamp_graph_artifacts.

An EXACT commit_sha match, deliberately not the $nin shape supersede_graph_artifacts uses: supersede asks "everything except the keeper", which would here sweep up older generations and field-absent rows and stamp genuinely dead code as live.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3676
3677
3678
3679
3680
3681
3682
3683
3684
3685
3686
3687
3688
3689
3690
3691
3692
3693
3694
3695
3696
3697
3698
3699
3700
def restamp_graph_artifacts(
    self, slug: str, *, from_commit: str, to_commit: str
) -> dict[str, int]:
    """See ``WikiStoreBase.restamp_graph_artifacts``.

    An EXACT ``commit_sha`` match, deliberately not the ``$nin`` shape
    ``supersede_graph_artifacts`` uses: supersede asks "everything except
    the keeper", which would here sweep up older generations and
    field-absent rows and stamp genuinely dead code as live.
    """
    self._ensure_graph_indexes()
    self._ensure_memory_indexes()
    prior = {"slug": slug, "commit_sha": from_commit}
    patch = {"$set": {"commit_sha": to_commit}}
    counts: dict[str, int] = {}
    for key, coll in (
        ("nodes", "wiki_graph_nodes"),
        ("edges", "wiki_graph_edges"),
        ("embeddings", "wiki_embeddings"),
        ("entities", "wiki_entities"),
        ("entity_edges", "wiki_entity_edges"),
        ("entity_embeddings", "wiki_entity_embeddings"),
    ):
        counts[key] = int(self._col(coll).update_many(prior, patch).modified_count)
    return counts

save_act_plan(job_id: str, plan: dict[str, Any]) -> None

Persist the act-stage record on the job doc; overwrites any previous one.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3139
3140
3141
3142
3143
3144
def save_act_plan(self, job_id: str, plan: dict[str, Any]) -> None:
    """Persist the act-stage record on the job doc; overwrites any previous one."""
    self._col("wiki_jobs").update_one(
        {"job_id": job_id},
        {"$set": {"act_plan": plan}},
    )

save_credentials(slug: str, blob: dict[str, Any]) -> None

Persist the encoded credential blob for slug (slug PK, upsert).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3251
3252
3253
3254
3255
def save_credentials(self, slug: str, blob: dict[str, Any]) -> None:
    """Persist the encoded credential blob for *slug* (slug PK, upsert)."""
    self._col("wiki_credentials").replace_one(
        {"slug": slug}, {"slug": slug, "blob": blob}, upsert=True
    )

save_entity_recommendation(slug: str, rec: EntityRecommendation) -> None

Upsert a recommendation for slug; dedup by (slug, id).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4202
4203
4204
4205
4206
4207
4208
4209
4210
4211
def save_entity_recommendation(
    self, slug: str, rec: EntityRecommendation
) -> None:
    """Upsert a recommendation for *slug*; dedup by (slug, id)."""
    self._ensure_memory_indexes()
    self._col("wiki_entity_recommendations").update_one(
        {"slug": slug, "id": rec.id},
        {"$set": {"slug": slug, **rec.model_dump(by_alias=False)}},
        upsert=True,
    )

save_job_plan(job_id: str, plan: list[dict[str, Any]]) -> None

Persist the page-plan list for job_id; overwrites any previous plan.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3109
3110
3111
3112
3113
3114
def save_job_plan(self, job_id: str, plan: list[dict[str, Any]]) -> None:
    """Persist the page-plan list for *job_id*; overwrites any previous plan."""
    self._col("wiki_jobs").update_one(
        {"job_id": job_id},
        {"$set": {"plan": plan}},
    )

save_job_submission(job_id: str, submission: dict[str, Any]) -> None

Persist the wizard submission dict for job_id (token must be absent).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3234
3235
3236
3237
3238
3239
def save_job_submission(self, job_id: str, submission: dict[str, Any]) -> None:
    """Persist the wizard submission dict for *job_id* (token must be absent)."""
    self._col("wiki_jobs").update_one(
        {"job_id": job_id},
        {"$set": {"submission": submission}},
    )

save_page(slug: str, page: WikiPage, *, commit_sha: str | None = None, job_id: str | None = None) -> None

Persist page for the project slug; overwrites if same page_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2937
2938
2939
2940
2941
2942
2943
2944
2945
2946
2947
2948
2949
2950
2951
2952
2953
2954
2955
def save_page(
    self,
    slug: str,
    page: WikiPage,
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Persist *page* for the project *slug*; overwrites if same page_id."""
    doc = {
        "slug": slug,
        "page_id": page.id,
        "commit_sha": commit_sha,
        "job_id": job_id,
        **page.model_dump(by_alias=False),
    }
    self._col("wiki_pages").replace_one(
        {"slug": slug, "page_id": page.id}, doc, upsert=True
    )

save_project_settings(slug: str, settings: ProjectSettings) -> None

Persist (upsert) the editable settings record for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
2847
2848
2849
2850
2851
def save_project_settings(self, slug: str, settings: ProjectSettings) -> None:
    """Persist (upsert) the editable settings record for *slug*."""
    doc = settings.model_dump(by_alias=False)
    doc["slug"] = slug
    self._col("wiki_settings").replace_one({"slug": slug}, doc, upsert=True)

save_qa(answer: QaAnswer) -> None

Persist a QA answer record (creation: resets event_count, no session yet).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3312
3313
3314
3315
3316
3317
def save_qa(self, answer: QaAnswer) -> None:
    """Persist a QA answer record (creation: resets event_count, no session yet)."""
    doc = {"event_count": 0, **answer.model_dump(by_alias=False)}
    self._col("wiki_qa").replace_one(
        {"answer_id": answer.answer_id}, doc, upsert=True
    )

save_resume_plan(job_id: str, plan: dict[str, Any]) -> None

Persist the resume-plan dict on the job doc; overwrites any previous one.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3124
3125
3126
3127
3128
3129
def save_resume_plan(self, job_id: str, plan: dict[str, Any]) -> None:
    """Persist the resume-plan dict on the job doc; overwrites any previous one."""
    self._col("wiki_jobs").update_one(
        {"job_id": job_id},
        {"$set": {"resume_plan": plan}},
    )

supersede_graph_artifacts(slug: str, *, keep_commit_sha: str) -> dict[str, int]

Drop prior-commit graph + entity artifacts, preserving None-stamped rows.

{"$nin": [None, keep]} matches a REAL commit other than keep while leaving both None-valued and field-absent rows untouched — the QA-minted / pre-isolation records supersede must not reap.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3652
3653
3654
3655
3656
3657
3658
3659
3660
3661
3662
3663
3664
3665
3666
3667
3668
3669
3670
3671
3672
3673
3674
def supersede_graph_artifacts(
    self, slug: str, *, keep_commit_sha: str
) -> dict[str, int]:
    """Drop prior-commit graph + entity artifacts, preserving ``None``-stamped rows.

    ``{"$nin": [None, keep]}`` matches a REAL commit other than *keep* while
    leaving both ``None``-valued and field-absent rows untouched — the
    QA-minted / pre-isolation records supersede must not reap.
    """
    self._ensure_graph_indexes()
    self._ensure_memory_indexes()
    stale = {"slug": slug, "commit_sha": {"$nin": [None, keep_commit_sha]}}
    counts: dict[str, int] = {}
    for key, coll in (
        ("nodes", "wiki_graph_nodes"),
        ("edges", "wiki_graph_edges"),
        ("embeddings", "wiki_embeddings"),
        ("entities", "wiki_entities"),
        ("entity_edges", "wiki_entity_edges"),
        ("entity_embeddings", "wiki_entity_embeddings"),
    ):
        counts[key] = int(self._col(coll).delete_many(stale).deleted_count)
    return counts

update_qa_fields(answer: QaAnswer) -> None

In-place $set of the QaAnswer fields only.

Leaves event_count + session_id (this backend packs both into the same doc) intact, unlike save_qa's full replace.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3319
3320
3321
3322
3323
3324
3325
3326
3327
3328
def update_qa_fields(self, answer: QaAnswer) -> None:
    """In-place ``$set`` of the QaAnswer fields only.

    Leaves ``event_count`` + ``session_id`` (this backend packs both into the
    same doc) intact, unlike ``save_qa``'s full replace.
    """
    self._col("wiki_qa").update_one(
        {"answer_id": answer.answer_id},
        {"$set": answer.model_dump(by_alias=False)},
    )

upsert_doc_notes(slug: str, notes: Iterable[DocPageNote]) -> None

Upsert doc-page notes for slug; dedup by (slug, page_id).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4017
4018
4019
4020
4021
4022
4023
4024
4025
4026
def upsert_doc_notes(self, slug: str, notes: Iterable[DocPageNote]) -> None:
    """Upsert doc-page notes for *slug*; dedup by (slug, page_id)."""
    self._ensure_memory_indexes()
    col = self._col("wiki_doc_notes")
    for note in notes:
        col.update_one(
            {"slug": slug, "page_id": note.page_id},
            {"$set": note.model_dump(by_alias=False)},
            upsert=True,
        )

upsert_edges(slug: str, edges: Iterable[GraphEdge], *, commit_sha: str | None = None, job_id: str | None = None, on_progress: Callable[[int, int], None] | None = None) -> None

Upsert graph edges for slug; dedup by (slug, source, target, type).

Cost: O(edges) — offline, the graph phase. Batched through _bulk_upsert, so the round-trip count is O(edges / batch).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3519
3520
3521
3522
3523
3524
3525
3526
3527
3528
3529
3530
3531
3532
3533
3534
3535
3536
3537
3538
3539
3540
3541
3542
3543
3544
3545
3546
3547
3548
3549
3550
3551
3552
3553
3554
3555
def upsert_edges(
    self,
    slug: str,
    edges: Iterable[GraphEdge],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
    on_progress: Callable[[int, int], None] | None = None,
) -> None:
    """Upsert graph edges for *slug*; dedup by (slug, source, target, type).

    Cost: ``O(edges)`` — offline, the ``graph`` phase. Batched through
    ``_bulk_upsert``, so the round-trip count is ``O(edges / batch)``.
    """
    self._ensure_graph_indexes()
    total_batches = (
        math.ceil(len(edges) / self._BULK_BATCH_SIZE)
        if isinstance(edges, Collection)
        else None
    )
    self._bulk_upsert(
        self._col("wiki_graph_edges"),
        (
            (
                {
                    "slug": slug,
                    "source": edge.source,
                    "target": edge.target,
                    "type": edge.type,
                },
                self._stamp_attribution(edge, commit_sha, job_id).model_dump(by_alias=False),
            )
            for edge in edges
        ),
        total_batches=total_batches,
        on_progress=on_progress,
    )

upsert_embeddings(slug: str, items: Iterable[Embedding], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert embedding vectors for slug; dedup by (slug, node_id).

Cost: O(vectors) — offline, the graph phase. Batched through _bulk_upsert, so the round-trip count is O(vectors / batch).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3557
3558
3559
3560
3561
3562
3563
3564
3565
3566
3567
3568
3569
3570
3571
3572
3573
3574
3575
3576
3577
3578
3579
3580
def upsert_embeddings(
    self,
    slug: str,
    items: Iterable[Embedding],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert embedding vectors for *slug*; dedup by (slug, node_id).

    Cost: ``O(vectors)`` — offline, the ``graph`` phase. Batched through
    ``_bulk_upsert``, so the round-trip count is ``O(vectors / batch)``.
    """
    self._ensure_graph_indexes()
    self._bulk_upsert(
        self._col("wiki_embeddings"),
        (
            (
                {"slug": slug, "node_id": item.node_id},
                self._embedding_doc(item, commit_sha, job_id),
            )
            for item in items
        ),
    )

upsert_entities(slug: str, entities: Iterable[Entity], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert entities for slug; dedup by (slug, id), stamp attribution.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4090
4091
4092
4093
4094
4095
4096
4097
4098
4099
4100
4101
4102
4103
4104
4105
4106
4107
def upsert_entities(
    self,
    slug: str,
    entities: Iterable[Entity],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert entities for *slug*; dedup by (slug, id), stamp attribution."""
    self._ensure_memory_indexes()
    col = self._col("wiki_entities")
    for entity in entities:
        stamped = self._stamp_attribution(entity, commit_sha, job_id)
        col.update_one(
            {"slug": slug, "id": entity.id},
            {"$set": {"slug": slug, **stamped.model_dump(by_alias=False)}},
            upsert=True,
        )

upsert_entity_edges(slug: str, edges: Iterable[EntityRelation], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert entity relations for slug; dedup by (slug, id).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4169
4170
4171
4172
4173
4174
4175
4176
4177
4178
4179
4180
4181
4182
4183
4184
4185
4186
def upsert_entity_edges(
    self,
    slug: str,
    edges: Iterable[EntityRelation],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert entity relations for *slug*; dedup by (slug, id)."""
    self._ensure_memory_indexes()
    col = self._col("wiki_entity_edges")
    for edge in edges:
        stamped = self._stamp_attribution(edge, commit_sha, job_id)
        col.update_one(
            {"slug": slug, "id": edge.id},
            {"$set": {"slug": slug, **stamped.model_dump(by_alias=False)}},
            upsert=True,
        )

upsert_entity_embeddings(slug: str, items: Iterable[EntityEmbedding], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert entity embedding vectors for slug; dedup by (slug, entity_id).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4140
4141
4142
4143
4144
4145
4146
4147
4148
4149
4150
4151
4152
4153
4154
4155
4156
4157
def upsert_entity_embeddings(
    self,
    slug: str,
    items: Iterable[EntityEmbedding],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert entity embedding vectors for *slug*; dedup by (slug, entity_id)."""
    self._ensure_memory_indexes()
    col = self._col("wiki_entity_embeddings")
    for item in items:
        stamped = self._stamp_attribution(item, commit_sha, job_id)
        col.update_one(
            {"slug": slug, "entity_id": item.entity_id},
            {"$set": stamped.model_dump(by_alias=False)},
            upsert=True,
        )

upsert_file_manifest(slug: str, entries: Iterable[FileManifest]) -> None

Upsert file-manifest entries for slug; dedup by (slug, path).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4051
4052
4053
4054
4055
4056
4057
4058
4059
4060
4061
4062
def upsert_file_manifest(
    self, slug: str, entries: Iterable[FileManifest]
) -> None:
    """Upsert file-manifest entries for *slug*; dedup by (slug, path)."""
    self._ensure_memory_indexes()
    col = self._col("wiki_file_manifest")
    for entry in entries:
        col.update_one(
            {"slug": slug, "path": entry.path},
            {"$set": entry.model_dump(by_alias=False)},
            upsert=True,
        )

upsert_memory_edges(slug: str, edges: Iterable[MemoryEdge]) -> None

Upsert memory edges for slug; dedup by (slug, source, target, type).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3926
3927
3928
3929
3930
3931
3932
3933
3934
3935
3936
def upsert_memory_edges(self, slug: str, edges: Iterable[MemoryEdge]) -> None:
    """Upsert memory edges for *slug*; dedup by (slug, source, target, type)."""
    self._ensure_memory_indexes()
    col = self._col("wiki_memory_edges")
    for edge in edges:
        col.update_one(
            {"slug": slug, "source": edge.source, "target": edge.target,
             "type": edge.type},
            {"$set": edge.model_dump(by_alias=False)},
            upsert=True,
        )

upsert_memory_embeddings(slug: str, items: Iterable[MemoryEmbedding]) -> None

Upsert memory embedding vectors for slug; dedup by (slug, node_id).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3987
3988
3989
3990
3991
3992
3993
3994
3995
3996
3997
3998
def upsert_memory_embeddings(
    self, slug: str, items: Iterable[MemoryEmbedding]
) -> None:
    """Upsert memory embedding vectors for *slug*; dedup by (slug, node_id)."""
    self._ensure_memory_indexes()
    col = self._col("wiki_memory_embeddings")
    for item in items:
        col.update_one(
            {"slug": slug, "node_id": item.node_id},
            {"$set": item.model_dump(by_alias=False)},
            upsert=True,
        )

upsert_memory_nodes(slug: str, nodes: Iterable[MemoryNode]) -> None

Upsert memory nodes for slug; dedup by (slug, node_id).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3886
3887
3888
3889
3890
3891
3892
3893
3894
3895
def upsert_memory_nodes(self, slug: str, nodes: Iterable[MemoryNode]) -> None:
    """Upsert memory nodes for *slug*; dedup by (slug, node_id)."""
    self._ensure_memory_indexes()
    col = self._col("wiki_memory_nodes")
    for node in nodes:
        col.update_one(
            {"slug": slug, "node_id": node.node_id},
            {"$set": node.model_dump(by_alias=False)},
            upsert=True,
        )

upsert_nodes(slug: str, nodes: Iterable[GraphNode], *, commit_sha: str | None = None, job_id: str | None = None, on_progress: Callable[[int, int], None] | None = None) -> None

Upsert graph nodes for slug; dedup by (slug, node_id), stamp attribution.

Cost: O(nodes) — offline, the graph phase. Batched through _bulk_upsert, so the round-trip count is O(nodes / batch).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3486
3487
3488
3489
3490
3491
3492
3493
3494
3495
3496
3497
3498
3499
3500
3501
3502
3503
3504
3505
3506
3507
3508
3509
3510
3511
3512
3513
3514
3515
3516
3517
def upsert_nodes(
    self,
    slug: str,
    nodes: Iterable[GraphNode],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
    on_progress: Callable[[int, int], None] | None = None,
) -> None:
    """Upsert graph nodes for *slug*; dedup by (slug, node_id), stamp attribution.

    Cost: ``O(nodes)`` — offline, the ``graph`` phase. Batched through
    ``_bulk_upsert``, so the round-trip count is ``O(nodes / batch)``.
    """
    self._ensure_graph_indexes()
    total_batches = (
        math.ceil(len(nodes) / self._BULK_BATCH_SIZE)
        if isinstance(nodes, Collection)
        else None
    )
    self._bulk_upsert(
        self._col("wiki_graph_nodes"),
        (
            (
                {"slug": slug, "node_id": node.node_id},
                self._stamp_attribution(node, commit_sha, job_id).model_dump(by_alias=False),
            )
            for node in nodes
        ),
        total_batches=total_batches,
        on_progress=on_progress,
    )

Return top-k embeddings for slug by cosine similarity.

Cost: O(embeddings for the slug) — every stored vector is still scored, so this remains the documented scale seam. What the packed path removes is the per-element Python cost of getting there: it reads only node_id + the _VEC_F32 buffer, scores the whole project as one NumPy matrix product, and materialises exactly k Embedding models.

Falls back to the pure-Python scan when NumPy is absent or any row lacks _VEC_F32 — never to a PARTIAL result.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
3702
3703
3704
3705
3706
3707
3708
3709
3710
3711
3712
3713
3714
3715
3716
3717
3718
3719
3720
3721
3722
3723
3724
3725
3726
3727
def vector_search(self, slug: str, qvec: list[float], k: int = 10) -> list[Embedding]:
    """Return top-k embeddings for *slug* by cosine similarity.

    Cost: ``O(embeddings for the slug)`` — every stored vector is still
    scored, so this remains the documented scale seam. What the packed path
    removes is the per-element Python cost of getting there: it reads only
    ``node_id`` + the ``_VEC_F32`` buffer, scores the whole project as one
    NumPy matrix product, and materialises exactly *k* ``Embedding`` models.

    Falls back to the pure-Python scan when NumPy is absent or any row
    lacks ``_VEC_F32`` — never to a PARTIAL result.
    """
    from .embedder import Embedder

    packed = self._vector_search_packed(slug, qvec, k)
    if packed is not None:
        return packed
    pool = [
        Embedding.model_validate(_clean_for_model(d, Embedding))
        for d in self._col("wiki_embeddings").find({"slug": slug})
    ]
    if not pool:
        return []
    scored = [(emb, Embedder.cosine(qvec, emb.vector)) for emb in pool]
    scored.sort(key=lambda t: t[1], reverse=True)
    return [emb for emb, _ in scored[:k]]

PageClaim dataclass

Outcome of a job claiming one page id: the new count + whether it was new.

One answer rather than two reads. The read-then-increment this replaces asked the store twice — "does this page exist" and then "bump the counter" — so two page-writers landing together could both see "new" and over-count, and a caller that saw "not new" had to go back for the count.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
@dataclass(frozen=True)
class PageClaim:
    """Outcome of a job claiming one page id: the new count + whether it was new.

    One answer rather than two reads. The read-then-increment this replaces
    asked the store twice — "does this page exist" and then "bump the counter" —
    so two page-writers landing together could both see "new" and over-count,
    and a caller that saw "not new" had to go back for the count.
    """

    count: int
    is_new: bool

WikiStoreBase

Bases: ABC

Abstract base for wiki persistence backends.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
 159
 160
 161
 162
 163
 164
 165
 166
 167
 168
 169
 170
 171
 172
 173
 174
 175
 176
 177
 178
 179
 180
 181
 182
 183
 184
 185
 186
 187
 188
 189
 190
 191
 192
 193
 194
 195
 196
 197
 198
 199
 200
 201
 202
 203
 204
 205
 206
 207
 208
 209
 210
 211
 212
 213
 214
 215
 216
 217
 218
 219
 220
 221
 222
 223
 224
 225
 226
 227
 228
 229
 230
 231
 232
 233
 234
 235
 236
 237
 238
 239
 240
 241
 242
 243
 244
 245
 246
 247
 248
 249
 250
 251
 252
 253
 254
 255
 256
 257
 258
 259
 260
 261
 262
 263
 264
 265
 266
 267
 268
 269
 270
 271
 272
 273
 274
 275
 276
 277
 278
 279
 280
 281
 282
 283
 284
 285
 286
 287
 288
 289
 290
 291
 292
 293
 294
 295
 296
 297
 298
 299
 300
 301
 302
 303
 304
 305
 306
 307
 308
 309
 310
 311
 312
 313
 314
 315
 316
 317
 318
 319
 320
 321
 322
 323
 324
 325
 326
 327
 328
 329
 330
 331
 332
 333
 334
 335
 336
 337
 338
 339
 340
 341
 342
 343
 344
 345
 346
 347
 348
 349
 350
 351
 352
 353
 354
 355
 356
 357
 358
 359
 360
 361
 362
 363
 364
 365
 366
 367
 368
 369
 370
 371
 372
 373
 374
 375
 376
 377
 378
 379
 380
 381
 382
 383
 384
 385
 386
 387
 388
 389
 390
 391
 392
 393
 394
 395
 396
 397
 398
 399
 400
 401
 402
 403
 404
 405
 406
 407
 408
 409
 410
 411
 412
 413
 414
 415
 416
 417
 418
 419
 420
 421
 422
 423
 424
 425
 426
 427
 428
 429
 430
 431
 432
 433
 434
 435
 436
 437
 438
 439
 440
 441
 442
 443
 444
 445
 446
 447
 448
 449
 450
 451
 452
 453
 454
 455
 456
 457
 458
 459
 460
 461
 462
 463
 464
 465
 466
 467
 468
 469
 470
 471
 472
 473
 474
 475
 476
 477
 478
 479
 480
 481
 482
 483
 484
 485
 486
 487
 488
 489
 490
 491
 492
 493
 494
 495
 496
 497
 498
 499
 500
 501
 502
 503
 504
 505
 506
 507
 508
 509
 510
 511
 512
 513
 514
 515
 516
 517
 518
 519
 520
 521
 522
 523
 524
 525
 526
 527
 528
 529
 530
 531
 532
 533
 534
 535
 536
 537
 538
 539
 540
 541
 542
 543
 544
 545
 546
 547
 548
 549
 550
 551
 552
 553
 554
 555
 556
 557
 558
 559
 560
 561
 562
 563
 564
 565
 566
 567
 568
 569
 570
 571
 572
 573
 574
 575
 576
 577
 578
 579
 580
 581
 582
 583
 584
 585
 586
 587
 588
 589
 590
 591
 592
 593
 594
 595
 596
 597
 598
 599
 600
 601
 602
 603
 604
 605
 606
 607
 608
 609
 610
 611
 612
 613
 614
 615
 616
 617
 618
 619
 620
 621
 622
 623
 624
 625
 626
 627
 628
 629
 630
 631
 632
 633
 634
 635
 636
 637
 638
 639
 640
 641
 642
 643
 644
 645
 646
 647
 648
 649
 650
 651
 652
 653
 654
 655
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
class WikiStoreBase(abc.ABC):
    """Abstract base for wiki persistence backends."""

    # Projects

    @abc.abstractmethod
    def create_project(self, project: Project) -> None:
        """Persist a new project record."""

    @abc.abstractmethod
    def get_project(self, slug: str) -> Project | None:
        """Return the project for *slug*, or None if absent."""

    @abc.abstractmethod
    def list_projects(self) -> list[Project]:
        """Return all projects sorted by indexed_at descending."""

    @abc.abstractmethod
    def delete_project(self, slug: str) -> bool:
        """Delete project *slug*; return True if deleted, False if absent."""

    # The ONLY ``Project`` fields a partial :meth:`update_project` may write.
    # Deliberately tiny, and not an oversight: every other field is SYSTEM-owned
    # — rebuilt wholesale by the next index run (``pages``/``indexed_at``/
    # ``landing_page_id``/``commit_sha``/``branch``/``maintainer_edited``/
    # ``graph_only``) or part of identity (``slug``/``source``/``host``/
    # ``repo_url``). Writing one here would either be silently clobbered at the
    # next finalize or make the record lie about what was actually indexed.
    # Settings that take effect on the NEXT index belong on ``ProjectSettings``.
    #
    # ``resolution`` is the one member written by an INDEX rather than by a
    # user edit: the graph phase is the only place that knows whether exact
    # cross-file symbol resolution ran, and it runs while the project record
    # for the previous index is still the current one. It is listed here
    # because that write is a partial update of an existing row, not because
    # the field is user-editable — nothing on the settings surface may set it.
    PROJECT_UPDATABLE: frozenset[str] = frozenset({"desc", "resolution"})

    def update_project(self, slug: str, fields: dict[str, Any]) -> Project | None:
        """Apply a partial update to *slug*'s Project; return the new state.

        Cost: ``O(one record)`` — one project read and one upsert of that same
        record, regardless of how much the project has indexed.

        Returns ``None`` when the project is absent. Only keys in
        :data:`PROJECT_UPDATABLE` are honoured — an unknown or ``None`` value is
        ignored, so a caller can hand over a whole PATCH body without pre-filtering
        (mirrors ``agentic_search.store.update_workspace``). ``None`` meaning "not
        supplied" is the OPPOSITE of ``update_job``'s rule, where a named ``None``
        is a value to be written — do not carry one convention onto the other.

        Concrete on the base rather than per-backend: ``create_project`` is an
        UPSERT in both drivers, so read → ``model_copy`` → upsert needs no
        duplicated JSON/Mongo pair that could drift. It is a read-modify-write with
        the same (non-)atomicity as every other Project write in this store.
        """
        project = self.get_project(slug)
        if project is None:
            return None
        updates = {
            key: value
            for key, value in fields.items()
            if key in self.PROJECT_UPDATABLE and value is not None
        }
        if not updates:
            return project
        updated = project.model_copy(update=updates)
        self.create_project(updated)
        return updated

    # Project settings (slug-keyed edit target — see ``ProjectSettings``)
    #
    # Its own surface, NOT the job-keyed submission sidecar: the sidecar records
    # what ONE job ran with (immutable history), while this records what the
    # project is CONFIGURED with (mutable, the PATCH target). Same separation the
    # recovery counter makes for the same reason.

    @abc.abstractmethod
    def save_project_settings(self, slug: str, settings: ProjectSettings) -> None:
        """Persist (upsert) the editable settings record for *slug*."""

    @abc.abstractmethod
    def get_project_settings(self, slug: str) -> ProjectSettings | None:
        """Return *slug*'s settings record, or None when it has never been written.

        ``None`` is a NORMAL state, not an error — the caller falls back to
        the per-job submission scan.
        """

    @abc.abstractmethod
    def delete_project_settings(self, slug: str) -> bool:
        """Delete *slug*'s settings record; return True if one existed.

        Called on project delete so a re-created slug can't inherit the dead
        project's settings (the rule the freshness cache eviction already follows).
        """

    def reap_slug(self, slug: str) -> dict[str, int]:
        """Delete every family this store persists for *slug*; return the counts.

        The exceptions are the three the delete-project ROUTE already owns
        directly: ``wiki_projects``
        (:meth:`delete_project`), ``wiki_settings`` (:meth:`delete_project_settings`),
        and a git credential (``CredentialStore.delete`` — repo-scope ONLY; a
        host-scoped credential is shared by every repo on that host and must
        NEVER cascade off one project's delete, so this method never touches
        credentials at all).

        Everything else a completed or in-flight index could have written under
        *slug* is deleted here: pages, the code graph (nodes/edges/embeddings),
        the entity layer (entities/entity-edges/entity-embeddings/
        recommendations), the memory layer (nodes/edges/embeddings), doc notes,
        the file manifest, the recovery-attempt counter, every indexing job
        (with its plan/resume/submission/session sidecars and its event log),
        and every QA answer for this slug (with its event log). Without this, a
        project delete orphans the large majority of what an index ever wrote —
        it stays reachable by nothing, forever, and a slug reused later inherits
        none of it (the same "clean slate" reasoning behind clearing settings/
        credentials/the freshness cache above, extended to the rest of the store).

        Returns per-family deleted-row counts (mirroring
        :meth:`supersede_graph_artifacts`'s shape) so a caller can report exactly
        what was removed. Idempotent: a second call for an already-reaped slug
        finds nothing and returns all-zero counts.
        """
        raise NotImplementedError("Graph backend is not implemented on this driver")

    # Pages

    @abc.abstractmethod
    def save_page(
        self,
        slug: str,
        page: WikiPage,
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Persist *page* for the project *slug*; overwrites if same page_id.

        ``commit_sha``/``job_id`` attribute the page to its owning index. Unlike
        the graph families, page attribution is a store-internal column, never a
        ``WikiPage`` field — the page IS a console wire type serialized whole, so
        keeping attribution off the model preserves that wire byte-for-byte. Page
        supersession is unaffected: ``wiki_finalize`` already prunes pages to the
        committed plan, so this is provenance, not the supersede mechanism.
        """

    def get_page(self, slug: str, page_id: str) -> WikiPage | None:
        """Return a single wiki page, or None if absent.

        THE single doc-content read seam: every upstream doc reader (the API
        page route, the ``wiki_read_page`` Q&A tool, the MCP ``read_wiki_page``
        facade over the route) funnels through here, so guarding it once makes a
        graph-only project's "no documentation" failure deterministic everywhere.
        A project indexed in graph-only (developer) mode carries
        ``graph_only=True`` and has ZERO pages — reading page content then raises
        :class:`DocumentationUnavailableError` from this ONE place rather than
        returning a confusing ``None``. The driver-specific fetch lives in
        :meth:`_get_page_raw`; this template method only adds the guard. The graph
        endpoint never calls this, so visualisation stays unaffected.
        """
        self._assert_docs_available(slug)
        return self._get_page_raw(slug, page_id)

    @abc.abstractmethod
    def _get_page_raw(self, slug: str, page_id: str) -> WikiPage | None:
        """Driver fetch of a single page (no graph-only guard)."""

    def _assert_docs_available(self, slug: str) -> None:
        """Raise :class:`DocumentationUnavailableError` for a graph-only project.

        Best-effort lookup: a missing/absent project is NOT a graph-only one, so
        the read proceeds (the caller handles the ``None`` page). Only a project
        record present AND flagged ``graph_only`` blocks doc reads.
        """
        from .errors import DocumentationUnavailableError  # noqa: PLC0415

        project = self.get_project(slug)
        if project is not None and getattr(project, "graph_only", False):
            raise DocumentationUnavailableError(slug)

    @abc.abstractmethod
    def list_pages(self, slug: str) -> list[WikiPage]:
        """Return all pages for project *slug*.

        NOT doc-guarded: ``list_pages`` is also the page-id roster used by the
        deterministic internal paths (finalize prune, resume, retriever,
        ``QaFinalizer.tag_page_citations``), so guarding it would break indexing
        and finalize themselves. A graph-only project simply returns ``[]`` here
        (it has no pages); the guard lives on the doc-CONTENT read (:meth:`get_page`).
        """

    def prune_pages(self, slug: str, keep: Iterable[str]) -> int:
        """Drop every page for *slug* whose ``page_id`` is not in *keep*.

        Default impl uses ``list_pages`` + per-page ``delete_page`` so
        backends only need a single primitive. Returns the number of
        pages dropped.
        """
        keep_set = set(keep)
        dropped = 0
        for page in self.list_pages(slug):
            if page.id not in keep_set:
                self.delete_page(slug, page.id)
                dropped += 1
        return dropped

    @abc.abstractmethod
    def delete_page(self, slug: str, page_id: str) -> bool:
        """Delete a single wiki page. Returns True if a page was removed."""

    # Indexing jobs

    @abc.abstractmethod
    def create_job(self, job: IndexingJob) -> None:
        """Persist a new indexing job."""

    @abc.abstractmethod
    def get_job(self, job_id: str) -> IndexingJob | None:
        """Return the indexing job, or None if absent."""

    def update_job(self, job_id: str, **fields: Any) -> IndexingJob:
        """Partially update *job_id* with *fields*; return the updated record.

        Concrete on the base because the rule both drivers owe their callers is
        ONE rule: only the fields the caller NAMED are written. Persisting the
        whole merged document instead turns every write into a read-modify-write
        over every field, so two writers that overlap lose one another's changes
        even when the fields they touch are disjoint. The phase tools write
        progress on a 50ms-to-5s cadence while an HTTP thread writes ``status``,
        so a cancel silently reverting to ``running`` needs no exotic timing to
        reproduce.

        The narrowing lives here; making the narrow write INDIVISIBLE is
        per-backend (:meth:`_write_job_patch`), because a lock and a
        field-scoped update are not the same cure.

        Raises ``KeyError`` when the job is absent and ``ValidationError`` when
        a field is unknown to the model or its value is wrong for it. Naming no
        field at all is a read: there is nothing to narrow, so nothing is
        written.
        """
        job = self.get_job(job_id)
        if job is None:
            raise KeyError(f"Job not found: {job_id}")
        if not fields:
            return job
        return self._write_job_patch(job_id, JobPatch.build(job, fields))

    @abc.abstractmethod
    def _write_job_patch(self, job_id: str, patch: JobPatch) -> IndexingJob:
        """Persist exactly *patch*'s fields onto *job_id*; return the STORED record.

        The backend's one job here is to make "read the record, set these
        fields, put it back" indivisible with respect to another writer, and to
        leave every field the patch does not name exactly as that other writer
        left it. Returns what is actually stored afterwards — concurrent writes
        included — never the caller's locally merged guess at it. Raises
        ``KeyError`` if the job vanished since the caller's read.
        """

    @abc.abstractmethod
    def list_jobs(self, slug: str | None = None) -> list[IndexingJob]:
        """Return all jobs, optionally filtered to *slug*."""

    def latest_job(
        self, slug: str, *, statuses: Iterable[str] | None = None
    ) -> IndexingJob | None:
        """Return *slug*'s most recent job, ordered by ``phase_started_at``.

        The ONE "what is the latest attempt for this slug" answer — a thin
        convenience over :meth:`list_jobs`, which already returns newest-first
        by ``phase_started_at`` (never ``job_id``, a ``uuid4`` hex that sorts
        RANDOMLY with respect to when a job actually ran). *statuses* narrows
        the candidates before taking the first, e.g. ``{"complete"}`` for a
        freshness or QA-source baseline. Returns ``None`` when no job
        (matching *statuses*, if given) exists.
        """
        candidates = self.list_jobs(slug=slug)
        if statuses is not None:
            wanted = set(statuses)
            candidates = [j for j in candidates if j.status in wanted]
        return candidates[0] if candidates else None

    def append_job_event(self, job_id: str, event: dict[str, Any]) -> int:
        """Validate then append one job event; return its monotonic index.

        The discriminated event model is the durable trust boundary: both the
        JSON and Mongo timelines, and therefore SSE, receive one known wire
        shape. Concrete drivers only own their atomic append mechanics.

        Cost: ``O(1)``.
        """
        payload = WikiJobEvent.parse(event).stored_payload()
        return self._append_job_event(job_id, payload)

    def _append_job_event(self, job_id: str, event: dict[str, Any]) -> int:
        """Append already-validated *event* with driver-specific atomicity.

        Kept concrete so lightweight store subclasses that implement the
        long-standing public ``append_job_event`` contract remain instantiable.
        Production drivers override this hook; the base has no durable timeline.
        """
        raise NotImplementedError

    @abc.abstractmethod
    def load_job_events(
        self, job_id: str, after_idx: int = -1
    ) -> list[dict[str, Any]]:
        """Return job events with idx > *after_idx* (-1 returns all)."""

    @abc.abstractmethod
    def cancel_job(self, job_id: str) -> bool:
        """Cancel *job_id*; return True on first cancel, False if already cancelled."""

    @abc.abstractmethod
    def attach_job_session(self, job_id: str, session_id: str) -> None:
        """Associate a Mewbo session_id with an indexing job (forward mapping)."""

    @abc.abstractmethod
    def get_job_session(self, job_id: str) -> str | None:
        """Return the session_id attached to *job_id*, or None."""

    @abc.abstractmethod
    def find_job_by_session(self, session_id: str) -> str | None:
        """Reverse lookup: return the job_id for *session_id*, or None."""

    # Job plan + extra metadata (not part of IndexingJob schema)

    @abc.abstractmethod
    def save_job_plan(self, job_id: str, plan: list[dict[str, Any]]) -> None:
        """Persist the page-plan list for *job_id*; overwrites any previous plan."""

    @abc.abstractmethod
    def get_job_plan(self, job_id: str) -> list[dict[str, Any]] | None:
        """Return the page-plan list, or None if no plan has been committed yet."""

    # Resume sidecar (checkpoint-aware recovery). A tiny dict computed
    # ONCE by ``ResumePlan.build`` at resume time; rebuilt cheaply per tool call
    # via ``ResumePlan.from_persisted`` so the phase skip-guards never re-query
    # the graph. Concrete defaults (no-op / None) so a backend that never persists
    # it simply degrades to a full rebuild on resume — never a crash.

    def save_resume_plan(self, job_id: str, plan: dict[str, Any]) -> None:
        """Persist the resume-plan dict for *job_id*; overwrites any previous one."""
        raise NotImplementedError

    def get_resume_plan(self, job_id: str) -> dict[str, Any] | None:
        """Return the persisted resume-plan dict, or None if the job isn't resuming."""
        return None

    # Act sidecar (the scoped refresh's second stage). ONE record carrying the
    # narrowed page-id work-list AND which stage the job reached, because the
    # two are read together and a job cannot express the second anywhere else:
    # ``refresh_decision`` is stamped once, before the run starts, so it can say
    # WHICH path a refresh took but never HOW FAR it got. Without this a job
    # that died in stage 2 re-ran clone + scan + the whole delta pass on resume.
    # Concrete defaults (no-op / None) for the same reason the resume sidecar
    # has them: a backend that never persists it degrades to re-driving stage 1,
    # which is the behaviour that existed before — never a crash.

    def save_act_plan(self, job_id: str, plan: dict[str, Any]) -> None:
        """Persist the act-stage record for *job_id*; overwrites any previous one."""
        raise NotImplementedError

    def get_act_plan(self, job_id: str) -> dict[str, Any] | None:
        """Return the persisted act-stage record, or None if stage 1 hasn't finished."""
        return None

    @abc.abstractmethod
    def get_job_submitted_count(self, job_id: str) -> int:
        """Return the number of pages submitted so far for *job_id*."""

    @abc.abstractmethod
    def claim_job_page(self, slug: str, job_id: str, page_id: str) -> PageClaim:
        """Atomically record *page_id* as written by *job_id*; return the claim.

        The submitted-pages counter belongs to the JOB, so its dedup key has to
        be the job's OWN set of claimed ids. Keying on the slug instead — "does a
        page with this id already exist for the slug" — answers yes for every
        page a PREVIOUS index of the same repository wrote, so a refresh would
        count zero new pages, emit no ``page_committed`` events, and leave the
        progress bar dead for the whole re-index.

        The returned count is the SIZE of that set, never a free-running
        increment. A counter that only ever counts distinct pages cannot drift
        past what the job actually wrote, where a read-then-``$inc`` can: a
        resumed job whose counter carried over from an earlier attempt reports
        90 pages written against a 50-page plan.

        A job with no claim record seeds its set from page attribution
        (:meth:`page_ids_for_job`) on first touch, so an interrupted index
        resumes with its earlier pages counted rather than from zero.

        Raises ``KeyError`` when *job_id* is unknown.
        """

    @abc.abstractmethod
    def get_job_page_ids(self, slug: str, job_id: str) -> frozenset[str]:
        """Return the page ids *job_id* wrote — its claim set, else attribution.

        The fallback is what makes a PRE-EXISTING interrupted job resumable: its
        meta carries a bare counter and no claim record, so reading the claim
        alone reported nothing done and the resume regenerated every page it had
        already written correctly — the exact waste the claim exists to prevent.
        """

    @abc.abstractmethod
    def page_ids_for_job(self, slug: str, job_id: str) -> frozenset[str]:
        """Page ids under *slug* whose stored attribution names *job_id*.

        ``save_page`` stamps attribution independently of the claim record,
        which is why this can answer for a job the claim record cannot.
        """

    @abc.abstractmethod
    def save_job_submission(self, job_id: str, submission: dict[str, Any]) -> None:
        """Persist the wizard submission dict for *job_id* (token must be absent)."""

    @abc.abstractmethod
    def get_job_submission(self, job_id: str) -> dict[str, Any] | None:
        """Return the persisted submission dict, or None if not yet saved."""

    # Repository credentials (isolated, per-slug, plaintext-at-rest)

    @abc.abstractmethod
    def save_credentials(self, slug: str, blob: dict[str, Any]) -> None:
        """Persist the (already encoded) credential *blob* for *slug*; overwrite."""

    @abc.abstractmethod
    def get_credentials(self, slug: str) -> dict[str, Any] | None:
        """Return the encoded credential blob for *slug*, or None if absent."""

    @abc.abstractmethod
    def delete_credentials(self, slug: str) -> bool:
        """Delete *slug*'s credential; return True if one was removed, else False."""

    @abc.abstractmethod
    def list_credentials(self) -> dict[str, dict[str, Any]]:
        """Return every stored credential blob keyed by scope (slug or bare host).

        The management surface (``CredentialStore.list`` → the ``/v1/git/credentials``
        route) reads this. Values are the raw encoded blobs; the caller decodes +
        redacts. The scope key is a full slug (``host/owner/repo``) or a bare host.
        """

    # Restart-recovery counter (slug-keyed, isolated from the submission sidecar)

    @abc.abstractmethod
    def get_recovery_attempts(self, slug: str) -> int:
        """Return the recovery-attempt count for *slug* (0 if never recovered)."""

    @abc.abstractmethod
    def bump_recovery_attempts(self, slug: str) -> int:
        """Atomically increment *slug*'s recovery counter; return the new value.

        Slug-keyed (not job-keyed) so the cap bounds re-drives across recovery
        generations / new job_ids. Lives on its OWN persistent surface so it
        never pollutes the wizard-submission sidecar (which validates strictly
        as a ``WizardSubmission``).
        """

    def reset_recovery_attempts(self, slug: str) -> None:
        """Clear *slug*'s recovery counter (a user-initiated resume gets a fresh budget).

        A human asking to retry an index must not be blocked by prior automatic
        re-drives, so the manual resume path resets the auto-recovery cap. Concrete
        default no-op so a backend that never tracks the counter is unaffected.
        """

    # QA

    @abc.abstractmethod
    def save_qa(self, answer: QaAnswer) -> None:
        """Persist a QA answer record. Use ONLY to create it (resets bookkeeping)."""

    @abc.abstractmethod
    def update_qa_fields(self, answer: QaAnswer) -> None:
        """Update a QA answer's content fields in place — NON-destructive.

        Persists every ``QaAnswer`` field but MUST NOT disturb store bookkeeping
        that some backends pack alongside the record: the ``event_count`` idx
        counter and the ``session_id`` mapping. ``save_qa`` does a FULL replace,
        which on Mongo resets ``event_count`` to 0 (so the next ``append_qa_event``
        collides at idx 0) AND drops ``session_id`` (breaking
        ``find_qa_by_session``). Mid-stream writers (``QaFinalizer``) MUST use
        this; ``save_qa`` is for creation only. (The JSON backend keeps session +
        events in separate files, so for it this is just an answer.json rewrite —
        the divergence is why a JSON-only test cannot catch the Mongo failure.)
        """

    @abc.abstractmethod
    def get_qa(self, answer_id: str) -> QaAnswer | None:
        """Return the QA answer, or None if absent."""

    @abc.abstractmethod
    def list_qa(self, status: str | None = None) -> list[QaAnswer]:
        """Return all QA answers, optionally filtered to *status*.

        Boot-time/offline use only (mirrors :meth:`list_jobs`) — an
        unfiltered call reads every persisted answer, so a caller on an
        interactive path must narrow with *status* rather than filtering the
        full list in Python.
        """

    @abc.abstractmethod
    def attach_qa_session(self, answer_id: str, session_id: str) -> None:
        """Associate a Mewbo session_id with a QA answer (forward mapping)."""

    @abc.abstractmethod
    def get_qa_session(self, answer_id: str) -> str | None:
        """Return the session_id attached to *answer_id*, or None."""

    @abc.abstractmethod
    def find_qa_by_session(self, session_id: str) -> str | None:
        """Reverse lookup: return the answer_id for *session_id*, or None."""

    @abc.abstractmethod
    def append_qa_event(self, answer_id: str, event: dict[str, Any]) -> int:
        """Append *event* to the QA event log; return the monotonic idx."""

    @abc.abstractmethod
    def load_qa_events(
        self, answer_id: str, after_idx: int = -1
    ) -> list[dict[str, Any]]:
        """Return QA events with idx > *after_idx* (-1 returns all)."""

    # Graph + embeddings (raise NotImplementedError in v1)
    #
    # ``commit_sha``/``job_id`` are the per-job/commit attribution the store
    # stamps onto every artifact (see ``GraphNodeBase``). A caller that owns a
    # job passes them so the write is attributed to the commit it indexed; a
    # commit-less path (catalog ingest) omits them and the row is stamped
    # ``None``. Keyword-only + defaulted so a caller that owns no job can omit them.

    @staticmethod
    def _stamp_attribution(item: _M, commit_sha: str | None, job_id: str | None) -> _M:
        """Return *item* carrying the write's ``commit_sha``/``job_id``.

        A no-op (identity) when neither is supplied — a commit-less catalog or
        Q&A write leaves the row ``None``-stamped, which is what supersede treats
        as "not a per-commit snapshot" and preserves. Otherwise a frozen-safe
        ``model_copy`` overwrites both, so the persisted attribution is the
        writer's, never whatever a constructed model happened to carry.
        """
        if commit_sha is None and job_id is None:
            return item
        return item.model_copy(update={"commit_sha": commit_sha, "job_id": job_id})

    def upsert_nodes(
        self,
        slug: str,
        nodes: Iterable[GraphNode],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
        on_progress: Callable[[int, int], None] | None = None,
    ) -> None:
        """Upsert code-graph nodes, reporting completed persistence batches."""
        raise NotImplementedError("Graph backend is not implemented on this driver")

    def upsert_edges(
        self,
        slug: str,
        edges: Iterable[GraphEdge],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
        on_progress: Callable[[int, int], None] | None = None,
    ) -> None:
        """Upsert code-graph edges, reporting completed persistence batches."""
        raise NotImplementedError("Graph backend is not implemented on this driver")

    def upsert_embeddings(
        self,
        slug: str,
        items: Iterable[Embedding],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert dense embedding vectors."""
        raise NotImplementedError("Embeddings are not implemented on this driver")

    def live_scope(self, slug: str) -> CommitScope:
        """The generation *slug*'s readers should see: the project's own commit.

        A project row carries the commit its last completed index built from,
        so that — not the union of every commit ever indexed — is what "the
        code as it is now" means. Lives here because the store is what holds
        the project row; every graph reader already has one, so nobody has to
        thread a commit sha through their own call chain to read correctly.

        A project with no ``commit_sha`` (a commit-less catalog ingestion, or a
        slug with no project row yet) has exactly one generation, so scoping it
        could only be a way to get it wrong — the union and the live generation
        are the same set, and ``every()`` says so without asserting a commit
        that does not exist.

        Cost: ``O(one record)``.
        """
        project = self.get_project(slug)
        if project is None or project.commit_sha is None:
            return CommitScope.every()
        return CommitScope.at(project.commit_sha)

    def query_graph(
        self,
        slug: str,
        *,
        scope: CommitScope,
        node_type: str | None = None,
        name_match: str | None = None,
        neighbors_of: str | None = None,
        node_ids: Collection[str] | None = None,
    ) -> list[GraphNode]:
        """Query the code graph for *slug*, restricted to *scope*'s generation.

        *node_ids* fetches an explicit set by id. It exists so a caller holding
        a bounded list of ids — the knowledge-graph view repairing entity
        anchors that point into a superseded generation — can look exactly
        those up instead of reading a whole generation to find a handful. An
        empty collection returns nothing; ``None`` means "no id filter".

        *scope* is REQUIRED and has no default on purpose. This store holds the
        UNION of every commit ever indexed for a slug, so "which generation"
        has no safe default — and a default of ``CommitScope.every()`` would be
        precisely the fail-open filter the root guidance forbids: a caller that
        forgot to scope would silently read every generation and still return
        ``200``. Required means a missed call site is a typecheck failure
        instead. See :class:`CommitScope` for why this is not a ``commit_sha``
        parameter (``count_graph_nodes`` already owns that name with the
        opposite meaning for ``None``).

        Cost: ``O(nodes for the slug in scope)`` — a read of one project's
        graph generation, not of all history.
        """
        raise NotImplementedError("Graph backend is not implemented on this driver")

    def count_graph_nodes(self, slug: str, *, commit_sha: str | None) -> int:
        """Count nodes for *slug* built by exactly *commit_sha* (``None`` matches None).

        The commit-scoped count the resume skip predicate keys on: "the graph
        for THIS commit is built" is ``count_graph_nodes(slug, commit_sha=X) >
        0``.

        ``commit_sha`` here is an EXACT match — ``None`` counts the rows stamped
        NULL, it does not mean "any commit". :class:`CommitScope` exists to keep
        that convention unambiguous on the read methods, which is why
        ``query_graph`` takes a scope rather than re-using this parameter name
        with the opposite sense; ``count_graph_nodes(slug, commit_sha=X)`` and
        ``len(query_graph(slug, scope=CommitScope.at(X)))`` agree by
        construction.
        """
        raise NotImplementedError("Graph backend is not implemented on this driver")

    def supersede_graph_artifacts(
        self, slug: str, *, keep_commit_sha: str
    ) -> dict[str, int]:
        """Reap prior-commit graph + entity artifacts once *keep_commit_sha* completes.

        Deletes every node/edge/embedding/entity/entity-edge/entity-embedding for
        *slug* whose ``commit_sha`` is a REAL value other than *keep_commit_sha*.
        Rows stamped ``None`` are PRESERVED — for entities that is a QA-minted or
        pre-isolation record (accretive memory, not a per-commit snapshot); for
        code nodes there are none on a git slug (every index stamps its commit).
        Returns per-collection delete counts. Idempotent: a second call for the
        same *keep_commit_sha* finds nothing to reap.
        """
        raise NotImplementedError("Graph backend is not implemented on this driver")

    def restamp_graph_artifacts(
        self, slug: str, *, from_commit: str, to_commit: str
    ) -> dict[str, int]:
        """Carry *from_commit*'s surviving artifacts forward onto *to_commit*.

        Concrete-raising rather than ``@abc.abstractmethod``, matching its
        sibling ``supersede_graph_artifacts`` and the rest of this graph family:
        an abstract method here is inherited by every partial test double of
        this base and stops it being instantiable at all, which turns a new
        store capability into a broad, unrelated test failure.

        The counterpart a SCOPED re-index needs and a full one does not. A full
        index re-stamps every artifact by rewriting them all, so afterwards one
        generation describes the whole slug and ``live_scope`` finds everything.
        An incremental refresh re-stamps only what it re-parsed, so without this
        the untouched majority keeps the PREVIOUS commit while the project row
        advances — and since ``live_scope`` is ``at(project.commit_sha)``, every
        reader (the graph view, the retriever, the agent graph tools) would see
        only the handful of files that happened to change. Not deleted, not
        erroring: invisible. That is the failure this method exists to prevent,
        and it is the precondition the delta indexer's own unscoped reads name.

        **The semantic is a claim, and it is a true one:** an artifact built
        from a file that did NOT change describes *to_commit* just as accurately
        as it described *from_commit*, so moving its stamp forward asserts
        nothing false.

        **Why this is a MOVE from a named generation rather than "stamp
        everything".** The store can hold generations older than *from_commit* —
        a supersede that never ran, or the field-absent rows the isolation
        backfill exists for. Those describe code that is genuinely gone.
        Blanket-stamping them onto *to_commit* would resurrect deleted symbols
        into the live view, which is a worse bug than the one being fixed. Rows
        stamped ``None`` are likewise left alone, matching
        ``supersede_graph_artifacts``' preserve rule.

        Returns per-family moved counts. Idempotent: a second call finds nothing
        left at *from_commit*.

        Cost: ``O(artifacts at from_commit)`` — an offline finalize step.
        """
        raise NotImplementedError

    def list_edges(self, slug: str, *, scope: CommitScope) -> list[GraphEdge]:
        """Return *slug*'s edges for *scope*'s generation (graph-viewer endpoint).

        *scope* is required for the same reason as on ``query_graph`` — and it
        must agree with the scope the nodes were read under, or the view is
        assembled from edges whose endpoints belong to a different generation.
        """
        raise NotImplementedError("Graph backend is not implemented on this driver")

    def vector_search(
        self, slug: str, qvec: list[float], k: int = 10
    ) -> list[Embedding]:
        """Nearest-neighbour vector search."""
        raise NotImplementedError("Embeddings are not implemented on this driver")

    # Scoped graph deletes (used by the incremental GraphDeltaIndexer)

    def delete_nodes_by_file(self, slug: str, file: str) -> int:
        """Delete every code node in *file* AND its vectors; return the NODE count.

        The embedding cascade is part of this method rather than a second call
        the caller makes first, because :class:`Embedding` carries no ``file``
        field: a vector is reachable only through the node it points at, so once
        that node is deleted the vector can no longer be found by file, by
        commit, or by anything else — it is simply unreachable, unreapable, and
        still scoreable by :meth:`vector_search`. Cascading here makes "no
        orphaned vectors" a property of the store instead of a call-ordering
        discipline every future caller has to rediscover.

        The return value stays the NODE count so the delete reads the same as
        every other scoped delete on this class.
        """
        raise NotImplementedError("Memory layer methods land in the memory store")

    def delete_edges_by_source_file(self, slug: str, file: str) -> int:
        """Delete edges originating from any node in *file*; return count.

        "Originating" = the edge ``source`` node_id belongs to a node whose
        ``file`` is *file*. Call BEFORE :meth:`delete_nodes_by_file` for the
        same file so the source nodes are still present to resolve.
        """
        raise NotImplementedError("Memory layer methods land in the memory store")

    # Memory layer — nodes / edges / embeddings (multiplex overlay)
    #
    # Default impls raise NotImplementedError so a backend opts in by
    # overriding (no separate MemoryStoreBase ABC — KISS). Both shipping
    # drivers (JSON, Mongo) implement the full surface.

    def upsert_memory_nodes(self, slug: str, nodes: Iterable[MemoryNode]) -> None:
        """Upsert memory nodes; dedup by ``node_id``."""
        raise NotImplementedError

    def get_memory_node(self, slug: str, node_id: str) -> MemoryNode | None:
        """Return a single memory node, or None if absent."""
        raise NotImplementedError

    def delete_memory_node(self, slug: str, node_id: str) -> bool:
        """Delete a memory node + its embedding; return True if one was removed.

        Edges are NOT touched (callers invalidate them separately so history
        survives). Used when a merge supersedes a note under a new identity.
        """
        raise NotImplementedError

    def query_memory(
        self, slug: str, *, filt: MemoryFilter | None = None
    ) -> list[MemoryNode]:
        """Return memory nodes matching *filt*'s node-level facets."""
        raise NotImplementedError

    def upsert_memory_edges(self, slug: str, edges: Iterable[MemoryEdge]) -> None:
        """Upsert memory edges; dedup by ``(source, target, type)``."""
        raise NotImplementedError

    def list_memory_edges(
        self,
        slug: str,
        *,
        node_id: str | None = None,
        include_invalidated: bool = False,
    ) -> list[MemoryEdge]:
        """Return memory edges, optionally scoped to ``source == node_id``.

        Invalidated edges (``invalid_at`` set) are excluded unless
        *include_invalidated* is True.
        """
        raise NotImplementedError

    def memories_anchored_to(
        self,
        slug: str,
        entity_keys: Iterable[EntityKey],
        *,
        include_invalidated: bool = False,
    ) -> list[str]:
        """Reverse ANCHORS lookup: entity_keys → distinct memory node_ids."""
        raise NotImplementedError

    def upsert_memory_embeddings(
        self, slug: str, items: Iterable[MemoryEmbedding]
    ) -> None:
        """Upsert memory embedding vectors; dedup by ``node_id``."""
        raise NotImplementedError

    def memory_vector_search(
        self,
        slug: str,
        qvec: list[float],
        k: int = 10,
        *,
        filt: MemoryFilter | None = None,
    ) -> list[MemoryEmbedding]:
        """Top-k memory embeddings by cosine, after applying *filt*.

        Scale seam — keep signature stable. v1 is brute-force cosine; IVF /
        Matryoshka / quantization slot in here without touching callers.
        """
        raise NotImplementedError

    def _live_anchored_ids(self, slug: str) -> set[str]:
        """Memory node_ids with ≥1 live ANCHORS edge (validity gate)."""
        raise NotImplementedError

    @staticmethod
    def _rank_embeddings(
        pool: list[_V], qvec: list[float], k: int
    ) -> list[_V]:
        """Cosine-rank a pre-loaded embedding pool and return the top-k.

        The single cosine-sort core shared by memory AND entity vector search:
        each driver loads (and, for memory, facet-filters) its own pool, then
        delegates here so scoring/ordering can never desync across families or
        backends. Generic over any row with a ``vector`` (`_HasVector`).
        """
        if not pool:
            return []
        scored = [(emb, _cosine(qvec, emb.vector)) for emb in pool]
        scored.sort(key=lambda t: t[1], reverse=True)
        return [emb for emb, _ in scored[:k]]

    def _rank_memory(
        self,
        slug: str,
        pool: list[MemoryEmbedding],
        qvec: list[float],
        k: int,
        filt: MemoryFilter | None,
    ) -> list[MemoryEmbedding]:
        """Facet-filter then cosine-rank a backend-loaded memory pool (top-k).

        The single ranking core both memory drivers share: each loads its own
        pool, then delegates here so facet/validity intersection can never
        desync across backends. The cosine-sort itself is `_rank_embeddings`.
        """
        if not pool:
            return []
        if filt is not None:
            # Only load + facet-filter nodes when a facet is actually set; the
            # common path (validity only) skips that O(N) node scan.
            allowed: set[str] | None = None
            if filt.corpus or filt.source or filt.kind or filt.labels:
                allowed = {n.node_id for n in self.query_memory(slug, filt=filt)}
            if filt.exclude_invalidated:
                live = self._live_anchored_ids(slug)
                allowed = live if allowed is None else (allowed & live)
            if allowed is not None:
                pool = [e for e in pool if e.node_id in allowed]
        return self._rank_embeddings(pool, qvec, k)

    # Documentation-page notes (docs-as-multiplex-nodes)

    def upsert_doc_notes(self, slug: str, notes: Iterable[DocPageNote]) -> None:
        """Upsert doc-page notes; dedup by ``page_id``."""
        raise NotImplementedError

    def get_doc_note(self, slug: str, page_id: str) -> DocPageNote | None:
        """Return a single doc-page note, or None if absent."""
        raise NotImplementedError

    def list_doc_notes(self, slug: str) -> list[DocPageNote]:
        """Return every doc-page note for *slug*."""
        raise NotImplementedError

    def delete_doc_note(self, slug: str, page_id: str) -> bool:
        """Delete a doc-page note; return True if one was removed."""
        raise NotImplementedError

    # File manifest (incremental retract index)

    def upsert_file_manifest(
        self, slug: str, entries: Iterable[FileManifest]
    ) -> None:
        """Upsert per-file manifest entries; dedup by ``path``."""
        raise NotImplementedError

    def get_file_manifest(self, slug: str, path: str) -> FileManifest | None:
        """Return a single file-manifest entry, or None if absent."""
        raise NotImplementedError

    def list_file_manifest(self, slug: str) -> list[FileManifest]:
        """Return every file-manifest entry for *slug*."""
        raise NotImplementedError

    def delete_file_manifest(self, slug: str, path: str) -> bool:
        """Delete a file-manifest entry; return True if one was removed."""
        raise NotImplementedError

    # Abstract-entity layer (multiplex overlay — same opt-in pattern as memory)
    #
    # Default impls raise NotImplementedError so a backend opts in by
    # overriding (no separate EntityStore ABC — KISS, one store). Both shipping
    # drivers (JSON, Mongo) implement the full surface; the base stays a
    # default-raise (not @abstractmethod) so existing partial test doubles keep
    # instantiating, exactly like the memory layer above.

    def upsert_entities(
        self,
        slug: str,
        entities: Iterable[Entity],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert entities; dedup by ``id`` (= sha1(normalized_name|type))."""
        raise NotImplementedError

    def get_entity(self, slug: str, entity_id: str) -> Entity | None:
        """Return a single entity, or None if absent."""
        raise NotImplementedError

    def query_entities(
        self, slug: str, *, filt: EntityFilter | None = None
    ) -> list[Entity]:
        """Return entities matching *filt*'s facets (no filter ⇒ all)."""
        raise NotImplementedError

    def count_entities(self, slug: str, *, commit_sha: str | None) -> int:
        """Count entities for *slug* minted by exactly *commit_sha* (``None`` matches None).

        The enrich-phase analogue of :meth:`count_graph_nodes`: "entities for
        THIS commit are minted" is ``count_entities(slug, commit_sha=X) > 0``,
        which the resume skip predicate needs instead of the union count.
        """
        raise NotImplementedError

    def upsert_entity_embeddings(
        self,
        slug: str,
        items: Iterable[EntityEmbedding],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert entity embedding vectors; dedup by ``entity_id``."""
        raise NotImplementedError

    def entity_vector_search(
        self, slug: str, qvec: list[float], k: int = 10
    ) -> list[EntityEmbedding]:
        """Top-k entity embeddings by cosine (the ANN block seam for ER)."""
        raise NotImplementedError

    def upsert_entity_edges(
        self,
        slug: str,
        edges: Iterable[EntityRelation],
        *,
        commit_sha: str | None = None,
        job_id: str | None = None,
    ) -> None:
        """Upsert entity relations; dedup by ``id`` (= source|type|target)."""
        raise NotImplementedError

    def list_entity_edges(
        self, slug: str, *, source_id: str | None = None
    ) -> list[EntityRelation]:
        """Return entity relations, optionally scoped to ``source_id``."""
        raise NotImplementedError

    def save_entity_recommendation(
        self, slug: str, rec: EntityRecommendation
    ) -> None:
        """Upsert a resolution-recommendation record (a prior for the next pass).

        Keyed on ``rec.id`` (deterministic over action + sorted subjects + type),
        NOT appended: these records are read back as priors by ``EntityResolver``,
        so an unkeyed insert let a replayed enrich pass state the same prior
        twice and re-weight the ladder purely by having run again.
        """
        raise NotImplementedError

    def get_entity_recommendations(self, slug: str) -> list[EntityRecommendation]:
        """Return every persisted entity recommendation for *slug*."""
        raise NotImplementedError

append_job_event(job_id: str, event: dict[str, Any]) -> int

Validate then append one job event; return its monotonic index.

The discriminated event model is the durable trust boundary: both the JSON and Mongo timelines, and therefore SSE, receive one known wire shape. Concrete drivers only own their atomic append mechanics.

Cost: O(1).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
444
445
446
447
448
449
450
451
452
453
454
def append_job_event(self, job_id: str, event: dict[str, Any]) -> int:
    """Validate then append one job event; return its monotonic index.

    The discriminated event model is the durable trust boundary: both the
    JSON and Mongo timelines, and therefore SSE, receive one known wire
    shape. Concrete drivers only own their atomic append mechanics.

    Cost: ``O(1)``.
    """
    payload = WikiJobEvent.parse(event).stored_payload()
    return self._append_job_event(job_id, payload)

append_qa_event(answer_id: str, event: dict[str, Any]) -> int abstractmethod

Append event to the QA event log; return the monotonic idx.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
677
678
679
@abc.abstractmethod
def append_qa_event(self, answer_id: str, event: dict[str, Any]) -> int:
    """Append *event* to the QA event log; return the monotonic idx."""

attach_job_session(job_id: str, session_id: str) -> None abstractmethod

Associate a Mewbo session_id with an indexing job (forward mapping).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
475
476
477
@abc.abstractmethod
def attach_job_session(self, job_id: str, session_id: str) -> None:
    """Associate a Mewbo session_id with an indexing job (forward mapping)."""

attach_qa_session(answer_id: str, session_id: str) -> None abstractmethod

Associate a Mewbo session_id with a QA answer (forward mapping).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
665
666
667
@abc.abstractmethod
def attach_qa_session(self, answer_id: str, session_id: str) -> None:
    """Associate a Mewbo session_id with a QA answer (forward mapping)."""

bump_recovery_attempts(slug: str) -> int abstractmethod

Atomically increment slug's recovery counter; return the new value.

Slug-keyed (not job-keyed) so the cap bounds re-drives across recovery generations / new job_ids. Lives on its OWN persistent surface so it never pollutes the wizard-submission sidecar (which validates strictly as a WizardSubmission).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
612
613
614
615
616
617
618
619
620
@abc.abstractmethod
def bump_recovery_attempts(self, slug: str) -> int:
    """Atomically increment *slug*'s recovery counter; return the new value.

    Slug-keyed (not job-keyed) so the cap bounds re-drives across recovery
    generations / new job_ids. Lives on its OWN persistent surface so it
    never pollutes the wizard-submission sidecar (which validates strictly
    as a ``WizardSubmission``).
    """

cancel_job(job_id: str) -> bool abstractmethod

Cancel job_id; return True on first cancel, False if already cancelled.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
471
472
473
@abc.abstractmethod
def cancel_job(self, job_id: str) -> bool:
    """Cancel *job_id*; return True on first cancel, False if already cancelled."""

claim_job_page(slug: str, job_id: str, page_id: str) -> PageClaim abstractmethod

Atomically record page_id as written by job_id; return the claim.

The submitted-pages counter belongs to the JOB, so its dedup key has to be the job's OWN set of claimed ids. Keying on the slug instead — "does a page with this id already exist for the slug" — answers yes for every page a PREVIOUS index of the same repository wrote, so a refresh would count zero new pages, emit no page_committed events, and leave the progress bar dead for the whole re-index.

The returned count is the SIZE of that set, never a free-running increment. A counter that only ever counts distinct pages cannot drift past what the job actually wrote, where a read-then-$inc can: a resumed job whose counter carried over from an earlier attempt reports 90 pages written against a 50-page plan.

A job with no claim record seeds its set from page attribution (:meth:page_ids_for_job) on first touch, so an interrupted index resumes with its earlier pages counted rather than from zero.

Raises KeyError when job_id is unknown.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
@abc.abstractmethod
def claim_job_page(self, slug: str, job_id: str, page_id: str) -> PageClaim:
    """Atomically record *page_id* as written by *job_id*; return the claim.

    The submitted-pages counter belongs to the JOB, so its dedup key has to
    be the job's OWN set of claimed ids. Keying on the slug instead — "does a
    page with this id already exist for the slug" — answers yes for every
    page a PREVIOUS index of the same repository wrote, so a refresh would
    count zero new pages, emit no ``page_committed`` events, and leave the
    progress bar dead for the whole re-index.

    The returned count is the SIZE of that set, never a free-running
    increment. A counter that only ever counts distinct pages cannot drift
    past what the job actually wrote, where a read-then-``$inc`` can: a
    resumed job whose counter carried over from an earlier attempt reports
    90 pages written against a 50-page plan.

    A job with no claim record seeds its set from page attribution
    (:meth:`page_ids_for_job`) on first touch, so an interrupted index
    resumes with its earlier pages counted rather than from zero.

    Raises ``KeyError`` when *job_id* is unknown.
    """

count_entities(slug: str, *, commit_sha: str | None) -> int

Count entities for slug minted by exactly commit_sha (None matches None).

The enrich-phase analogue of :meth:count_graph_nodes: "entities for THIS commit are minted" is count_entities(slug, commit_sha=X) > 0, which the resume skip predicate needs instead of the union count.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1111
1112
1113
1114
1115
1116
1117
1118
def count_entities(self, slug: str, *, commit_sha: str | None) -> int:
    """Count entities for *slug* minted by exactly *commit_sha* (``None`` matches None).

    The enrich-phase analogue of :meth:`count_graph_nodes`: "entities for
    THIS commit are minted" is ``count_entities(slug, commit_sha=X) > 0``,
    which the resume skip predicate needs instead of the union count.
    """
    raise NotImplementedError

count_graph_nodes(slug: str, *, commit_sha: str | None) -> int

Count nodes for slug built by exactly commit_sha (None matches None).

The commit-scoped count the resume skip predicate keys on: "the graph for THIS commit is built" is count_graph_nodes(slug, commit_sha=X) > 0.

commit_sha here is an EXACT match — None counts the rows stamped NULL, it does not mean "any commit". :class:CommitScope exists to keep that convention unambiguous on the read methods, which is why query_graph takes a scope rather than re-using this parameter name with the opposite sense; count_graph_nodes(slug, commit_sha=X) and len(query_graph(slug, scope=CommitScope.at(X))) agree by construction.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
def count_graph_nodes(self, slug: str, *, commit_sha: str | None) -> int:
    """Count nodes for *slug* built by exactly *commit_sha* (``None`` matches None).

    The commit-scoped count the resume skip predicate keys on: "the graph
    for THIS commit is built" is ``count_graph_nodes(slug, commit_sha=X) >
    0``.

    ``commit_sha`` here is an EXACT match — ``None`` counts the rows stamped
    NULL, it does not mean "any commit". :class:`CommitScope` exists to keep
    that convention unambiguous on the read methods, which is why
    ``query_graph`` takes a scope rather than re-using this parameter name
    with the opposite sense; ``count_graph_nodes(slug, commit_sha=X)`` and
    ``len(query_graph(slug, scope=CommitScope.at(X)))`` agree by
    construction.
    """
    raise NotImplementedError("Graph backend is not implemented on this driver")

create_job(job: IndexingJob) -> None abstractmethod

Persist a new indexing job.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
373
374
375
@abc.abstractmethod
def create_job(self, job: IndexingJob) -> None:
    """Persist a new indexing job."""

create_project(project: Project) -> None abstractmethod

Persist a new project record.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
164
165
166
@abc.abstractmethod
def create_project(self, project: Project) -> None:
    """Persist a new project record."""

delete_credentials(slug: str) -> bool abstractmethod

Delete slug's credential; return True if one was removed, else False.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
593
594
595
@abc.abstractmethod
def delete_credentials(self, slug: str) -> bool:
    """Delete *slug*'s credential; return True if one was removed, else False."""

delete_doc_note(slug: str, page_id: str) -> bool

Delete a doc-page note; return True if one was removed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1058
1059
1060
def delete_doc_note(self, slug: str, page_id: str) -> bool:
    """Delete a doc-page note; return True if one was removed."""
    raise NotImplementedError

delete_edges_by_source_file(slug: str, file: str) -> int

Delete edges originating from any node in file; return count.

"Originating" = the edge source node_id belongs to a node whose file is file. Call BEFORE :meth:delete_nodes_by_file for the same file so the source nodes are still present to resolve.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
908
909
910
911
912
913
914
915
def delete_edges_by_source_file(self, slug: str, file: str) -> int:
    """Delete edges originating from any node in *file*; return count.

    "Originating" = the edge ``source`` node_id belongs to a node whose
    ``file`` is *file*. Call BEFORE :meth:`delete_nodes_by_file` for the
    same file so the source nodes are still present to resolve.
    """
    raise NotImplementedError("Memory layer methods land in the memory store")

delete_file_manifest(slug: str, path: str) -> bool

Delete a file-manifest entry; return True if one was removed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1078
1079
1080
def delete_file_manifest(self, slug: str, path: str) -> bool:
    """Delete a file-manifest entry; return True if one was removed."""
    raise NotImplementedError

delete_memory_node(slug: str, node_id: str) -> bool

Delete a memory node + its embedding; return True if one was removed.

Edges are NOT touched (callers invalidate them separately so history survives). Used when a merge supersedes a note under a new identity.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
931
932
933
934
935
936
937
def delete_memory_node(self, slug: str, node_id: str) -> bool:
    """Delete a memory node + its embedding; return True if one was removed.

    Edges are NOT touched (callers invalidate them separately so history
    survives). Used when a merge supersedes a note under a new identity.
    """
    raise NotImplementedError

delete_nodes_by_file(slug: str, file: str) -> int

Delete every code node in file AND its vectors; return the NODE count.

The embedding cascade is part of this method rather than a second call the caller makes first, because :class:Embedding carries no file field: a vector is reachable only through the node it points at, so once that node is deleted the vector can no longer be found by file, by commit, or by anything else — it is simply unreachable, unreapable, and still scoreable by :meth:vector_search. Cascading here makes "no orphaned vectors" a property of the store instead of a call-ordering discipline every future caller has to rediscover.

The return value stays the NODE count so the delete reads the same as every other scoped delete on this class.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
def delete_nodes_by_file(self, slug: str, file: str) -> int:
    """Delete every code node in *file* AND its vectors; return the NODE count.

    The embedding cascade is part of this method rather than a second call
    the caller makes first, because :class:`Embedding` carries no ``file``
    field: a vector is reachable only through the node it points at, so once
    that node is deleted the vector can no longer be found by file, by
    commit, or by anything else — it is simply unreachable, unreapable, and
    still scoreable by :meth:`vector_search`. Cascading here makes "no
    orphaned vectors" a property of the store instead of a call-ordering
    discipline every future caller has to rediscover.

    The return value stays the NODE count so the delete reads the same as
    every other scoped delete on this class.
    """
    raise NotImplementedError("Memory layer methods land in the memory store")

delete_page(slug: str, page_id: str) -> bool abstractmethod

Delete a single wiki page. Returns True if a page was removed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
367
368
369
@abc.abstractmethod
def delete_page(self, slug: str, page_id: str) -> bool:
    """Delete a single wiki page. Returns True if a page was removed."""

delete_project(slug: str) -> bool abstractmethod

Delete project slug; return True if deleted, False if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
176
177
178
@abc.abstractmethod
def delete_project(self, slug: str) -> bool:
    """Delete project *slug*; return True if deleted, False if absent."""

delete_project_settings(slug: str) -> bool abstractmethod

Delete slug's settings record; return True if one existed.

Called on project delete so a re-created slug can't inherit the dead project's settings (the rule the freshness cache eviction already follows).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
248
249
250
251
252
253
254
@abc.abstractmethod
def delete_project_settings(self, slug: str) -> bool:
    """Delete *slug*'s settings record; return True if one existed.

    Called on project delete so a re-created slug can't inherit the dead
    project's settings (the rule the freshness cache eviction already follows).
    """

Top-k entity embeddings by cosine (the ANN block seam for ER).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1131
1132
1133
1134
1135
def entity_vector_search(
    self, slug: str, qvec: list[float], k: int = 10
) -> list[EntityEmbedding]:
    """Top-k entity embeddings by cosine (the ANN block seam for ER)."""
    raise NotImplementedError

find_job_by_session(session_id: str) -> str | None abstractmethod

Reverse lookup: return the job_id for session_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
483
484
485
@abc.abstractmethod
def find_job_by_session(self, session_id: str) -> str | None:
    """Reverse lookup: return the job_id for *session_id*, or None."""

find_qa_by_session(session_id: str) -> str | None abstractmethod

Reverse lookup: return the answer_id for session_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
673
674
675
@abc.abstractmethod
def find_qa_by_session(self, session_id: str) -> str | None:
    """Reverse lookup: return the answer_id for *session_id*, or None."""

get_act_plan(job_id: str) -> dict[str, Any] | None

Return the persisted act-stage record, or None if stage 1 hasn't finished.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
525
526
527
def get_act_plan(self, job_id: str) -> dict[str, Any] | None:
    """Return the persisted act-stage record, or None if stage 1 hasn't finished."""
    return None

get_credentials(slug: str) -> dict[str, Any] | None abstractmethod

Return the encoded credential blob for slug, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
589
590
591
@abc.abstractmethod
def get_credentials(self, slug: str) -> dict[str, Any] | None:
    """Return the encoded credential blob for *slug*, or None if absent."""

get_doc_note(slug: str, page_id: str) -> DocPageNote | None

Return a single doc-page note, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1050
1051
1052
def get_doc_note(self, slug: str, page_id: str) -> DocPageNote | None:
    """Return a single doc-page note, or None if absent."""
    raise NotImplementedError

get_entity(slug: str, entity_id: str) -> Entity | None

Return a single entity, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1101
1102
1103
def get_entity(self, slug: str, entity_id: str) -> Entity | None:
    """Return a single entity, or None if absent."""
    raise NotImplementedError

get_entity_recommendations(slug: str) -> list[EntityRecommendation]

Return every persisted entity recommendation for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1166
1167
1168
def get_entity_recommendations(self, slug: str) -> list[EntityRecommendation]:
    """Return every persisted entity recommendation for *slug*."""
    raise NotImplementedError

get_file_manifest(slug: str, path: str) -> FileManifest | None

Return a single file-manifest entry, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1070
1071
1072
def get_file_manifest(self, slug: str, path: str) -> FileManifest | None:
    """Return a single file-manifest entry, or None if absent."""
    raise NotImplementedError

get_job(job_id: str) -> IndexingJob | None abstractmethod

Return the indexing job, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
377
378
379
@abc.abstractmethod
def get_job(self, job_id: str) -> IndexingJob | None:
    """Return the indexing job, or None if absent."""

get_job_page_ids(slug: str, job_id: str) -> frozenset[str] abstractmethod

Return the page ids job_id wrote — its claim set, else attribution.

The fallback is what makes a PRE-EXISTING interrupted job resumable: its meta carries a bare counter and no claim record, so reading the claim alone reported nothing done and the resume regenerated every page it had already written correctly — the exact waste the claim exists to prevent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
557
558
559
560
561
562
563
564
565
@abc.abstractmethod
def get_job_page_ids(self, slug: str, job_id: str) -> frozenset[str]:
    """Return the page ids *job_id* wrote — its claim set, else attribution.

    The fallback is what makes a PRE-EXISTING interrupted job resumable: its
    meta carries a bare counter and no claim record, so reading the claim
    alone reported nothing done and the resume regenerated every page it had
    already written correctly — the exact waste the claim exists to prevent.
    """

get_job_plan(job_id: str) -> list[dict[str, Any]] | None abstractmethod

Return the page-plan list, or None if no plan has been committed yet.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
493
494
495
@abc.abstractmethod
def get_job_plan(self, job_id: str) -> list[dict[str, Any]] | None:
    """Return the page-plan list, or None if no plan has been committed yet."""

get_job_session(job_id: str) -> str | None abstractmethod

Return the session_id attached to job_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
479
480
481
@abc.abstractmethod
def get_job_session(self, job_id: str) -> str | None:
    """Return the session_id attached to *job_id*, or None."""

get_job_submission(job_id: str) -> dict[str, Any] | None abstractmethod

Return the persisted submission dict, or None if not yet saved.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
579
580
581
@abc.abstractmethod
def get_job_submission(self, job_id: str) -> dict[str, Any] | None:
    """Return the persisted submission dict, or None if not yet saved."""

get_job_submitted_count(job_id: str) -> int abstractmethod

Return the number of pages submitted so far for job_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
529
530
531
@abc.abstractmethod
def get_job_submitted_count(self, job_id: str) -> int:
    """Return the number of pages submitted so far for *job_id*."""

get_memory_node(slug: str, node_id: str) -> MemoryNode | None

Return a single memory node, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
927
928
929
def get_memory_node(self, slug: str, node_id: str) -> MemoryNode | None:
    """Return a single memory node, or None if absent."""
    raise NotImplementedError

get_page(slug: str, page_id: str) -> WikiPage | None

Return a single wiki page, or None if absent.

THE single doc-content read seam: every upstream doc reader (the API page route, the wiki_read_page Q&A tool, the MCP read_wiki_page facade over the route) funnels through here, so guarding it once makes a graph-only project's "no documentation" failure deterministic everywhere. A project indexed in graph-only (developer) mode carries graph_only=True and has ZERO pages — reading page content then raises :class:DocumentationUnavailableError from this ONE place rather than returning a confusing None. The driver-specific fetch lives in :meth:_get_page_raw; this template method only adds the guard. The graph endpoint never calls this, so visualisation stays unaffected.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
def get_page(self, slug: str, page_id: str) -> WikiPage | None:
    """Return a single wiki page, or None if absent.

    THE single doc-content read seam: every upstream doc reader (the API
    page route, the ``wiki_read_page`` Q&A tool, the MCP ``read_wiki_page``
    facade over the route) funnels through here, so guarding it once makes a
    graph-only project's "no documentation" failure deterministic everywhere.
    A project indexed in graph-only (developer) mode carries
    ``graph_only=True`` and has ZERO pages — reading page content then raises
    :class:`DocumentationUnavailableError` from this ONE place rather than
    returning a confusing ``None``. The driver-specific fetch lives in
    :meth:`_get_page_raw`; this template method only adds the guard. The graph
    endpoint never calls this, so visualisation stays unaffected.
    """
    self._assert_docs_available(slug)
    return self._get_page_raw(slug, page_id)

get_project(slug: str) -> Project | None abstractmethod

Return the project for slug, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
168
169
170
@abc.abstractmethod
def get_project(self, slug: str) -> Project | None:
    """Return the project for *slug*, or None if absent."""

get_project_settings(slug: str) -> ProjectSettings | None abstractmethod

Return slug's settings record, or None when it has never been written.

None is a NORMAL state, not an error — the caller falls back to the per-job submission scan.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
240
241
242
243
244
245
246
@abc.abstractmethod
def get_project_settings(self, slug: str) -> ProjectSettings | None:
    """Return *slug*'s settings record, or None when it has never been written.

    ``None`` is a NORMAL state, not an error — the caller falls back to
    the per-job submission scan.
    """

get_qa(answer_id: str) -> QaAnswer | None abstractmethod

Return the QA answer, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
651
652
653
@abc.abstractmethod
def get_qa(self, answer_id: str) -> QaAnswer | None:
    """Return the QA answer, or None if absent."""

get_qa_session(answer_id: str) -> str | None abstractmethod

Return the session_id attached to answer_id, or None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
669
670
671
@abc.abstractmethod
def get_qa_session(self, answer_id: str) -> str | None:
    """Return the session_id attached to *answer_id*, or None."""

get_recovery_attempts(slug: str) -> int abstractmethod

Return the recovery-attempt count for slug (0 if never recovered).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
608
609
610
@abc.abstractmethod
def get_recovery_attempts(self, slug: str) -> int:
    """Return the recovery-attempt count for *slug* (0 if never recovered)."""

get_resume_plan(job_id: str) -> dict[str, Any] | None

Return the persisted resume-plan dict, or None if the job isn't resuming.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
507
508
509
def get_resume_plan(self, job_id: str) -> dict[str, Any] | None:
    """Return the persisted resume-plan dict, or None if the job isn't resuming."""
    return None

latest_job(slug: str, *, statuses: Iterable[str] | None = None) -> IndexingJob | None

Return slug's most recent job, ordered by phase_started_at.

The ONE "what is the latest attempt for this slug" answer — a thin convenience over :meth:list_jobs, which already returns newest-first by phase_started_at (never job_id, a uuid4 hex that sorts RANDOMLY with respect to when a job actually ran). statuses narrows the candidates before taking the first, e.g. {"complete"} for a freshness or QA-source baseline. Returns None when no job (matching statuses, if given) exists.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
def latest_job(
    self, slug: str, *, statuses: Iterable[str] | None = None
) -> IndexingJob | None:
    """Return *slug*'s most recent job, ordered by ``phase_started_at``.

    The ONE "what is the latest attempt for this slug" answer — a thin
    convenience over :meth:`list_jobs`, which already returns newest-first
    by ``phase_started_at`` (never ``job_id``, a ``uuid4`` hex that sorts
    RANDOMLY with respect to when a job actually ran). *statuses* narrows
    the candidates before taking the first, e.g. ``{"complete"}`` for a
    freshness or QA-source baseline. Returns ``None`` when no job
    (matching *statuses*, if given) exists.
    """
    candidates = self.list_jobs(slug=slug)
    if statuses is not None:
        wanted = set(statuses)
        candidates = [j for j in candidates if j.status in wanted]
    return candidates[0] if candidates else None

list_credentials() -> dict[str, dict[str, Any]] abstractmethod

Return every stored credential blob keyed by scope (slug or bare host).

The management surface (CredentialStore.list → the /v1/git/credentials route) reads this. Values are the raw encoded blobs; the caller decodes + redacts. The scope key is a full slug (host/owner/repo) or a bare host.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
597
598
599
600
601
602
603
604
@abc.abstractmethod
def list_credentials(self) -> dict[str, dict[str, Any]]:
    """Return every stored credential blob keyed by scope (slug or bare host).

    The management surface (``CredentialStore.list`` → the ``/v1/git/credentials``
    route) reads this. Values are the raw encoded blobs; the caller decodes +
    redacts. The scope key is a full slug (``host/owner/repo``) or a bare host.
    """

list_doc_notes(slug: str) -> list[DocPageNote]

Return every doc-page note for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1054
1055
1056
def list_doc_notes(self, slug: str) -> list[DocPageNote]:
    """Return every doc-page note for *slug*."""
    raise NotImplementedError

list_edges(slug: str, *, scope: CommitScope) -> list[GraphEdge]

Return slug's edges for scope's generation (graph-viewer endpoint).

scope is required for the same reason as on query_graph — and it must agree with the scope the nodes were read under, or the view is assembled from edges whose endpoints belong to a different generation.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
874
875
876
877
878
879
880
881
def list_edges(self, slug: str, *, scope: CommitScope) -> list[GraphEdge]:
    """Return *slug*'s edges for *scope*'s generation (graph-viewer endpoint).

    *scope* is required for the same reason as on ``query_graph`` — and it
    must agree with the scope the nodes were read under, or the view is
    assembled from edges whose endpoints belong to a different generation.
    """
    raise NotImplementedError("Graph backend is not implemented on this driver")

list_entity_edges(slug: str, *, source_id: str | None = None) -> list[EntityRelation]

Return entity relations, optionally scoped to source_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1148
1149
1150
1151
1152
def list_entity_edges(
    self, slug: str, *, source_id: str | None = None
) -> list[EntityRelation]:
    """Return entity relations, optionally scoped to ``source_id``."""
    raise NotImplementedError

list_file_manifest(slug: str) -> list[FileManifest]

Return every file-manifest entry for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1074
1075
1076
def list_file_manifest(self, slug: str) -> list[FileManifest]:
    """Return every file-manifest entry for *slug*."""
    raise NotImplementedError

list_jobs(slug: str | None = None) -> list[IndexingJob] abstractmethod

Return all jobs, optionally filtered to slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
421
422
423
@abc.abstractmethod
def list_jobs(self, slug: str | None = None) -> list[IndexingJob]:
    """Return all jobs, optionally filtered to *slug*."""

list_memory_edges(slug: str, *, node_id: str | None = None, include_invalidated: bool = False) -> list[MemoryEdge]

Return memory edges, optionally scoped to source == node_id.

Invalidated edges (invalid_at set) are excluded unless include_invalidated is True.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
949
950
951
952
953
954
955
956
957
958
959
960
961
def list_memory_edges(
    self,
    slug: str,
    *,
    node_id: str | None = None,
    include_invalidated: bool = False,
) -> list[MemoryEdge]:
    """Return memory edges, optionally scoped to ``source == node_id``.

    Invalidated edges (``invalid_at`` set) are excluded unless
    *include_invalidated* is True.
    """
    raise NotImplementedError

list_pages(slug: str) -> list[WikiPage] abstractmethod

Return all pages for project slug.

NOT doc-guarded: list_pages is also the page-id roster used by the deterministic internal paths (finalize prune, resume, retriever, QaFinalizer.tag_page_citations), so guarding it would break indexing and finalize themselves. A graph-only project simply returns [] here (it has no pages); the guard lives on the doc-CONTENT read (:meth:get_page).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
341
342
343
344
345
346
347
348
349
350
@abc.abstractmethod
def list_pages(self, slug: str) -> list[WikiPage]:
    """Return all pages for project *slug*.

    NOT doc-guarded: ``list_pages`` is also the page-id roster used by the
    deterministic internal paths (finalize prune, resume, retriever,
    ``QaFinalizer.tag_page_citations``), so guarding it would break indexing
    and finalize themselves. A graph-only project simply returns ``[]`` here
    (it has no pages); the guard lives on the doc-CONTENT read (:meth:`get_page`).
    """

list_projects() -> list[Project] abstractmethod

Return all projects sorted by indexed_at descending.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
172
173
174
@abc.abstractmethod
def list_projects(self) -> list[Project]:
    """Return all projects sorted by indexed_at descending."""

list_qa(status: str | None = None) -> list[QaAnswer] abstractmethod

Return all QA answers, optionally filtered to status.

Boot-time/offline use only (mirrors :meth:list_jobs) — an unfiltered call reads every persisted answer, so a caller on an interactive path must narrow with status rather than filtering the full list in Python.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
655
656
657
658
659
660
661
662
663
@abc.abstractmethod
def list_qa(self, status: str | None = None) -> list[QaAnswer]:
    """Return all QA answers, optionally filtered to *status*.

    Boot-time/offline use only (mirrors :meth:`list_jobs`) — an
    unfiltered call reads every persisted answer, so a caller on an
    interactive path must narrow with *status* rather than filtering the
    full list in Python.
    """

live_scope(slug: str) -> CommitScope

The generation slug's readers should see: the project's own commit.

A project row carries the commit its last completed index built from, so that — not the union of every commit ever indexed — is what "the code as it is now" means. Lives here because the store is what holds the project row; every graph reader already has one, so nobody has to thread a commit sha through their own call chain to read correctly.

A project with no commit_sha (a commit-less catalog ingestion, or a slug with no project row yet) has exactly one generation, so scoping it could only be a way to get it wrong — the union and the live generation are the same set, and every() says so without asserting a commit that does not exist.

Cost: O(one record).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
def live_scope(self, slug: str) -> CommitScope:
    """The generation *slug*'s readers should see: the project's own commit.

    A project row carries the commit its last completed index built from,
    so that — not the union of every commit ever indexed — is what "the
    code as it is now" means. Lives here because the store is what holds
    the project row; every graph reader already has one, so nobody has to
    thread a commit sha through their own call chain to read correctly.

    A project with no ``commit_sha`` (a commit-less catalog ingestion, or a
    slug with no project row yet) has exactly one generation, so scoping it
    could only be a way to get it wrong — the union and the live generation
    are the same set, and ``every()`` says so without asserting a commit
    that does not exist.

    Cost: ``O(one record)``.
    """
    project = self.get_project(slug)
    if project is None or project.commit_sha is None:
        return CommitScope.every()
    return CommitScope.at(project.commit_sha)

load_job_events(job_id: str, after_idx: int = -1) -> list[dict[str, Any]] abstractmethod

Return job events with idx > after_idx (-1 returns all).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
465
466
467
468
469
@abc.abstractmethod
def load_job_events(
    self, job_id: str, after_idx: int = -1
) -> list[dict[str, Any]]:
    """Return job events with idx > *after_idx* (-1 returns all)."""

load_qa_events(answer_id: str, after_idx: int = -1) -> list[dict[str, Any]] abstractmethod

Return QA events with idx > after_idx (-1 returns all).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
681
682
683
684
685
@abc.abstractmethod
def load_qa_events(
    self, answer_id: str, after_idx: int = -1
) -> list[dict[str, Any]]:
    """Return QA events with idx > *after_idx* (-1 returns all)."""

memories_anchored_to(slug: str, entity_keys: Iterable[EntityKey], *, include_invalidated: bool = False) -> list[str]

Reverse ANCHORS lookup: entity_keys → distinct memory node_ids.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
963
964
965
966
967
968
969
970
971
def memories_anchored_to(
    self,
    slug: str,
    entity_keys: Iterable[EntityKey],
    *,
    include_invalidated: bool = False,
) -> list[str]:
    """Reverse ANCHORS lookup: entity_keys → distinct memory node_ids."""
    raise NotImplementedError

Top-k memory embeddings by cosine, after applying filt.

Scale seam — keep signature stable. v1 is brute-force cosine; IVF / Matryoshka / quantization slot in here without touching callers.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
979
980
981
982
983
984
985
986
987
988
989
990
991
992
def memory_vector_search(
    self,
    slug: str,
    qvec: list[float],
    k: int = 10,
    *,
    filt: MemoryFilter | None = None,
) -> list[MemoryEmbedding]:
    """Top-k memory embeddings by cosine, after applying *filt*.

    Scale seam — keep signature stable. v1 is brute-force cosine; IVF /
    Matryoshka / quantization slot in here without touching callers.
    """
    raise NotImplementedError

page_ids_for_job(slug: str, job_id: str) -> frozenset[str] abstractmethod

Page ids under slug whose stored attribution names job_id.

save_page stamps attribution independently of the claim record, which is why this can answer for a job the claim record cannot.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
567
568
569
570
571
572
573
@abc.abstractmethod
def page_ids_for_job(self, slug: str, job_id: str) -> frozenset[str]:
    """Page ids under *slug* whose stored attribution names *job_id*.

    ``save_page`` stamps attribution independently of the claim record,
    which is why this can answer for a job the claim record cannot.
    """

prune_pages(slug: str, keep: Iterable[str]) -> int

Drop every page for slug whose page_id is not in keep.

Default impl uses list_pages + per-page delete_page so backends only need a single primitive. Returns the number of pages dropped.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
352
353
354
355
356
357
358
359
360
361
362
363
364
365
def prune_pages(self, slug: str, keep: Iterable[str]) -> int:
    """Drop every page for *slug* whose ``page_id`` is not in *keep*.

    Default impl uses ``list_pages`` + per-page ``delete_page`` so
    backends only need a single primitive. Returns the number of
    pages dropped.
    """
    keep_set = set(keep)
    dropped = 0
    for page in self.list_pages(slug):
        if page.id not in keep_set:
            self.delete_page(slug, page.id)
            dropped += 1
    return dropped

query_entities(slug: str, *, filt: EntityFilter | None = None) -> list[Entity]

Return entities matching filt's facets (no filter ⇒ all).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1105
1106
1107
1108
1109
def query_entities(
    self, slug: str, *, filt: EntityFilter | None = None
) -> list[Entity]:
    """Return entities matching *filt*'s facets (no filter ⇒ all)."""
    raise NotImplementedError

query_graph(slug: str, *, scope: CommitScope, node_type: str | None = None, name_match: str | None = None, neighbors_of: str | None = None, node_ids: Collection[str] | None = None) -> list[GraphNode]

Query the code graph for slug, restricted to scope's generation.

node_ids fetches an explicit set by id. It exists so a caller holding a bounded list of ids — the knowledge-graph view repairing entity anchors that point into a superseded generation — can look exactly those up instead of reading a whole generation to find a handful. An empty collection returns nothing; None means "no id filter".

scope is REQUIRED and has no default on purpose. This store holds the UNION of every commit ever indexed for a slug, so "which generation" has no safe default — and a default of CommitScope.every() would be precisely the fail-open filter the root guidance forbids: a caller that forgot to scope would silently read every generation and still return 200. Required means a missed call site is a typecheck failure instead. See :class:CommitScope for why this is not a commit_sha parameter (count_graph_nodes already owns that name with the opposite meaning for None).

Cost: O(nodes for the slug in scope) — a read of one project's graph generation, not of all history.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
def query_graph(
    self,
    slug: str,
    *,
    scope: CommitScope,
    node_type: str | None = None,
    name_match: str | None = None,
    neighbors_of: str | None = None,
    node_ids: Collection[str] | None = None,
) -> list[GraphNode]:
    """Query the code graph for *slug*, restricted to *scope*'s generation.

    *node_ids* fetches an explicit set by id. It exists so a caller holding
    a bounded list of ids — the knowledge-graph view repairing entity
    anchors that point into a superseded generation — can look exactly
    those up instead of reading a whole generation to find a handful. An
    empty collection returns nothing; ``None`` means "no id filter".

    *scope* is REQUIRED and has no default on purpose. This store holds the
    UNION of every commit ever indexed for a slug, so "which generation"
    has no safe default — and a default of ``CommitScope.every()`` would be
    precisely the fail-open filter the root guidance forbids: a caller that
    forgot to scope would silently read every generation and still return
    ``200``. Required means a missed call site is a typecheck failure
    instead. See :class:`CommitScope` for why this is not a ``commit_sha``
    parameter (``count_graph_nodes`` already owns that name with the
    opposite meaning for ``None``).

    Cost: ``O(nodes for the slug in scope)`` — a read of one project's
    graph generation, not of all history.
    """
    raise NotImplementedError("Graph backend is not implemented on this driver")

query_memory(slug: str, *, filt: MemoryFilter | None = None) -> list[MemoryNode]

Return memory nodes matching filt's node-level facets.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
939
940
941
942
943
def query_memory(
    self, slug: str, *, filt: MemoryFilter | None = None
) -> list[MemoryNode]:
    """Return memory nodes matching *filt*'s node-level facets."""
    raise NotImplementedError

reap_slug(slug: str) -> dict[str, int]

Delete every family this store persists for slug; return the counts.

The exceptions are the three the delete-project ROUTE already owns directly: wiki_projects (:meth:delete_project), wiki_settings (:meth:delete_project_settings), and a git credential (CredentialStore.delete — repo-scope ONLY; a host-scoped credential is shared by every repo on that host and must NEVER cascade off one project's delete, so this method never touches credentials at all).

Everything else a completed or in-flight index could have written under slug is deleted here: pages, the code graph (nodes/edges/embeddings), the entity layer (entities/entity-edges/entity-embeddings/ recommendations), the memory layer (nodes/edges/embeddings), doc notes, the file manifest, the recovery-attempt counter, every indexing job (with its plan/resume/submission/session sidecars and its event log), and every QA answer for this slug (with its event log). Without this, a project delete orphans the large majority of what an index ever wrote — it stays reachable by nothing, forever, and a slug reused later inherits none of it (the same "clean slate" reasoning behind clearing settings/ credentials/the freshness cache above, extended to the rest of the store).

Returns per-family deleted-row counts (mirroring :meth:supersede_graph_artifacts's shape) so a caller can report exactly what was removed. Idempotent: a second call for an already-reaped slug finds nothing and returns all-zero counts.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
def reap_slug(self, slug: str) -> dict[str, int]:
    """Delete every family this store persists for *slug*; return the counts.

    The exceptions are the three the delete-project ROUTE already owns
    directly: ``wiki_projects``
    (:meth:`delete_project`), ``wiki_settings`` (:meth:`delete_project_settings`),
    and a git credential (``CredentialStore.delete`` — repo-scope ONLY; a
    host-scoped credential is shared by every repo on that host and must
    NEVER cascade off one project's delete, so this method never touches
    credentials at all).

    Everything else a completed or in-flight index could have written under
    *slug* is deleted here: pages, the code graph (nodes/edges/embeddings),
    the entity layer (entities/entity-edges/entity-embeddings/
    recommendations), the memory layer (nodes/edges/embeddings), doc notes,
    the file manifest, the recovery-attempt counter, every indexing job
    (with its plan/resume/submission/session sidecars and its event log),
    and every QA answer for this slug (with its event log). Without this, a
    project delete orphans the large majority of what an index ever wrote —
    it stays reachable by nothing, forever, and a slug reused later inherits
    none of it (the same "clean slate" reasoning behind clearing settings/
    credentials/the freshness cache above, extended to the rest of the store).

    Returns per-family deleted-row counts (mirroring
    :meth:`supersede_graph_artifacts`'s shape) so a caller can report exactly
    what was removed. Idempotent: a second call for an already-reaped slug
    finds nothing and returns all-zero counts.
    """
    raise NotImplementedError("Graph backend is not implemented on this driver")

reset_recovery_attempts(slug: str) -> None

Clear slug's recovery counter (a user-initiated resume gets a fresh budget).

A human asking to retry an index must not be blocked by prior automatic re-drives, so the manual resume path resets the auto-recovery cap. Concrete default no-op so a backend that never tracks the counter is unaffected.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
622
623
624
625
626
627
628
def reset_recovery_attempts(self, slug: str) -> None:
    """Clear *slug*'s recovery counter (a user-initiated resume gets a fresh budget).

    A human asking to retry an index must not be blocked by prior automatic
    re-drives, so the manual resume path resets the auto-recovery cap. Concrete
    default no-op so a backend that never tracks the counter is unaffected.
    """

restamp_graph_artifacts(slug: str, *, from_commit: str, to_commit: str) -> dict[str, int]

Carry from_commit's surviving artifacts forward onto to_commit.

Concrete-raising rather than @abc.abstractmethod, matching its sibling supersede_graph_artifacts and the rest of this graph family: an abstract method here is inherited by every partial test double of this base and stops it being instantiable at all, which turns a new store capability into a broad, unrelated test failure.

The counterpart a SCOPED re-index needs and a full one does not. A full index re-stamps every artifact by rewriting them all, so afterwards one generation describes the whole slug and live_scope finds everything. An incremental refresh re-stamps only what it re-parsed, so without this the untouched majority keeps the PREVIOUS commit while the project row advances — and since live_scope is at(project.commit_sha), every reader (the graph view, the retriever, the agent graph tools) would see only the handful of files that happened to change. Not deleted, not erroring: invisible. That is the failure this method exists to prevent, and it is the precondition the delta indexer's own unscoped reads name.

The semantic is a claim, and it is a true one: an artifact built from a file that did NOT change describes to_commit just as accurately as it described from_commit, so moving its stamp forward asserts nothing false.

Why this is a MOVE from a named generation rather than "stamp everything". The store can hold generations older than from_commit — a supersede that never ran, or the field-absent rows the isolation backfill exists for. Those describe code that is genuinely gone. Blanket-stamping them onto to_commit would resurrect deleted symbols into the live view, which is a worse bug than the one being fixed. Rows stamped None are likewise left alone, matching supersede_graph_artifacts' preserve rule.

Returns per-family moved counts. Idempotent: a second call finds nothing left at from_commit.

Cost: O(artifacts at from_commit) — an offline finalize step.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
831
832
833
834
835
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
865
866
867
868
869
870
871
872
def restamp_graph_artifacts(
    self, slug: str, *, from_commit: str, to_commit: str
) -> dict[str, int]:
    """Carry *from_commit*'s surviving artifacts forward onto *to_commit*.

    Concrete-raising rather than ``@abc.abstractmethod``, matching its
    sibling ``supersede_graph_artifacts`` and the rest of this graph family:
    an abstract method here is inherited by every partial test double of
    this base and stops it being instantiable at all, which turns a new
    store capability into a broad, unrelated test failure.

    The counterpart a SCOPED re-index needs and a full one does not. A full
    index re-stamps every artifact by rewriting them all, so afterwards one
    generation describes the whole slug and ``live_scope`` finds everything.
    An incremental refresh re-stamps only what it re-parsed, so without this
    the untouched majority keeps the PREVIOUS commit while the project row
    advances — and since ``live_scope`` is ``at(project.commit_sha)``, every
    reader (the graph view, the retriever, the agent graph tools) would see
    only the handful of files that happened to change. Not deleted, not
    erroring: invisible. That is the failure this method exists to prevent,
    and it is the precondition the delta indexer's own unscoped reads name.

    **The semantic is a claim, and it is a true one:** an artifact built
    from a file that did NOT change describes *to_commit* just as accurately
    as it described *from_commit*, so moving its stamp forward asserts
    nothing false.

    **Why this is a MOVE from a named generation rather than "stamp
    everything".** The store can hold generations older than *from_commit* —
    a supersede that never ran, or the field-absent rows the isolation
    backfill exists for. Those describe code that is genuinely gone.
    Blanket-stamping them onto *to_commit* would resurrect deleted symbols
    into the live view, which is a worse bug than the one being fixed. Rows
    stamped ``None`` are likewise left alone, matching
    ``supersede_graph_artifacts``' preserve rule.

    Returns per-family moved counts. Idempotent: a second call finds nothing
    left at *from_commit*.

    Cost: ``O(artifacts at from_commit)`` — an offline finalize step.
    """
    raise NotImplementedError

save_act_plan(job_id: str, plan: dict[str, Any]) -> None

Persist the act-stage record for job_id; overwrites any previous one.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
521
522
523
def save_act_plan(self, job_id: str, plan: dict[str, Any]) -> None:
    """Persist the act-stage record for *job_id*; overwrites any previous one."""
    raise NotImplementedError

save_credentials(slug: str, blob: dict[str, Any]) -> None abstractmethod

Persist the (already encoded) credential blob for slug; overwrite.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
585
586
587
@abc.abstractmethod
def save_credentials(self, slug: str, blob: dict[str, Any]) -> None:
    """Persist the (already encoded) credential *blob* for *slug*; overwrite."""

save_entity_recommendation(slug: str, rec: EntityRecommendation) -> None

Upsert a resolution-recommendation record (a prior for the next pass).

Keyed on rec.id (deterministic over action + sorted subjects + type), NOT appended: these records are read back as priors by EntityResolver, so an unkeyed insert let a replayed enrich pass state the same prior twice and re-weight the ladder purely by having run again.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
def save_entity_recommendation(
    self, slug: str, rec: EntityRecommendation
) -> None:
    """Upsert a resolution-recommendation record (a prior for the next pass).

    Keyed on ``rec.id`` (deterministic over action + sorted subjects + type),
    NOT appended: these records are read back as priors by ``EntityResolver``,
    so an unkeyed insert let a replayed enrich pass state the same prior
    twice and re-weight the ladder purely by having run again.
    """
    raise NotImplementedError

save_job_plan(job_id: str, plan: list[dict[str, Any]]) -> None abstractmethod

Persist the page-plan list for job_id; overwrites any previous plan.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
489
490
491
@abc.abstractmethod
def save_job_plan(self, job_id: str, plan: list[dict[str, Any]]) -> None:
    """Persist the page-plan list for *job_id*; overwrites any previous plan."""

save_job_submission(job_id: str, submission: dict[str, Any]) -> None abstractmethod

Persist the wizard submission dict for job_id (token must be absent).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
575
576
577
@abc.abstractmethod
def save_job_submission(self, job_id: str, submission: dict[str, Any]) -> None:
    """Persist the wizard submission dict for *job_id* (token must be absent)."""

save_page(slug: str, page: WikiPage, *, commit_sha: str | None = None, job_id: str | None = None) -> None abstractmethod

Persist page for the project slug; overwrites if same page_id.

commit_sha/job_id attribute the page to its owning index. Unlike the graph families, page attribution is a store-internal column, never a WikiPage field — the page IS a console wire type serialized whole, so keeping attribution off the model preserves that wire byte-for-byte. Page supersession is unaffected: wiki_finalize already prunes pages to the committed plan, so this is provenance, not the supersede mechanism.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
@abc.abstractmethod
def save_page(
    self,
    slug: str,
    page: WikiPage,
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Persist *page* for the project *slug*; overwrites if same page_id.

    ``commit_sha``/``job_id`` attribute the page to its owning index. Unlike
    the graph families, page attribution is a store-internal column, never a
    ``WikiPage`` field — the page IS a console wire type serialized whole, so
    keeping attribution off the model preserves that wire byte-for-byte. Page
    supersession is unaffected: ``wiki_finalize`` already prunes pages to the
    committed plan, so this is provenance, not the supersede mechanism.
    """

save_project_settings(slug: str, settings: ProjectSettings) -> None abstractmethod

Persist (upsert) the editable settings record for slug.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
236
237
238
@abc.abstractmethod
def save_project_settings(self, slug: str, settings: ProjectSettings) -> None:
    """Persist (upsert) the editable settings record for *slug*."""

save_qa(answer: QaAnswer) -> None abstractmethod

Persist a QA answer record. Use ONLY to create it (resets bookkeeping).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
632
633
634
@abc.abstractmethod
def save_qa(self, answer: QaAnswer) -> None:
    """Persist a QA answer record. Use ONLY to create it (resets bookkeeping)."""

save_resume_plan(job_id: str, plan: dict[str, Any]) -> None

Persist the resume-plan dict for job_id; overwrites any previous one.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
503
504
505
def save_resume_plan(self, job_id: str, plan: dict[str, Any]) -> None:
    """Persist the resume-plan dict for *job_id*; overwrites any previous one."""
    raise NotImplementedError

supersede_graph_artifacts(slug: str, *, keep_commit_sha: str) -> dict[str, int]

Reap prior-commit graph + entity artifacts once keep_commit_sha completes.

Deletes every node/edge/embedding/entity/entity-edge/entity-embedding for slug whose commit_sha is a REAL value other than keep_commit_sha. Rows stamped None are PRESERVED — for entities that is a QA-minted or pre-isolation record (accretive memory, not a per-commit snapshot); for code nodes there are none on a git slug (every index stamps its commit). Returns per-collection delete counts. Idempotent: a second call for the same keep_commit_sha finds nothing to reap.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
816
817
818
819
820
821
822
823
824
825
826
827
828
829
def supersede_graph_artifacts(
    self, slug: str, *, keep_commit_sha: str
) -> dict[str, int]:
    """Reap prior-commit graph + entity artifacts once *keep_commit_sha* completes.

    Deletes every node/edge/embedding/entity/entity-edge/entity-embedding for
    *slug* whose ``commit_sha`` is a REAL value other than *keep_commit_sha*.
    Rows stamped ``None`` are PRESERVED — for entities that is a QA-minted or
    pre-isolation record (accretive memory, not a per-commit snapshot); for
    code nodes there are none on a git slug (every index stamps its commit).
    Returns per-collection delete counts. Idempotent: a second call for the
    same *keep_commit_sha* finds nothing to reap.
    """
    raise NotImplementedError("Graph backend is not implemented on this driver")

update_job(job_id: str, **fields: Any) -> IndexingJob

Partially update job_id with fields; return the updated record.

Concrete on the base because the rule both drivers owe their callers is ONE rule: only the fields the caller NAMED are written. Persisting the whole merged document instead turns every write into a read-modify-write over every field, so two writers that overlap lose one another's changes even when the fields they touch are disjoint. The phase tools write progress on a 50ms-to-5s cadence while an HTTP thread writes status, so a cancel silently reverting to running needs no exotic timing to reproduce.

The narrowing lives here; making the narrow write INDIVISIBLE is per-backend (:meth:_write_job_patch), because a lock and a field-scoped update are not the same cure.

Raises KeyError when the job is absent and ValidationError when a field is unknown to the model or its value is wrong for it. Naming no field at all is a read: there is nothing to narrow, so nothing is written.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
def update_job(self, job_id: str, **fields: Any) -> IndexingJob:
    """Partially update *job_id* with *fields*; return the updated record.

    Concrete on the base because the rule both drivers owe their callers is
    ONE rule: only the fields the caller NAMED are written. Persisting the
    whole merged document instead turns every write into a read-modify-write
    over every field, so two writers that overlap lose one another's changes
    even when the fields they touch are disjoint. The phase tools write
    progress on a 50ms-to-5s cadence while an HTTP thread writes ``status``,
    so a cancel silently reverting to ``running`` needs no exotic timing to
    reproduce.

    The narrowing lives here; making the narrow write INDIVISIBLE is
    per-backend (:meth:`_write_job_patch`), because a lock and a
    field-scoped update are not the same cure.

    Raises ``KeyError`` when the job is absent and ``ValidationError`` when
    a field is unknown to the model or its value is wrong for it. Naming no
    field at all is a read: there is nothing to narrow, so nothing is
    written.
    """
    job = self.get_job(job_id)
    if job is None:
        raise KeyError(f"Job not found: {job_id}")
    if not fields:
        return job
    return self._write_job_patch(job_id, JobPatch.build(job, fields))

update_project(slug: str, fields: dict[str, Any]) -> Project | None

Apply a partial update to slug's Project; return the new state.

Cost: O(one record) — one project read and one upsert of that same record, regardless of how much the project has indexed.

Returns None when the project is absent. Only keys in :data:PROJECT_UPDATABLE are honoured — an unknown or None value is ignored, so a caller can hand over a whole PATCH body without pre-filtering (mirrors agentic_search.store.update_workspace). None meaning "not supplied" is the OPPOSITE of update_job's rule, where a named None is a value to be written — do not carry one convention onto the other.

Concrete on the base rather than per-backend: create_project is an UPSERT in both drivers, so read → model_copy → upsert needs no duplicated JSON/Mongo pair that could drift. It is a read-modify-write with the same (non-)atomicity as every other Project write in this store.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
def update_project(self, slug: str, fields: dict[str, Any]) -> Project | None:
    """Apply a partial update to *slug*'s Project; return the new state.

    Cost: ``O(one record)`` — one project read and one upsert of that same
    record, regardless of how much the project has indexed.

    Returns ``None`` when the project is absent. Only keys in
    :data:`PROJECT_UPDATABLE` are honoured — an unknown or ``None`` value is
    ignored, so a caller can hand over a whole PATCH body without pre-filtering
    (mirrors ``agentic_search.store.update_workspace``). ``None`` meaning "not
    supplied" is the OPPOSITE of ``update_job``'s rule, where a named ``None``
    is a value to be written — do not carry one convention onto the other.

    Concrete on the base rather than per-backend: ``create_project`` is an
    UPSERT in both drivers, so read → ``model_copy`` → upsert needs no
    duplicated JSON/Mongo pair that could drift. It is a read-modify-write with
    the same (non-)atomicity as every other Project write in this store.
    """
    project = self.get_project(slug)
    if project is None:
        return None
    updates = {
        key: value
        for key, value in fields.items()
        if key in self.PROJECT_UPDATABLE and value is not None
    }
    if not updates:
        return project
    updated = project.model_copy(update=updates)
    self.create_project(updated)
    return updated

update_qa_fields(answer: QaAnswer) -> None abstractmethod

Update a QA answer's content fields in place — NON-destructive.

Persists every QaAnswer field but MUST NOT disturb store bookkeeping that some backends pack alongside the record: the event_count idx counter and the session_id mapping. save_qa does a FULL replace, which on Mongo resets event_count to 0 (so the next append_qa_event collides at idx 0) AND drops session_id (breaking find_qa_by_session). Mid-stream writers (QaFinalizer) MUST use this; save_qa is for creation only. (The JSON backend keeps session + events in separate files, so for it this is just an answer.json rewrite — the divergence is why a JSON-only test cannot catch the Mongo failure.)

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
636
637
638
639
640
641
642
643
644
645
646
647
648
649
@abc.abstractmethod
def update_qa_fields(self, answer: QaAnswer) -> None:
    """Update a QA answer's content fields in place — NON-destructive.

    Persists every ``QaAnswer`` field but MUST NOT disturb store bookkeeping
    that some backends pack alongside the record: the ``event_count`` idx
    counter and the ``session_id`` mapping. ``save_qa`` does a FULL replace,
    which on Mongo resets ``event_count`` to 0 (so the next ``append_qa_event``
    collides at idx 0) AND drops ``session_id`` (breaking
    ``find_qa_by_session``). Mid-stream writers (``QaFinalizer``) MUST use
    this; ``save_qa`` is for creation only. (The JSON backend keeps session +
    events in separate files, so for it this is just an answer.json rewrite —
    the divergence is why a JSON-only test cannot catch the Mongo failure.)
    """

upsert_doc_notes(slug: str, notes: Iterable[DocPageNote]) -> None

Upsert doc-page notes; dedup by page_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1046
1047
1048
def upsert_doc_notes(self, slug: str, notes: Iterable[DocPageNote]) -> None:
    """Upsert doc-page notes; dedup by ``page_id``."""
    raise NotImplementedError

upsert_edges(slug: str, edges: Iterable[GraphEdge], *, commit_sha: str | None = None, job_id: str | None = None, on_progress: Callable[[int, int], None] | None = None) -> None

Upsert code-graph edges, reporting completed persistence batches.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
721
722
723
724
725
726
727
728
729
730
731
def upsert_edges(
    self,
    slug: str,
    edges: Iterable[GraphEdge],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
    on_progress: Callable[[int, int], None] | None = None,
) -> None:
    """Upsert code-graph edges, reporting completed persistence batches."""
    raise NotImplementedError("Graph backend is not implemented on this driver")

upsert_embeddings(slug: str, items: Iterable[Embedding], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert dense embedding vectors.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
733
734
735
736
737
738
739
740
741
742
def upsert_embeddings(
    self,
    slug: str,
    items: Iterable[Embedding],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert dense embedding vectors."""
    raise NotImplementedError("Embeddings are not implemented on this driver")

upsert_entities(slug: str, entities: Iterable[Entity], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert entities; dedup by id (= sha1(normalized_name|type)).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
def upsert_entities(
    self,
    slug: str,
    entities: Iterable[Entity],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert entities; dedup by ``id`` (= sha1(normalized_name|type))."""
    raise NotImplementedError

upsert_entity_edges(slug: str, edges: Iterable[EntityRelation], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert entity relations; dedup by id (= source|type|target).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
def upsert_entity_edges(
    self,
    slug: str,
    edges: Iterable[EntityRelation],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert entity relations; dedup by ``id`` (= source|type|target)."""
    raise NotImplementedError

upsert_entity_embeddings(slug: str, items: Iterable[EntityEmbedding], *, commit_sha: str | None = None, job_id: str | None = None) -> None

Upsert entity embedding vectors; dedup by entity_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
def upsert_entity_embeddings(
    self,
    slug: str,
    items: Iterable[EntityEmbedding],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
) -> None:
    """Upsert entity embedding vectors; dedup by ``entity_id``."""
    raise NotImplementedError

upsert_file_manifest(slug: str, entries: Iterable[FileManifest]) -> None

Upsert per-file manifest entries; dedup by path.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
1064
1065
1066
1067
1068
def upsert_file_manifest(
    self, slug: str, entries: Iterable[FileManifest]
) -> None:
    """Upsert per-file manifest entries; dedup by ``path``."""
    raise NotImplementedError

upsert_memory_edges(slug: str, edges: Iterable[MemoryEdge]) -> None

Upsert memory edges; dedup by (source, target, type).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
945
946
947
def upsert_memory_edges(self, slug: str, edges: Iterable[MemoryEdge]) -> None:
    """Upsert memory edges; dedup by ``(source, target, type)``."""
    raise NotImplementedError

upsert_memory_embeddings(slug: str, items: Iterable[MemoryEmbedding]) -> None

Upsert memory embedding vectors; dedup by node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
973
974
975
976
977
def upsert_memory_embeddings(
    self, slug: str, items: Iterable[MemoryEmbedding]
) -> None:
    """Upsert memory embedding vectors; dedup by ``node_id``."""
    raise NotImplementedError

upsert_memory_nodes(slug: str, nodes: Iterable[MemoryNode]) -> None

Upsert memory nodes; dedup by node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
923
924
925
def upsert_memory_nodes(self, slug: str, nodes: Iterable[MemoryNode]) -> None:
    """Upsert memory nodes; dedup by ``node_id``."""
    raise NotImplementedError

upsert_nodes(slug: str, nodes: Iterable[GraphNode], *, commit_sha: str | None = None, job_id: str | None = None, on_progress: Callable[[int, int], None] | None = None) -> None

Upsert code-graph nodes, reporting completed persistence batches.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
709
710
711
712
713
714
715
716
717
718
719
def upsert_nodes(
    self,
    slug: str,
    nodes: Iterable[GraphNode],
    *,
    commit_sha: str | None = None,
    job_id: str | None = None,
    on_progress: Callable[[int, int], None] | None = None,
) -> None:
    """Upsert code-graph nodes, reporting completed persistence batches."""
    raise NotImplementedError("Graph backend is not implemented on this driver")

Nearest-neighbour vector search.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
883
884
885
886
887
def vector_search(
    self, slug: str, qvec: list[float], k: int = 10
) -> list[Embedding]:
    """Nearest-neighbour vector search."""
    raise NotImplementedError("Embeddings are not implemented on this driver")

create_wiki_store() -> WikiStoreBase

Return the configured wiki store driver.

Reads storage.driver from the app config. Defaults to "json" (filesystem). Set to "mongodb" to use MongoDB.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4227
4228
4229
4230
4231
4232
4233
4234
4235
4236
def create_wiki_store() -> WikiStoreBase:
    """Return the configured wiki store driver.

    Reads ``storage.driver`` from the app config. Defaults to ``"json"``
    (filesystem). Set to ``"mongodb"`` to use MongoDB.
    """
    driver = get_config_value("storage", "driver", default="json")
    if driver == "mongodb":
        return MongoWikiStore()
    return JsonWikiStore()

get_wiki_store() -> WikiStoreBase

Return the process-wide wiki store, constructing it on first use.

The single instance both the API routes and the wiki SessionTools share — the same singleton+factory+reset_for_tests shape as the SCG store and the run store. It lets the relocated plugins reach the store down through this factory instead of up through the API runtime; the JSON/Mongo backend is config-addressed, so a fresh instance still sees the same data.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4246
4247
4248
4249
4250
4251
4252
4253
4254
4255
4256
4257
4258
def get_wiki_store() -> WikiStoreBase:
    """Return the process-wide wiki store, constructing it on first use.

    The single instance both the API routes and the wiki SessionTools share
    — the same singleton+factory+``reset_for_tests`` shape as the SCG store
    and the run store. It lets the relocated plugins reach the store **down**
    through this factory instead of up through the API runtime; the JSON/Mongo
    backend is config-addressed, so a fresh instance still sees the same data.
    """
    global _WIKI_STORE
    if _WIKI_STORE is None:
        _WIKI_STORE = create_wiki_store()
    return _WIKI_STORE

reset_for_tests(root_dir: str | Path | None = None) -> WikiStoreBase

Swap in a fresh JSON store (under root_dir if given) for test isolation.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4267
4268
4269
4270
4271
def reset_for_tests(root_dir: str | Path | None = None) -> WikiStoreBase:
    """Swap in a fresh JSON store (under *root_dir* if given) for test isolation."""
    store = JsonWikiStore(root_dir=root_dir) if root_dir is not None else JsonWikiStore()
    set_wiki_store(store)
    return store

set_wiki_store(store: WikiStoreBase | None) -> None

Pin the process-wide wiki store (API startup wiring / test injection).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/store.py
4261
4262
4263
4264
def set_wiki_store(store: WikiStoreBase | None) -> None:
    """Pin the process-wide wiki store (API startup wiring / test injection)."""
    global _WIKI_STORE
    _WIKI_STORE = store

mewbo_graph.wiki.types

Pydantic v2 mirrors of the frontend wiki API wire types.

Every model here corresponds 1-to-1 with a TypeScript interface or type alias declared in apps/mewbo_console/src/components/wiki/api/types.ts.

Conventions: - model_config = ConfigDict(extra="forbid", populate_by_name=True) - Python attributes are snake_case; camelCase wire names use Field(alias=...). - Discriminated unions are wrapped in RootModel for model_validate access.

AccordionBlock

Bases: BaseModel

Accordion (collapsible) block.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
308
309
310
311
312
313
314
class AccordionBlock(BaseModel):
    """Accordion (collapsible) block."""

    model_config = _CFG
    kind: Literal["accordion"]
    title: str
    items: list[str]

BlockCloseEvent

Bases: BaseModel

Emitted when the current block is finalised.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1704
1705
1706
1707
1708
1709
class BlockCloseEvent(BaseModel):
    """Emitted when the current block is finalised."""

    model_config = _CFG
    type: Literal["block_close"]
    index: int

BlockDeltaEvent

Bases: BaseModel

Emitted for each text chunk appended to the current block.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1695
1696
1697
1698
1699
1700
1701
class BlockDeltaEvent(BaseModel):
    """Emitted for each text chunk appended to the current block."""

    model_config = _CFG
    type: Literal["block_delta"]
    index: int
    text_append: str = Field(alias="textAppend")

BlockOpenEvent

Bases: BaseModel

Emitted when a new block starts streaming.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1686
1687
1688
1689
1690
1691
1692
class BlockOpenEvent(BaseModel):
    """Emitted when a new block starts streaming."""

    model_config = _CFG
    type: Literal["block_open"]
    index: int
    block: BlockUnion

BlockUnion

Bases: RootModel[_BlockAnnotated]

Discriminated union of all block kinds; use BlockUnion.model_validate.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
356
357
class BlockUnion(RootModel[_BlockAnnotated]):
    """Discriminated union of all block kinds; use ``BlockUnion.model_validate``."""

CancelledEvent

Bases: BaseModel

Terminal event: indexing was cancelled.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1471
1472
1473
1474
1475
class CancelledEvent(BaseModel):
    """Terminal event: indexing was cancelled."""

    model_config = _CFG
    type: Literal["cancelled"]

CatalogDocument

Bases: BaseModel

One programmatically-ingested catalog record (a product, FAQ, doc, …).

The wire shape POST /v1/wiki/projects/{slug}/documents accepts. Each record becomes BOTH a WikiPage (BM25 + wiki_search_pages) AND a graph node carrying the text (embeddings + wiki_code_search) so the existing :class:HybridRetriever grounds it with no pipeline change.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
class CatalogDocument(BaseModel):
    """One programmatically-ingested catalog record (a product, FAQ, doc, …).

    The wire shape ``POST /v1/wiki/projects/{slug}/documents`` accepts. Each
    record becomes BOTH a ``WikiPage`` (BM25 + ``wiki_search_pages``) AND a
    graph node carrying the text (embeddings + ``wiki_code_search``) so the
    existing :class:`HybridRetriever` grounds it with no pipeline change.
    """

    model_config = _CFG

    id: str = Field(..., min_length=1, description="stable document id (idempotent upsert)")
    title: str = Field(..., min_length=1)
    text: str = Field(..., min_length=1, description="full body — the grounding corpus")
    metadata: dict[str, str] = Field(default_factory=dict)

CatalogIngestReport

Bases: BaseModel

Outcome of a :class:CatalogIngestor.ingest call.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
class CatalogIngestReport(BaseModel):
    """Outcome of a :class:`CatalogIngestor.ingest` call."""

    model_config = _CFG

    slug: str
    ingested: int = Field(description="number of documents written this call")
    embedded: int = Field(default=0, description="documents whose node was embedded")
    total_documents: int = Field(
        default=0, alias="totalDocuments", description="catalog size after this call"
    )
    bm25_only: bool = Field(
        default=False,
        alias="bm25Only",
        description="True when the embedder was absent → lexical-only grounding",
    )
    landing_page_id: str = Field(alias="landingPageId")

ClassNode

Bases: GraphNodeBase

A class / struct definition.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1959
1960
1961
1962
class ClassNode(GraphNodeBase):
    """A class / struct definition."""

    type: Literal["Class"] = "Class"

CodeGraph

Bases: BaseModel

Whole-graph validated bundle of nodes + edges (schema v2).

Assembled + validated ONCE at ingest (build_graph_core) before the store persists the flat lists. The model_validator enforces three invariants a per-node/edge check can't: (1) node-id uniqueness; (2) referential integrity — a non-synthetic edge's endpoints must both resolve to a node (a synthetic cross-file edge carrying target_name may point out-of-repo, and its source may itself be a by-name id); (3) CPG-style per-edge-type endpoint rules. An already-trusted persisted graph can skip re-validation via model_construct.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
2115
2116
2117
2118
2119
2120
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
2143
2144
2145
2146
2147
2148
2149
2150
2151
2152
2153
2154
2155
2156
2157
2158
2159
2160
2161
class CodeGraph(BaseModel):
    """Whole-graph validated bundle of nodes + edges (schema v2).

    Assembled + validated ONCE at ingest (``build_graph_core``) before the store
    persists the flat lists. The ``model_validator`` enforces three invariants a
    per-node/edge check can't: (1) node-id uniqueness; (2) referential integrity
    — a non-synthetic edge's endpoints must both resolve to a node (a synthetic
    cross-file edge carrying ``target_name`` may point out-of-repo, and its
    source may itself be a by-name id); (3) CPG-style per-edge-type endpoint
    rules. An already-trusted persisted graph can skip re-validation via
    ``model_construct``.
    """

    model_config = _CFG

    schema_version: Literal["1"] = "1"
    nodes: list[_GraphNodeAnnotated] = Field(default_factory=list)
    edges: list[GraphEdge] = Field(default_factory=list)

    @model_validator(mode="after")
    def _validate_graph(self) -> CodeGraph:
        """Enforce id-uniqueness + referential integrity + endpoint rules."""
        id_type: dict[str, str] = {}
        for n in self.nodes:
            if n.node_id in id_type:
                raise ValueError(f"duplicate node_id: {n.node_id!r}")
            id_type[n.node_id] = n.type
        for e in self.edges:
            synthetic = e.target_name is not None
            if not synthetic:
                if e.source not in id_type:
                    raise ValueError(
                        f"{e.type} edge source {e.source!r} does not resolve to a node"
                    )
                if e.target not in id_type:
                    raise ValueError(
                        f"{e.type} edge target {e.target!r} does not resolve to a node "
                        "(and carries no target_name)"
                    )
            allowed = _EDGE_SOURCE_KINDS.get(e.type)
            src_type = id_type.get(e.source)
            if allowed is not None and src_type is not None and src_type not in allowed:
                raise ValueError(
                    f"{e.type} edge illegal source kind {src_type!r} "
                    f"(allowed: {sorted(allowed)})"
                )
        return self

CommitScope dataclass

Which generation of a slug's persisted graph a read covers.

The store holds the UNION of every commit ever indexed for a slug, so a reader has to say which generation it means. There are exactly two answers, they are not orderable, and one of them is not expressible as a commit sha — hence a type rather than another optional string:

  • CommitScope.at(sha) — rows stamped exactly sha. at(None) matches rows stamped NULL (a commit-less catalog node, or any unstamped node), which is a real generation, not "no filter".
  • CommitScope.every() — every generation ever indexed.

Why this is not a commit_sha: str | None parameter. count_graph_nodes and supersede_graph_artifacts already take that parameter, and None there means "stamped NULL" — an exact match, which is precisely how they count and preserve the pre-isolation generation. Adding commit_sha: str | None = None to query_graph with "unscoped" semantics would give one parameter name OPPOSITE meanings on two methods of the same class, so a reader who learned it on one would be wrong on the other with nothing to warn them.

Both projections live here rather than in the drivers because the two drivers filter through different mechanisms — the JSON driver tests a loaded row, Mongo narrows a query document — and a rule typed twice is a rule that drifts. filter_fields returns plain field equality, not a Mongo operator, so it stays storage-agnostic.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
@dataclass(frozen=True, slots=True)
class CommitScope:
    """Which generation of a slug's persisted graph a read covers.

    The store holds the UNION of every commit ever indexed for a slug, so a
    reader has to say which generation it means. There are exactly two
    answers, they are not orderable, and one of them is not expressible as a
    commit sha — hence a type rather than another optional string:

    - ``CommitScope.at(sha)`` — rows stamped exactly *sha*. ``at(None)``
      matches rows stamped NULL (a commit-less catalog node, or any unstamped
      node), which is a real generation, not "no filter".
    - ``CommitScope.every()`` — every generation ever indexed.

    **Why this is not a ``commit_sha: str | None`` parameter.**
    ``count_graph_nodes`` and ``supersede_graph_artifacts`` already take that
    parameter, and ``None`` there means "stamped NULL" — an exact match, which
    is precisely how they count and preserve the pre-isolation generation.
    Adding ``commit_sha: str | None = None`` to ``query_graph`` with "unscoped"
    semantics would give one parameter name OPPOSITE meanings on two methods of
    the same class, so a reader who learned it on one would be wrong on the
    other with nothing to warn them.

    Both projections live here rather than in the drivers because the two
    drivers filter through different mechanisms — the JSON driver tests a
    loaded row, Mongo narrows a query document — and a rule typed twice is a
    rule that drifts. ``filter_fields`` returns plain field equality, not a
    Mongo operator, so it stays storage-agnostic.
    """

    sha: str | None = None
    scoped: bool = False

    @classmethod
    def at(cls, sha: str | None) -> CommitScope:
        """Scope to rows stamped exactly *sha* (``None`` matches NULL-stamped)."""
        return cls(sha=sha, scoped=True)

    @classmethod
    def every(cls) -> CommitScope:
        """Every generation ever indexed for the slug (the union)."""
        return cls(sha=None, scoped=False)

    def matches(self, row_commit: str | None) -> bool:
        """Does a row stamped *row_commit* fall in this scope?"""
        return row_commit == self.sha if self.scoped else True

    def filter_fields(self) -> dict[str, str | None]:
        """Field-equality predicate for this scope; empty dict when unscoped."""
        return {"commit_sha": self.sha} if self.scoped else {}

at(sha: str | None) -> CommitScope classmethod

Scope to rows stamped exactly sha (None matches NULL-stamped).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1840
1841
1842
1843
@classmethod
def at(cls, sha: str | None) -> CommitScope:
    """Scope to rows stamped exactly *sha* (``None`` matches NULL-stamped)."""
    return cls(sha=sha, scoped=True)

every() -> CommitScope classmethod

Every generation ever indexed for the slug (the union).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1845
1846
1847
1848
@classmethod
def every(cls) -> CommitScope:
    """Every generation ever indexed for the slug (the union)."""
    return cls(sha=None, scoped=False)

filter_fields() -> dict[str, str | None]

Field-equality predicate for this scope; empty dict when unscoped.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1854
1855
1856
def filter_fields(self) -> dict[str, str | None]:
    """Field-equality predicate for this scope; empty dict when unscoped."""
    return {"commit_sha": self.sha} if self.scoped else {}

matches(row_commit: str | None) -> bool

Does a row stamped row_commit fall in this scope?

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1850
1851
1852
def matches(self, row_commit: str | None) -> bool:
    """Does a row stamped *row_commit* fall in this scope?"""
    return row_commit == self.sha if self.scoped else True

CompleteEvent

Bases: BaseModel

Terminal event: indexing succeeded.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1462
1463
1464
1465
1466
1467
1468
class CompleteEvent(BaseModel):
    """Terminal event: indexing succeeded."""

    model_config = _CFG
    type: Literal["complete"]
    landing_page_id: str = Field(alias="landingPageId")
    page_count: int = Field(alias="pageCount")

DiagramBlock

Bases: BaseModel

Mermaid diagram reference block.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
334
335
336
337
338
339
class DiagramBlock(BaseModel):
    """Mermaid diagram reference block."""

    model_config = _CFG
    kind: Literal["diagram"]
    id: str

Embedding

Bases: BaseModel

Dense embedding vector for a graph node.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
class Embedding(BaseModel):
    """Dense embedding vector for a graph node."""

    model_config = _CFG

    slug: str
    node_id: str
    vector: list[float]
    model: str  # embedding model id
    dim: int
    # Per-job/commit attribution — see ``GraphNodeBase``. An embedding is reaped
    # with the node it vectorises when a completed re-index supersedes its commit.
    commit_sha: str | None = None
    job_id: str | None = None

ErrorEvent

Bases: BaseModel

Terminal event: indexing failed with an error.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1478
1479
1480
1481
1482
1483
class ErrorEvent(BaseModel):
    """Terminal event: indexing failed with an error."""

    model_config = _CFG
    type: Literal["error"]
    error: WikiError

ExternalNode

Bases: GraphNodeBase

VIEW-only convergence node for an unresolved out-of-repo symbol.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1995
1996
1997
1998
class ExternalNode(GraphNodeBase):
    """VIEW-only convergence node for an unresolved out-of-repo symbol."""

    type: Literal["External"] = "External"

FileNode

Bases: GraphNodeBase

A source file — the container every in-file symbol hangs off.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1947
1948
1949
1950
class FileNode(GraphNodeBase):
    """A source file — the container every in-file symbol hangs off."""

    type: Literal["File"] = "File"

FinalizingEvent

Bases: BaseModel

Emitted when all files are scanned and final pages are being written.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1453
1454
1455
1456
1457
1458
1459
class FinalizingEvent(BaseModel):
    """Emitted when all files are scanned and final pages are being written."""

    model_config = _CFG
    type: Literal["finalizing"]
    scanned_count: int = Field(alias="scannedCount")
    total_count: int = Field(alias="totalCount")

FingerprintDecision

Bases: BaseModel

The three-state, honest verdict from comparing two fingerprints.

Mirrors RepoFreshness.check: "no prior fingerprint to compare against" (reason="unknown") must never collapse into "compared and it matched" (reason="match") — the first means a full rebuild is the safe default, the second means reuse is permitted, and rendering them the same would be exactly the false-green RepoFreshness already refuses to produce.

can_reuse is DERIVED from reason, never independently settable — a plain @property, not a computed_field: this model is frozen (a computed verdict, not mutable state), and a derived key inside model_dump would be a second writer of the same fact reason already carries.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
891
892
893
894
895
896
897
898
899
900
901
902
903
904
905
906
907
908
909
910
911
912
913
914
915
916
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
class FingerprintDecision(BaseModel):
    """The three-state, honest verdict from comparing two fingerprints.

    Mirrors ``RepoFreshness.check``: "no prior fingerprint to compare against"
    (``reason="unknown"``) must never collapse into "compared and it matched"
    (``reason="match"``) — the first means a full rebuild is the safe default,
    the second means reuse is permitted, and rendering them the same would be
    exactly the false-green ``RepoFreshness`` already refuses to produce.

    ``can_reuse`` is DERIVED from ``reason``, never independently settable —
    a plain ``@property``, not a ``computed_field``: this model is frozen (a
    computed verdict, not mutable state), and a derived key inside
    ``model_dump`` would be a second writer of the same fact ``reason``
    already carries.
    """

    model_config = ConfigDict(extra="forbid", populate_by_name=True, frozen=True)

    reason: Literal["unknown", "match", "mismatch"]
    mismatches: list[FingerprintMismatch] = Field(default_factory=list)

    @property
    def can_reuse(self) -> bool:
        """True only when a prior fingerprint was compared and every field matched."""
        return self.reason == "match"

    @classmethod
    def compute(
        cls, current: IndexFingerprint, prior: IndexFingerprint | None
    ) -> FingerprintDecision:
        """Compare *current* against *prior* — ``prior=None`` reads ``unknown``.

        Accumulates EVERY mismatching field rather than stopping at the
        first, so a caller (a log line, a console readout) can report every
        reason a rebuild is needed in one pass, per :data:`_FINGERPRINT_FIELDS`.
        """
        if prior is None:
            return cls(reason="unknown")
        mismatches = [
            FingerprintMismatch(
                field=f, expected=getattr(prior, f), actual=getattr(current, f)
            )
            for f in _FINGERPRINT_FIELDS
            if getattr(prior, f) != getattr(current, f)
        ]
        if mismatches:
            return cls(reason="mismatch", mismatches=mismatches)
        return cls(reason="match")

can_reuse: bool property

True only when a prior fingerprint was compared and every field matched.

compute(current: IndexFingerprint, prior: IndexFingerprint | None) -> FingerprintDecision classmethod

Compare current against prior — prior=None reads unknown.

Accumulates EVERY mismatching field rather than stopping at the first, so a caller (a log line, a console readout) can report every reason a rebuild is needed in one pass, per :data:_FINGERPRINT_FIELDS.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
917
918
919
920
921
922
923
924
925
926
927
928
929
930
931
932
933
934
935
936
937
938
@classmethod
def compute(
    cls, current: IndexFingerprint, prior: IndexFingerprint | None
) -> FingerprintDecision:
    """Compare *current* against *prior* — ``prior=None`` reads ``unknown``.

    Accumulates EVERY mismatching field rather than stopping at the
    first, so a caller (a log line, a console readout) can report every
    reason a rebuild is needed in one pass, per :data:`_FINGERPRINT_FIELDS`.
    """
    if prior is None:
        return cls(reason="unknown")
    mismatches = [
        FingerprintMismatch(
            field=f, expected=getattr(prior, f), actual=getattr(current, f)
        )
        for f in _FINGERPRINT_FIELDS
        if getattr(prior, f) != getattr(current, f)
    ]
    if mismatches:
        return cls(reason="mismatch", mismatches=mismatches)
    return cls(reason="match")

FingerprintMismatch

Bases: BaseModel

One field where a prior fingerprint and the current one disagree.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
881
882
883
884
885
886
887
888
class FingerprintMismatch(BaseModel):
    """One field where a prior fingerprint and the current one disagree."""

    model_config = _CFG

    field: FingerprintField
    expected: str | bool | None
    actual: str | bool | None

FolderNode

Bases: GraphNodeBase

VIEW-only directory supernode (hierarchy wire mode).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
2001
2002
2003
2004
class FolderNode(GraphNodeBase):
    """VIEW-only directory supernode (hierarchy wire mode)."""

    type: Literal["Folder"] = "Folder"

Frontmatter

Bases: BaseModel

Parsed frontmatter from a wiki page.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
371
372
373
374
375
376
377
378
379
380
class Frontmatter(BaseModel):
    """Parsed frontmatter from a wiki page."""

    model_config = _CFG
    title: str
    slug: str
    relevant_sources: list[SourceRef] | None = Field(
        default=None, alias="relevantSources"
    )
    sources: list[SourceRef] | None = None

FunctionNode

Bases: GraphNodeBase

A top-level (unbound) function.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1971
1972
1973
1974
class FunctionNode(GraphNodeBase):
    """A top-level (unbound) function."""

    type: Literal["Function"] = "Function"

GraphEdge

Bases: BaseModel

Directed edge in the code graph.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
class GraphEdge(BaseModel):
    """Directed edge in the code graph."""

    model_config = _CFG

    slug: str
    source: str  # node_id
    target: str  # node_id
    type: GraphEdgeType
    # Carry for cross-file IMPORTS/CALLS/EXTENDS whose target is NOT an in-repo
    # node. ``target`` then holds a synthetic external id and ``target_name`` the
    # raw symbol name, so the view can converge every reference to one named
    # ``External`` node (a view concern — the persisted node table stays
    # real-in-repo-symbols only). ``None`` for ordinary in-repo edges.
    target_name: str | None = None
    # Same open subkind + namespaced extension bag as the node axes.
    # Defaults keep a persisted edge lacking them valid.
    subkind: str | None = None
    attributes: dict[str, _JsonScalar] = Field(default_factory=dict)
    # Per-job/commit attribution — see ``GraphNodeBase``. Stamped at write and
    # superseded per commit alongside the nodes an edge connects.
    commit_sha: str | None = None
    job_id: str | None = None

    @field_validator("attributes")
    @classmethod
    def _ns_attributes(cls, v: dict[str, _JsonScalar]) -> dict[str, _JsonScalar]:
        """Enforce namespaced attribute keys (see :func:`_require_namespaced`)."""
        return _require_namespaced(v)

GraphNodeBase

Bases: BaseModel

Shared identity + provenance for every code-graph node kind.

frozen — a node is a value object the extractor emits once and every downstream layer only READS (the view stamps hierarchy onto the wire dict, never the node), so immutability is free and makes nodes hashable/cacheable. Per-kind subclasses narrow type to a Literal; the discriminated :data:GraphNode union dispatches on it. subkind/attributes default, so a persisted node lacking them still validates.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
class GraphNodeBase(BaseModel):
    """Shared identity + provenance for every code-graph node kind.

    ``frozen`` — a node is a value object the extractor emits once and every
    downstream layer only READS (the view stamps hierarchy onto the wire dict,
    never the node), so immutability is free and makes nodes hashable/cacheable.
    Per-kind subclasses narrow ``type`` to a Literal; the discriminated
    :data:`GraphNode` union dispatches on it. ``subkind``/``attributes``
    default, so a persisted node lacking them still validates.
    """

    model_config = _NODE_CFG

    slug: str
    node_id: str
    type: GraphNodeType
    name: str
    file: str
    range: tuple[int, int]
    docstring: str | None = None
    # Open per-kind refinement (Kythe subkind): e.g. ``companion`` for a Kotlin
    # companion object, ``const`` for a constant Property. ``None`` = unrefined.
    subkind: str | None = None
    # Namespaced extension bag (``<lang|tool>.<name>`` keys) — the zero-schema-
    # change seam for language-specific facts. Defaults empty so a persisted
    # node lacking it still validates.
    attributes: dict[str, _JsonScalar] = Field(default_factory=dict)
    # Per-job/commit attribution. ``commit_sha`` is the git commit whose index
    # produced this node; ``job_id`` the job that wrote it. The store stamps both
    # at write from the owning job ctx. ``commit_sha`` is what makes "the graph
    # for THIS commit is built" expressible (the resume skip predicate keys on
    # it) and what a completed re-index supersedes on: every prior-commit node is
    # reaped, so the store stops being the UNION of every commit ever indexed for
    # a slug. ``None`` on a commit-less catalog node and on any unstamped node
    # (the backfill stamps those); the default keeps those valid. NOT part of
    # the wire shape — the graph
    # view assembles its Cytoscape payload field-by-field and never dumps a node.
    commit_sha: str | None = None
    job_id: str | None = None

    @field_validator("attributes")
    @classmethod
    def _ns_attributes(cls, v: dict[str, _JsonScalar]) -> dict[str, _JsonScalar]:
        """Enforce namespaced attribute keys (see :func:`_require_namespaced`)."""
        return _require_namespaced(v)

    @property
    def embedding_text(self) -> str:
        """The text an embedder vectorises this node as.

        Lives ON the node because it is a projection of the node's own fields,
        and because two paths embed the same graph — the full index and the
        scoped incremental refresh. Held as a helper beside one of them, the
        other silently embeds a DIFFERENT string, and the same symbol lands in a
        different vector neighbourhood depending on which path last touched its
        file. The ``file`` segment is dropped when it merely repeats ``name``
        (a File node names itself) so it never counts twice.
        """
        parts = [self.name]
        if self.docstring:
            parts.append(self.docstring)
        if self.file and self.file != self.name:
            parts.append(self.file)
        return " — ".join(parts)

embedding_text: str property

The text an embedder vectorises this node as.

Lives ON the node because it is a projection of the node's own fields, and because two paths embed the same graph — the full index and the scoped incremental refresh. Held as a helper beside one of them, the other silently embeds a DIFFERENT string, and the same symbol lands in a different vector neighbourhood depending on which path last touched its file. The file segment is dropped when it merely repeats name (a File node names itself) so it never counts twice.

GraphResolution

Bases: BaseModel

What exact cross-file symbol resolution achieved for one indexed commit.

A code graph built without exact resolution is not visibly broken — it is fully populated, passes validation and renders — so "were this graph's cross-file edges resolved exactly, or guessed by name?" cannot be answered from the graph itself. It is answerable from this record, which the graph phase writes onto the project it indexed.

Lives on TWO records for the same reason IndexFingerprint does: IndexingJob.resolution is what THIS run achieved, Project.resolution is the snapshot describing the graph currently in the store.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
class GraphResolution(BaseModel):
    """What exact cross-file symbol resolution achieved for one indexed commit.

    A code graph built without exact resolution is not visibly broken — it is
    fully populated, passes validation and renders — so "were this graph's
    cross-file edges resolved exactly, or guessed by name?" cannot be answered
    from the graph itself. It is answerable from this record, which the graph
    phase writes onto the project it indexed.

    Lives on TWO records for the same reason ``IndexFingerprint`` does:
    ``IndexingJob.resolution`` is what THIS run achieved, ``Project.resolution``
    is the snapshot describing the graph currently in the store.
    """

    model_config = _CFG

    # False when the resolver could not run at all (its backend is not
    # installed on this deployment), which is a different repair from one that
    # ran and covered only part of the repository — hence a field of its own
    # rather than an inference from a zero count.
    available: bool
    # Package roots the resolver found in the repository, and how many of them
    # it indexed. Fewer indexed than discovered is a PARTIAL pass: the roots it
    # missed keep their name-matched edges, so the two counts together are what
    # separates "exact everywhere" from "exact in places".
    roots_discovered: int = Field(default=0, ge=0, alias="rootsDiscovered")
    roots_indexed: int = Field(default=0, ge=0, alias="rootsIndexed")
    # Cross-file edges contributed by exact resolution. Zero alongside a
    # complete pass means the repository genuinely has no resolvable
    # cross-file references in the resolver's language.
    resolved_edges: int = Field(default=0, ge=0, alias="resolvedEdges")

    # Both questions below are plain properties, NOT ``computed_field``: this
    # model is ``extra="forbid"`` and round-trips through the store, so a
    # derived key inside ``model_dump`` would make every persisted record fail
    # to re-validate — the reasoning ``FingerprintDecision.can_reuse`` already
    # states for the same trade.

    @property
    def faithful(self) -> bool:
        """True when exact resolution ran and covered every root it discovered.

        This is the one condition under which a name-matched cross-file edge
        may be dropped in favour of a resolved one: a partial pass leaves the
        roots it missed with no exact edges at all, so dropping theirs would
        remove the only edges those files have.
        """
        return (
            self.available
            and self.roots_indexed > 0
            and self.roots_indexed >= self.roots_discovered
        )

    @property
    def degraded(self) -> bool:
        """True when the stored graph's cross-file edges are not exact ones.

        The single question a reader asks of this record: a graph whose
        resolution never ran, covered part of the repository, or produced no
        edges answers every "who calls this" by name matching. Reading it off
        the project record is what makes that visible without counting edges by
        hand.
        """
        return not self.faithful or self.resolved_edges == 0

    def describe(self) -> str:
        """One line naming the outcome — for a job log or an operator readout.

        On the model rather than at the call site because every surface that
        reports a pass wants the same sentence, and the counters only mean
        something together.
        """
        if not self.available:
            return "exact symbol resolution unavailable — cross-file edges are name matches"
        coverage = f"{self.roots_indexed}/{self.roots_discovered} project roots indexed"
        if self.roots_indexed == 0:
            return (
                f"exact symbol resolution ran but indexed nothing ({coverage}) — "
                "cross-file edges are name matches"
            )
        if self.faithful:
            return f"exact symbol resolution: {coverage}, {self.resolved_edges} edges"
        return (
            f"exact symbol resolution covered part of the repository ({coverage}, "
            f"{self.resolved_edges} edges) — the roots it missed keep their "
            "name-matched edges"
        )

degraded: bool property

True when the stored graph's cross-file edges are not exact ones.

The single question a reader asks of this record: a graph whose resolution never ran, covered part of the repository, or produced no edges answers every "who calls this" by name matching. Reading it off the project record is what makes that visible without counting edges by hand.

faithful: bool property

True when exact resolution ran and covered every root it discovered.

This is the one condition under which a name-matched cross-file edge may be dropped in favour of a resolved one: a partial pass leaves the roots it missed with no exact edges at all, so dropping theirs would remove the only edges those files have.

describe() -> str

One line naming the outcome — for a job log or an operator readout.

On the model rather than at the call site because every surface that reports a pass wants the same sentence, and the counters only mean something together.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
def describe(self) -> str:
    """One line naming the outcome — for a job log or an operator readout.

    On the model rather than at the call site because every surface that
    reports a pass wants the same sentence, and the counters only mean
    something together.
    """
    if not self.available:
        return "exact symbol resolution unavailable — cross-file edges are name matches"
    coverage = f"{self.roots_indexed}/{self.roots_discovered} project roots indexed"
    if self.roots_indexed == 0:
        return (
            f"exact symbol resolution ran but indexed nothing ({coverage}) — "
            "cross-file edges are name matches"
        )
    if self.faithful:
        return f"exact symbol resolution: {coverage}, {self.resolved_edges} edges"
    return (
        f"exact symbol resolution covered part of the repository ({coverage}, "
        f"{self.resolved_edges} edges) — the roots it missed keep their "
        "name-matched edges"
    )

H2Block

Bases: BaseModel

Level-2 heading block.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
275
276
277
278
279
280
281
class H2Block(BaseModel):
    """Level-2 heading block."""

    model_config = _CFG
    kind: Literal["h2"]
    id: str | None = None
    text: str

H3Block

Bases: BaseModel

Level-3 heading block.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
284
285
286
287
288
289
290
class H3Block(BaseModel):
    """Level-3 heading block."""

    model_config = _CFG
    kind: Literal["h3"]
    id: str | None = None
    text: str

HeartbeatEvent

Bases: BaseModel

Keep-alive event; consumers must ignore it.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1486
1487
1488
1489
1490
class HeartbeatEvent(BaseModel):
    """Keep-alive event; consumers must ignore it."""

    model_config = _CFG
    type: Literal["heartbeat"]

HrBlock

Bases: BaseModel

Horizontal-rule block.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
293
294
295
296
297
class HrBlock(BaseModel):
    """Horizontal-rule block."""

    model_config = _CFG
    kind: Literal["hr"]

IndexFingerprint

Bases: BaseModel

The non-content inputs that can invalidate a hash-identical index.

Stamped by build_graph_core at the moment it actually builds the graph — never re-derived later, and never re-probed live at finalize (a live probe would describe "now", not "what built the artifacts actually in the store" — see the comment at the stamp site for the resume-skip case this avoids). Lives on TWO records for the same reason commit_sha already does: IndexingJob.fingerprint is what THIS run used, Project.fingerprint is a copy taken at finalize — the snapshot of what produced the index currently in the store.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
836
837
838
839
840
841
842
843
844
845
846
847
848
849
850
851
852
853
854
855
856
857
858
859
860
861
862
863
864
class IndexFingerprint(BaseModel):
    """The non-content inputs that can invalidate a hash-identical index.

    Stamped by ``build_graph_core`` at the moment it actually builds the
    graph — never re-derived later, and never re-probed live at finalize (a
    live probe would describe "now", not "what built the artifacts actually
    in the store" — see the comment at the stamp site for the resume-skip
    case this avoids). Lives on TWO records for the same reason ``commit_sha``
    already does: ``IndexingJob.fingerprint`` is what THIS run used,
    ``Project.fingerprint`` is a copy taken at finalize — the snapshot of
    what produced the index currently in the store.
    """

    model_config = _CFG

    # ``None`` means no embedding was made this run — embeddings disabled
    # (``wiki.embedding.enabled=false``), or the embed pass failed before any
    # vector was actually persisted. Deliberately NOT a config fallback: a
    # fingerprint records what happened, not what would have run. An index
    # with zero vectors is genuinely stale against one that has them, and a
    # captured value must be able to say so rather than paper over it with
    # "what config says right now".
    embedding_model: str | None = Field(default=None, alias="embeddingModel")
    graph_schema_version: str = Field(alias="graphSchemaVersion")
    # ``None`` when the installed ``tree-sitter-language-pack`` distribution
    # metadata isn't readable (the ``treesitter`` extra absent) — an honest
    # absence, matched against itself below, never coerced into a mismatch.
    grammar_pack_version: str | None = Field(default=None, alias="grammarPackVersion")
    resolver_available: bool = Field(alias="resolverAvailable")

IndexingEventUnion

Bases: RootModel[_IndexingEventAnnotated]

Discriminated union of all indexing SSE events.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1564
1565
class IndexingEventUnion(RootModel[_IndexingEventAnnotated]):
    """Discriminated union of all indexing SSE events."""

IndexingJob

Bases: BaseModel

Snapshot of an in-progress or finished indexing job.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
class IndexingJob(BaseModel):
    """Snapshot of an in-progress or finished indexing job."""

    model_config = _CFG

    job_id: str = Field(alias="jobId")
    slug: str
    status: IndexingStatus
    scanned_count: int = Field(alias="scannedCount")
    total_count: int = Field(alias="totalCount")
    current_file: str | None = Field(alias="currentFile")
    landing_page_id: str | None = Field(default=None, alias="landingPageId")
    # Platform of record (gitea, github, …). Hydrated from the wizard
    # submission; lets the FE compose canonical URLs without a round-trip.
    platform: PlatformId | None = None
    # DNS host the repo lives on — first-class for enterprise/self-hosted.
    host: str | None = None
    # LLM model authoring this wiki — surfaced for user transparency.
    model: str | None = None
    # ── Phase-weighted progress ────────────────────────────────────────
    # The coarse 6-state ``status`` is the lifecycle bucket; ``phase`` is
    # the fine-grained progress state. Both the landing card and the
    # indexing page read these to render a single honest progress bar.
    phase: IndexingPhase | None = None
    total_pages: int | None = Field(default=None, alias="totalPages")
    pages_submitted: int = Field(default=0, alias="pagesSubmitted")
    # ISO timestamp at which the current ``phase`` started. Used by the
    # FE to extrapolate an ETA inside the active phase.
    phase_started_at: str | None = Field(default=None, alias="phaseStartedAt")
    # Generic per-phase progress — ONE mechanism for every phase beyond
    # scan/pages (today ``graph`` and ``enrich``; a future phase needs no new
    # field). Per-phase field pairs were the alternative and they are how
    # ``scanned_count``/``current_file`` became untrustworthy: N pairs for N
    # phases, each left frozen at its phase's last value and still readable as
    # if it described the current one.
    #
    # THE INVARIANT that makes these trustworthy: ``emit_phase`` clears all
    # three on every transition, so a non-null ``phase_progress_current``
    # always belongs to the phase named in ``phase``. A ``None`` total means "a
    # running count with no knowable total" — a real status line but not a
    # fraction. ``unit`` is the plural noun the reader renders ("files",
    # "nodes", "entities"); absent, a consumer falls back to a generic label.
    # The durable per-step model. It is bounded by DECLARED steps, never by
    # repository units, so every snapshot read stays ``O(1)`` in repository
    # size. It supersedes the legacy triple below, retained only so an in-flight
    # job and older clients keep working during the migration.
    progress: ProgressLedger | None = Field(default=None)
    phase_progress_current: int | None = Field(default=None, alias="phaseProgressCurrent")
    phase_progress_total: int | None = Field(default=None, alias="phaseProgressTotal")
    phase_progress_unit: str | None = Field(default=None, alias="phaseProgressUnit")
    # ISO timestamp of the most recent write by WHATEVER phase is running — the
    # one honest "this job is still moving" signal. ``scanned_count`` and
    # ``current_file`` cannot answer that question: they are the scan phase's
    # private bookkeeping and nothing touches them again until finalize, so they
    # sit frozen — and freshly plausible — through graph, enrich, plan and
    # pages. Read via :meth:`seconds_since_progress`, which is what separates a
    # long phase from a dead one.
    last_progress_at: str | None = Field(default=None, alias="lastProgressAt")
    # Git snapshot resolved at clone time. ``finalize`` reads these off
    # the snapshot when persisting the Project record — no extra args
    # threaded through the tool chain.
    branch: str | None = None
    commit_sha: str | None = Field(default=None, alias="commitSha")
    # forward ref to WikiError — resolved by IndexingJob.model_rebuild() below
    error: WikiError | None = None
    # Non-content invalidators captured when THIS run last built the graph
    # (``build_graph_core``, at the end of the ``graph`` phase) — never
    # re-probed at finalize. ``update_job`` is an unlocked read-modify-write
    # on both backends (a known, separately-tracked defect) and a concurrent
    # writer — a Cancel from the request thread is the documented case — can
    # lose this stamp exactly as it can lose any other field here. That
    # failure mode is SAFE by construction: an absent fingerprint reads as
    # ``FingerprintDecision(reason="unknown")``, which forces a full rebuild
    # rather than silently permitting reuse of artifacts nothing actually
    # fingerprinted. Do not "optimise" a missing value into an assumed match.
    fingerprint: IndexFingerprint | None = None
    # What exact cross-file symbol resolution achieved when THIS run built the
    # graph (``build_graph_core``, at the end of the ``graph`` phase). Written
    # beside the fingerprint, and lost to the same known unlocked
    # read-modify-write in ``update_job`` under a concurrent writer — safe by
    # construction for the same reason: an absent record reads as unknown,
    # never as a pass that resolved everything.
    resolution: GraphResolution | None = None
    # Which refresh path this job took, and why — stamped ONCE at creation by
    # ``WikiIndexingJob.refresh`` and never rewritten. ``None`` means the job is
    # a FIRST index rather than a refresh, which is a third state and not the
    # same as "full": nothing was reused because there was nothing to reuse.
    # Two readers depend on it beyond display — the resume branch, which must
    # re-drive a stranded scoped job the way it ran the first time, and the
    # console, which renders the full-rebuild reason beside the scope preview.
    refresh_decision: RefreshDecision | None = Field(
        default=None, alias="refreshDecision"
    )
    # The committed scope of a scoped refresh, written at the end of its delta
    # pass through the SAME ``emit_*`` seam that writes the event — one write,
    # two transports, so the live stream and the snapshot cannot disagree.
    # ``None`` on every full rebuild: a full index has no delta to preview, and
    # rendering zeros there would claim it examined a scope and found nothing.
    scope_preview: ScopePreview | None = Field(default=None, alias="scopePreview")

    # ── Lifecycle questions ────────────────────────────────────────────
    # Plain properties, NOT ``computed_field``: this model is ``extra="forbid"``
    # and round-trips through the store, so a derived value in ``model_dump``
    # would make every persisted snapshot fail to re-validate. The wire gets
    # the derived flag stamped at its one serialisation seam instead.

    @property
    def is_active(self) -> bool:
        """True while the job is still working — the "Indexing now" question.

        A restart-stranded (``interrupted``) job counts as active: it is
        awaiting recovery, not finished, and hiding it would leave a repository
        looking un-indexed while its job is still queued for a re-drive.
        """
        return self.status in _ACTIVE_STATUSES

    @property
    def is_recoverable(self) -> bool:
        """True when restart recovery should re-drive this job on boot.

        Deliberately excludes ``failed`` even though a failed job may still hold
        reusable checkpoints: an automatic re-drive of a job that already
        exhausted its retry budget is how a dying index loops the API. A human
        can still resume it — see :attr:`is_resumable`.
        """
        return self.status in _RECOVERABLE_STATUSES

    @property
    def is_terminal(self) -> bool:
        """True when the job settled on its own terms (finished, or stopped).

        The question a session-end reconciler asks: *may I leave this alone?*
        ``failed`` answers False on purpose — a session ending cleanly on a
        failed job is a mismatch between what the run believed and what it
        built, and that mismatch is worth reporting.
        """
        return self.status in _SETTLED_STATUSES

    @property
    def is_resumable(self) -> bool:
        """True when a checkpoint resume is worth attempting.

        Broader than :attr:`is_recoverable`: a ``failed`` job is not re-driven
        automatically but a user may still ask to resume it, and ``ResumePlan``
        decides what of it can actually be reused.
        """
        return self.status not in _NON_RESUMABLE_STATUSES

    def regresses_to(self, phase: str) -> bool:
        """True when *phase* sits BEFORE the one this job already reached.

        A resume is told to re-clone and re-scan so the source is back on disk
        before pages are written, so those tools legitimately run again and
        re-stamp phases the job passed long ago. Both progress surfaces read
        ``phase`` off this snapshot, so without this question being asked the
        bar walks backwards mid-resume — a job that had already built its graph
        reported ``scan`` again, with a ``phase_started_at`` later than the
        graph build's.

        Unknown phase names are never a regression: an unrecognised value is a
        vocabulary the caller knows about and this model does not, and silently
        swallowing its transition would hide real progress.
        """
        if self.phase is None or phase == self.phase:
            return False
        try:
            return PHASE_SEQUENCE.index(phase) < PHASE_SEQUENCE.index(self.phase)
        except ValueError:
            return False

    @staticmethod
    def format_stamp(now: datetime) -> str:
        """Render *now* in the one spelling this model's timestamp fields use.

        The write half of :meth:`seconds_since_progress`. Both live here so the
        format is stated once: a writer that spells it differently produces a
        stamp its own reader cannot parse, and the reader's failure mode is a
        silent ``None`` rather than an exception.
        """
        return now.strftime(_TS_FORMAT)

    def seconds_since_progress(self, now: datetime) -> float | None:
        """Seconds between *now* and the last progress write, or ``None``.

        ``None`` means "no usable baseline" — either nothing has reported
        progress yet or the stored stamp is unreadable. Both answers say the
        same thing to a caller (there is nothing to compare against), and
        neither may be reported as "0 seconds ago", which would read as a job
        that just moved.

        The clock arrives as an ARGUMENT: this model is persisted and wired, and
        a model that reads a clock cannot be tested without patching one.
        """
        if not self.last_progress_at:
            return None
        try:
            stamp = datetime.strptime(self.last_progress_at, _TS_FORMAT).replace(
                tzinfo=timezone.utc
            )
        except ValueError:
            return None
        return (now - stamp).total_seconds()

is_active: bool property

True while the job is still working — the "Indexing now" question.

A restart-stranded (interrupted) job counts as active: it is awaiting recovery, not finished, and hiding it would leave a repository looking un-indexed while its job is still queued for a re-drive.

is_recoverable: bool property

True when restart recovery should re-drive this job on boot.

Deliberately excludes failed even though a failed job may still hold reusable checkpoints: an automatic re-drive of a job that already exhausted its retry budget is how a dying index loops the API. A human can still resume it — see :attr:is_resumable.

is_resumable: bool property

True when a checkpoint resume is worth attempting.

Broader than :attr:is_recoverable: a failed job is not re-driven automatically but a user may still ask to resume it, and ResumePlan decides what of it can actually be reused.

is_terminal: bool property

True when the job settled on its own terms (finished, or stopped).

The question a session-end reconciler asks: may I leave this alone? failed answers False on purpose — a session ending cleanly on a failed job is a mismatch between what the run believed and what it built, and that mismatch is worth reporting.

format_stamp(now: datetime) -> str staticmethod

Render now in the one spelling this model's timestamp fields use.

The write half of :meth:seconds_since_progress. Both live here so the format is stated once: a writer that spells it differently produces a stamp its own reader cannot parse, and the reader's failure mode is a silent None rather than an exception.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
@staticmethod
def format_stamp(now: datetime) -> str:
    """Render *now* in the one spelling this model's timestamp fields use.

    The write half of :meth:`seconds_since_progress`. Both live here so the
    format is stated once: a writer that spells it differently produces a
    stamp its own reader cannot parse, and the reader's failure mode is a
    silent ``None`` rather than an exception.
    """
    return now.strftime(_TS_FORMAT)

regresses_to(phase: str) -> bool

True when phase sits BEFORE the one this job already reached.

A resume is told to re-clone and re-scan so the source is back on disk before pages are written, so those tools legitimately run again and re-stamp phases the job passed long ago. Both progress surfaces read phase off this snapshot, so without this question being asked the bar walks backwards mid-resume — a job that had already built its graph reported scan again, with a phase_started_at later than the graph build's.

Unknown phase names are never a regression: an unrecognised value is a vocabulary the caller knows about and this model does not, and silently swallowing its transition would hide real progress.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
def regresses_to(self, phase: str) -> bool:
    """True when *phase* sits BEFORE the one this job already reached.

    A resume is told to re-clone and re-scan so the source is back on disk
    before pages are written, so those tools legitimately run again and
    re-stamp phases the job passed long ago. Both progress surfaces read
    ``phase`` off this snapshot, so without this question being asked the
    bar walks backwards mid-resume — a job that had already built its graph
    reported ``scan`` again, with a ``phase_started_at`` later than the
    graph build's.

    Unknown phase names are never a regression: an unrecognised value is a
    vocabulary the caller knows about and this model does not, and silently
    swallowing its transition would hide real progress.
    """
    if self.phase is None or phase == self.phase:
        return False
    try:
        return PHASE_SEQUENCE.index(phase) < PHASE_SEQUENCE.index(self.phase)
    except ValueError:
        return False

seconds_since_progress(now: datetime) -> float | None

Seconds between now and the last progress write, or None.

None means "no usable baseline" — either nothing has reported progress yet or the stored stamp is unreadable. Both answers say the same thing to a caller (there is nothing to compare against), and neither may be reported as "0 seconds ago", which would read as a job that just moved.

The clock arrives as an ARGUMENT: this model is persisted and wired, and a model that reads a clock cannot be tested without patching one.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
def seconds_since_progress(self, now: datetime) -> float | None:
    """Seconds between *now* and the last progress write, or ``None``.

    ``None`` means "no usable baseline" — either nothing has reported
    progress yet or the stored stamp is unreadable. Both answers say the
    same thing to a caller (there is nothing to compare against), and
    neither may be reported as "0 seconds ago", which would read as a job
    that just moved.

    The clock arrives as an ARGUMENT: this model is persisted and wired, and
    a model that reads a clock cannot be tested without patching one.
    """
    if not self.last_progress_at:
        return None
    try:
        stamp = datetime.strptime(self.last_progress_at, _TS_FORMAT).replace(
            tzinfo=timezone.utc
        )
    except ValueError:
        return None
    return (now - stamp).total_seconds()

InlineNode

Bases: RootModel[str | list['InlineNode'] | dict]

Recursive inline rich-text node.

Valid root values:

  • str — plain text
  • list[InlineNode] — sequence of inline nodes
  • {"code": str} — inline code span
  • {"link": str, "text": str} — hyperlink
  • {"kind": "src", "path": str, "lines"?: str} — source reference
Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
248
249
250
251
252
253
254
255
256
257
258
class InlineNode(RootModel[str | list["InlineNode"] | dict]):
    """Recursive inline rich-text node.

    Valid root values:

    - ``str`` — plain text
    - ``list[InlineNode]`` — sequence of inline nodes
    - ``{"code": str}`` — inline code span
    - ``{"link": str, "text": str}`` — hyperlink
    - ``{"kind": "src", "path": str, "lines"?: str}`` — source reference
    """

InterfaceNode

Bases: GraphNodeBase

An interface / trait / protocol definition.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1965
1966
1967
1968
class InterfaceNode(GraphNodeBase):
    """An interface / trait / protocol definition."""

    type: Literal["Interface"] = "Interface"

Language

Bases: BaseModel

Language option shown in the wizard.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
206
207
208
209
210
211
212
213
class Language(BaseModel):
    """Language option shown in the wizard."""

    model_config = _CFG

    id: str
    label: str
    subtle: str | None = None

LogEvent

Bases: BaseModel

Free-form milestone line shown in the indexing timeline.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1538
1539
1540
1541
1542
1543
1544
class LogEvent(BaseModel):
    """Free-form milestone line shown in the indexing timeline."""

    model_config = _CFG
    type: Literal["log"]
    level: LogLevel
    text: str

MetaEvent

Bases: BaseModel

First QA event of a turn, carrying answer ID and chosen model.

Emitted once per turn — at the start of :meth:WikiQaSession.start AND :meth:WikiQaSession.follow_up — so it also marks where each turn's events begin in the append-only log (see QaFinalizer.current_turn_events).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
class MetaEvent(BaseModel):
    """First QA event of a turn, carrying answer ID and chosen model.

    Emitted once per turn — at the start of :meth:`WikiQaSession.start` AND
    :meth:`WikiQaSession.follow_up` — so it also marks where each turn's
    events begin in the append-only log (see ``QaFinalizer.current_turn_events``).
    """

    model_config = _CFG
    type: Literal["meta"]
    answer_id: str = Field(alias="answerId")
    model: str
    from_page_id: str = Field(alias="fromPageId")
    # The backing Mewbo session id — exposed so continuation is addressable /
    # traceable. Defaulted so an older persisted event replays
    # unchanged.
    session_id: str = Field(default="", alias="sessionId")

MethodNode

Bases: GraphNodeBase

A method bound to a class/struct/object.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1977
1978
1979
1980
class MethodNode(GraphNodeBase):
    """A method bound to a class/struct/object."""

    type: Literal["Method"] = "Method"

ModuleNode

Bases: GraphNodeBase

An imported module target (synthetic cross-file IMPORTS endpoint).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1953
1954
1955
1956
class ModuleNode(GraphNodeBase):
    """An imported module target (synthetic cross-file IMPORTS endpoint)."""

    type: Literal["Module"] = "Module"

NavEntry

Bases: BaseModel

Sidebar navigation entry.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
219
220
221
222
223
224
225
226
227
class NavEntry(BaseModel):
    """Sidebar navigation entry."""

    model_config = _CFG

    id: str
    label: str
    lvl: Literal[1, 2, 3]
    parent: str | None = None

ObjectNode

Bases: GraphNodeBase

A singleton object (Kotlin object / companion object; Scala later).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1983
1984
1985
1986
class ObjectNode(GraphNodeBase):
    """A singleton object (Kotlin ``object`` / ``companion object``; Scala later)."""

    type: Literal["Object"] = "Object"

PBlock

Bases: BaseModel

Paragraph block.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
267
268
269
270
271
272
class PBlock(BaseModel):
    """Paragraph block."""

    model_config = _CFG
    kind: Literal["p"]
    text: InlineNode

PageCommittedEvent

Bases: BaseModel

One page just landed; index is 0-based.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1525
1526
1527
1528
1529
1530
1531
1532
class PageCommittedEvent(BaseModel):
    """One page just landed; ``index`` is 0-based."""

    model_config = _CFG
    type: Literal["page_committed"]
    page_id: str = Field(alias="pageId")
    index: int
    total_pages: int = Field(alias="totalPages")

PagePlan

Bases: BaseModel

Planned wiki page — used by the indexing pipeline before writing.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
class PagePlan(BaseModel):
    """Planned wiki page — used by the indexing pipeline before writing."""

    model_config = _CFG

    id: str
    title: str
    description: str = ""
    importance: Literal["high", "medium", "low"] = "medium"
    relevant_files: list[str] = Field(default_factory=list, alias="relevantFiles")
    related_pages: list[str] = Field(default_factory=list, alias="relatedPages")
    parent: str | None = None

PhaseEvent

Bases: BaseModel

Coarse-phase transition; drives the phase-weighted progress bar.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1509
1510
1511
1512
1513
1514
class PhaseEvent(BaseModel):
    """Coarse-phase transition; drives the phase-weighted progress bar."""

    model_config = _CFG
    type: Literal["phase"]
    name: IndexingPhase

PlanCommittedEvent

Bases: BaseModel

Plan has landed — drives the denominator of the page-write bar.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1517
1518
1519
1520
1521
1522
class PlanCommittedEvent(BaseModel):
    """Plan has landed — drives the denominator of the page-write bar."""

    model_config = _CFG
    type: Literal["plan_committed"]
    total_pages: int = Field(alias="totalPages")

Platform

Bases: BaseModel

Git-hosting platform descriptor (used in wizard).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
class Platform(BaseModel):
    """Git-hosting platform descriptor (used in wizard)."""

    model_config = _CFG

    id: PlatformId
    name: str
    mono: str
    color: str
    short: str
    hosts: list[str]
    token_label: str = Field(alias="tokenLabel")
    token_scope: str = Field(alias="tokenScope")
    token_url: str | None = Field(alias="tokenUrl")
    token_steps: list[str] = Field(alias="tokenSteps")

Project

Bases: BaseModel

Landing-card model for a wiki project.

Slug is fully qualified — host/owner/repo — so the identity is unambiguous across self-hosted and enterprise instances. A two-segment slug (owner/repo) also reads, with host then None.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
class Project(BaseModel):
    """Landing-card model for a wiki project.

    Slug is fully qualified — ``host/owner/repo`` — so the identity is
    unambiguous across self-hosted and enterprise instances. A two-segment
    slug (``owner/repo``) also reads, with ``host`` then ``None``.
    """

    model_config = _CFG

    slug: str
    source: PlatformId
    lang: str
    indexed_at: str = Field(alias="indexedAt")
    pages: int
    primary: bool | None = None
    desc: str
    landing_page_id: str | None = Field(default=None, alias="landingPageId")
    repo_url: str | None = Field(default=None, alias="repoUrl")
    # DNS host the repo lives on (github.com, a self-hosted git.example.com, …).
    # First-class so enterprise instances need no fallback heuristics.
    host: str | None = None
    # Git snapshot the wiki was generated from. Populated by ``finalize``
    # from the IndexingJob; absent when no job stamped one (the FE atomic
    # class hides absent values).
    branch: str | None = None
    commit_sha: str | None = Field(default=None, alias="commitSha")
    commit_short: str | None = Field(default=None, alias="commitShort")
    # True when the cloned repo carried a ``.mewbo/wiki.json`` or
    # ``.devin/wiki.json`` grounder file at finalize time. Sole driver of
    # the "Maintainer Edited" badge — absent means un-edited.
    maintainer_edited: bool = Field(default=False, alias="maintainerEdited")
    # True when the project was indexed in graph-only (developer) mode: the AST
    # code graph was built with NO documentation pages and NO LLM. Stamped at
    # finalize by ``GraphOnlyIndexer``; drives the console's "No documentation
    # available" empty state and makes the doc-content read seam raise
    # ``DocumentationUnavailableError``. Absent means documented.
    graph_only: bool = Field(default=False, alias="graphOnly")
    # Copied from ``IndexingJob.fingerprint`` at finalize — see
    # ``IndexFingerprint`` (defined below; forward ref resolved by
    # ``Project.model_rebuild()`` right after it). ``None`` when no job
    # recorded one — read downstream as "cannot compare, full rebuild",
    # never as a silent match.
    fingerprint: IndexFingerprint | None = None
    # What exact cross-file symbol resolution achieved for the indexed commit
    # — see ``GraphResolution`` (defined below; forward ref resolved by
    # ``Project.model_rebuild()``). Written by the graph phase. ``None`` means
    # the question was never recorded for this project, which reads as unknown
    # and never as a healthy pass.
    resolution: GraphResolution | None = None
    # The prior completed index's observed step costs. One row per declared step,
    # never per repository unit, so loading a project stays ``O(one record)``.
    step_measurements: dict[str, StepMeasurement] = Field(
        default_factory=dict, alias="stepMeasurements"
    )

    def measured_steps(
        self,
        records: list[StepRecord],
        *,
        declared_keys: set[str],
        now: datetime,
    ) -> dict[str, StepMeasurement]:
        """Blend this run's completed records into the calibrated step costs.

        Cost: ``O(declared steps)``. A 75/25 rolling blend retains most of the
        prior reading while admitting a repository's current shape; one run is
        noisy, but a repository can also change size between indexes. Only
        declared, completed records participate, so resume-skipped steps retain
        the previous reading instead of becoming falsely free.
        """
        measurements = {
            key: value
            for key, value in self.step_measurements.items()
            if key in declared_keys
        }
        for record in records:
            if record.key not in declared_keys:
                continue
            observed = StepMeasurement.from_record(record, now)
            if observed is None:
                continue
            previous = measurements.get(record.key)
            measurements[record.key] = (
                observed if previous is None else previous.blended_with(observed)
            )
        return measurements

measured_steps(records: list[StepRecord], *, declared_keys: set[str], now: datetime) -> dict[str, StepMeasurement]

Blend this run's completed records into the calibrated step costs.

Cost: O(declared steps). A 75/25 rolling blend retains most of the prior reading while admitting a repository's current shape; one run is noisy, but a repository can also change size between indexes. Only declared, completed records participate, so resume-skipped steps retain the previous reading instead of becoming falsely free.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
def measured_steps(
    self,
    records: list[StepRecord],
    *,
    declared_keys: set[str],
    now: datetime,
) -> dict[str, StepMeasurement]:
    """Blend this run's completed records into the calibrated step costs.

    Cost: ``O(declared steps)``. A 75/25 rolling blend retains most of the
    prior reading while admitting a repository's current shape; one run is
    noisy, but a repository can also change size between indexes. Only
    declared, completed records participate, so resume-skipped steps retain
    the previous reading instead of becoming falsely free.
    """
    measurements = {
        key: value
        for key, value in self.step_measurements.items()
        if key in declared_keys
    }
    for record in records:
        if record.key not in declared_keys:
            continue
        observed = StepMeasurement.from_record(record, now)
        if observed is None:
            continue
        previous = measurements.get(record.key)
        measurements[record.key] = (
            observed if previous is None else previous.blended_with(observed)
        )
    return measurements

ProjectSettings

Bases: BaseModel

The EDITABLE settings of an INDEXED wiki project, keyed by slug.

Why this exists. :class:Project is a DISPLAY snapshot — wiki_finalize / GraphOnlyIndexer rebuild it WHOLESALE on every successful (re)index, so any field written directly onto it is silently wiped by the next reindex. The settings a project is actually re-indexed WITH have always been the :class:WizardSubmission — but that was persisted as a JOB-keyed sidecar, which gave an editor no stable write target (and made "latest submission" a scan over jobs).

This record is that target: ONE per slug, holding the submission contract minus the never-persisted token, plus a desc display override. WikiIndexingJob.refresh consults it FIRST, falling back to the per-job scan when a project has no record — which is what makes an edit actually take effect on the next index.

Two fields are deliberately absent. slug is the store key for pages, jobs, credentials and freshness, so it is immutable — there is no rename primitive. token never lands here: credentials resolve through the ONE registry (mewbo_graph.wiki.credentials), and a secret in this record would be a third source of truth.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
class ProjectSettings(BaseModel):
    """The EDITABLE settings of an INDEXED wiki project, keyed by slug.

    Why this exists. :class:`Project` is a DISPLAY snapshot —
    ``wiki_finalize`` / ``GraphOnlyIndexer`` rebuild it WHOLESALE on every
    successful (re)index, so any field written directly onto it is silently wiped
    by the next reindex. The settings a project is actually re-indexed WITH have
    always been the :class:`WizardSubmission` — but that was persisted as a
    JOB-keyed sidecar, which gave an editor no stable write target (and made
    "latest submission" a scan over jobs).

    This record is that target: ONE per slug, holding the submission contract
    minus the never-persisted ``token``, plus a ``desc`` display override.
    ``WikiIndexingJob.refresh`` consults it FIRST, falling back to the per-job
    scan when a project has no record — which is what makes an edit actually
    take effect on the next index.

    Two fields are deliberately absent. ``slug`` is the store key for pages, jobs,
    credentials and freshness, so it is immutable — there is no rename primitive.
    ``token`` never lands here: credentials resolve through the ONE registry
    (``mewbo_graph.wiki.credentials``), and a secret in this record would be a
    third source of truth.
    """

    model_config = _CFG

    slug: str
    repo_url: str | None = Field(default=None, alias="repoUrl")
    platform: PlatformId
    model: str
    depth: DepthMode
    language: str
    filter_mode: FilterMode = Field(alias="filterMode")
    dirs: list[str] = Field(default_factory=list)
    files: list[str] = Field(default_factory=list)
    graph_only: bool = Field(default=False, alias="graphOnly")
    ref: str | None = None
    # Ordered model ladder the next index falls back through; ``None`` = inherit
    # the configured policy. It MUST round-trip through
    # ``from_submission``/``to_submission``: ``WikiIndexingJob.refresh`` rebuilds
    # its submission from THIS record, so a ladder the record cannot carry is
    # silently dropped from every index after the first — including for a project
    # that was first indexed with one.
    fallback_models: list[str] | None = Field(default=None, alias="fallbackModels")
    # Embedding model the next index builds this project's vectors with, and that
    # every read of them must embed its query with. ``None`` = inherit the
    # deployment default. Same round-trip obligation as the ladder above; see
    # ``WizardSubmission.embedding_model`` for why it is per-project.
    embedding_model: str | None = Field(default=None, alias="embeddingModel")
    # Operator guidance appended to the indexer playbook on the next index, and
    # the external MCP servers attached to it. Both carry the SAME round-trip
    # obligation the ladder above spells out — a value this record cannot carry
    # is dropped from every index after the first.
    custom_instructions: str | None = Field(default=None, alias="customInstructions")
    mcp_servers: dict[str, dict] | None = Field(default=None, alias="mcpServers")
    # User-set description override. ``None`` = no override, so finalize's
    # platform-API fetch wins (today's behaviour, unchanged). A non-empty value
    # SURVIVES a reindex — that is the read-preserve contract implemented once in
    # ``plugins.wiki.finalize._resolve_project_desc`` and shared by both indexers.
    desc: str | None = None
    updated_at: str | None = Field(default=None, alias="updatedAt")

    @field_validator("custom_instructions")
    @classmethod
    def _check_instructions(cls, v: str | None) -> str | None:
        """Delegate to the submission's rule — this record is also PATCH-written.

        It is a validation boundary in its own right, not merely a projection of
        a submission that was already checked: ``ProjectSettingsPatch`` writes
        here directly, and the store re-reads here on every refresh.
        """
        return WizardSubmission.check_custom_instructions(v)

    @field_validator("mcp_servers")
    @classmethod
    def _check_servers(cls, v: dict[str, dict] | None) -> dict[str, dict] | None:
        """Delegate to the submission's rule — see :meth:`_check_instructions`."""
        return WizardSubmission.check_mcp_servers(v)

    @classmethod
    def from_submission(
        cls, sub: WizardSubmission, *, desc: str | None = None
    ) -> ProjectSettings:
        """Project a :class:`WizardSubmission` onto the durable settings record.

        ``sub.token`` is dropped (the credential registry owns it). *desc* carries
        an existing override forward, so re-seeding this record from a submission
        — which every ``WikiIndexingJob.start`` does, including the one a refresh
        drives — cannot clobber a user's edited description.
        """
        return cls(
            slug=sub.slug,
            repoUrl=sub.repo_url,
            platform=sub.platform,
            model=sub.model,
            depth=sub.depth,
            language=sub.language,
            filterMode=sub.filter_mode,
            dirs=list(sub.dirs),
            files=list(sub.files),
            graphOnly=sub.graph_only,
            ref=sub.ref,
            fallbackModels=(
                list(sub.fallback_models) if sub.fallback_models is not None else None
            ),
            embeddingModel=sub.embedding_model,
            customInstructions=sub.custom_instructions,
            mcpServers=(
                dict(sub.mcp_servers) if sub.mcp_servers is not None else None
            ),
            desc=desc,
        )

    def to_submission(self) -> WizardSubmission:
        """Rebuild the :class:`WizardSubmission` a re-index replays.

        Always token-less: the clone tool's ``resolve_chain`` reads the durable
        credential itself at clone time, so a refresh never needs to carry one.
        """
        return WizardSubmission(
            repoUrl=self.repo_url,
            slug=self.slug,
            platform=self.platform,
            token=None,
            depth=self.depth,
            language=self.language,
            model=self.model,
            filterMode=self.filter_mode,
            dirs=list(self.dirs),
            files=list(self.files),
            graphOnly=self.graph_only,
            ref=self.ref,
            fallbackModels=(
                list(self.fallback_models) if self.fallback_models is not None else None
            ),
            embeddingModel=self.embedding_model,
            customInstructions=self.custom_instructions,
            mcpServers=(
                dict(self.mcp_servers) if self.mcp_servers is not None else None
            ),
        )

from_submission(sub: WizardSubmission, *, desc: str | None = None) -> ProjectSettings classmethod

Project a :class:WizardSubmission onto the durable settings record.

sub.token is dropped (the credential registry owns it). desc carries an existing override forward, so re-seeding this record from a submission — which every WikiIndexingJob.start does, including the one a refresh drives — cannot clobber a user's edited description.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
@classmethod
def from_submission(
    cls, sub: WizardSubmission, *, desc: str | None = None
) -> ProjectSettings:
    """Project a :class:`WizardSubmission` onto the durable settings record.

    ``sub.token`` is dropped (the credential registry owns it). *desc* carries
    an existing override forward, so re-seeding this record from a submission
    — which every ``WikiIndexingJob.start`` does, including the one a refresh
    drives — cannot clobber a user's edited description.
    """
    return cls(
        slug=sub.slug,
        repoUrl=sub.repo_url,
        platform=sub.platform,
        model=sub.model,
        depth=sub.depth,
        language=sub.language,
        filterMode=sub.filter_mode,
        dirs=list(sub.dirs),
        files=list(sub.files),
        graphOnly=sub.graph_only,
        ref=sub.ref,
        fallbackModels=(
            list(sub.fallback_models) if sub.fallback_models is not None else None
        ),
        embeddingModel=sub.embedding_model,
        customInstructions=sub.custom_instructions,
        mcpServers=(
            dict(sub.mcp_servers) if sub.mcp_servers is not None else None
        ),
        desc=desc,
    )

to_submission() -> WizardSubmission

Rebuild the :class:WizardSubmission a re-index replays.

Always token-less: the clone tool's resolve_chain reads the durable credential itself at clone time, so a refresh never needs to carry one.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
def to_submission(self) -> WizardSubmission:
    """Rebuild the :class:`WizardSubmission` a re-index replays.

    Always token-less: the clone tool's ``resolve_chain`` reads the durable
    credential itself at clone time, so a refresh never needs to carry one.
    """
    return WizardSubmission(
        repoUrl=self.repo_url,
        slug=self.slug,
        platform=self.platform,
        token=None,
        depth=self.depth,
        language=self.language,
        model=self.model,
        filterMode=self.filter_mode,
        dirs=list(self.dirs),
        files=list(self.files),
        graphOnly=self.graph_only,
        ref=self.ref,
        fallbackModels=(
            list(self.fallback_models) if self.fallback_models is not None else None
        ),
        embeddingModel=self.embedding_model,
        customInstructions=self.custom_instructions,
        mcpServers=(
            dict(self.mcp_servers) if self.mcp_servers is not None else None
        ),
    )

PropertyNode

Bases: GraphNodeBase

A field / property / constant.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1989
1990
1991
1992
class PropertyNode(GraphNodeBase):
    """A field / property / constant."""

    type: Literal["Property"] = "Property"

QaAnswer

Bases: BaseModel

Complete Q&A answer returned after streaming finishes.

Top-level question/blocks/summary_sources/accessed_sources/ models_used/status always describe the LATEST turn — the shape a single-shot consumer (MCP ask_wiki, a fresh GET) already expects, byte-compatible with before turns existed. turns additively carries every PRIOR completed turn once a session is continued via WikiQaSession.follow_up; a never-followed-up answer has an empty turns list.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
class QaAnswer(BaseModel):
    """Complete Q&A answer returned after streaming finishes.

    Top-level ``question``/``blocks``/``summary_sources``/``accessed_sources``/
    ``models_used``/``status`` always describe the LATEST turn — the shape a
    single-shot consumer (MCP ``ask_wiki``, a fresh ``GET``) already expects,
    byte-compatible with before ``turns`` existed. ``turns`` additively carries
    every PRIOR completed turn once a session is continued via
    ``WikiQaSession.follow_up``; a never-followed-up answer has an
    empty ``turns`` list.
    """

    model_config = _CFG

    answer_id: str = Field(alias="answerId")
    from_page_id: str = Field(alias="fromPageId")
    # The current/latest turn's question text. Empty when unrecorded.
    question: str = Field(default="")
    summary_sources: list[str] = Field(alias="summarySources")
    model: str
    blocks: list[BlockUnion]
    # Deterministic provenance of the CURRENT turn (NOT the LLM's hand-picked
    # citations). ``accessed_sources`` is the de-duplicated trail of every graph
    # node + source file + page the probes actually touched (the probe tools
    # record an ``access`` event per call; the finalizer folds them). ``models_used``
    # is the distinct set of models that ran across the hypervisor + its probes,
    # across every turn (session-wide, never reset — see ``QaFinalizer.enrich``).
    # Both surface so the UI can show "what was read" + "which models" alongside
    # the answer. Defaulted so older persisted answers validate unchanged.
    accessed_sources: list[str] = Field(default_factory=list, alias="accessedSources")
    models_used: list[str] = Field(default_factory=list, alias="modelsUsed")
    # Run lifecycle of the CURRENT turn on the persisted snapshot — ``running``
    # until a terminal event finalizes the run. ``QaFinalizer.close`` sets
    # ``complete``; ``WikiQaSession.cancel`` sets ``cancelled``. This is the
    # field the MCP ``ask_wiki`` poll keys off of (no fragile "blocks
    # unchanged" guess).
    status: QaStatus = Field(default="running")
    # Which Q&A agent shape ran: ``deep`` is the hypervisor that fans out
    # retrieval probes; ``fast`` is a single root holding the retrieval surface
    # itself. Fixed for the life of an answer — a follow-up turn keeps the mode
    # its session started in, which is why ``QaTurn`` carries no copy. An
    # answer with no stored mode ran the probe fan-out, hence the ``deep``
    # default; the NEW-request default is a separate decision made at the wire
    # boundary.
    mode: Literal["fast", "deep"] = Field(default="deep")
    # Project slug that owns this answer. Persisted so ``resolve_qa_ctx``
    # can recover it after a process restart or any read-back path. NOT
    # ``exclude=True``: that would keep it off the wire but also out of the
    # store, leaving ``slug=""`` on every ctx lookup and breaking
    # ``wiki_search_pages`` (empty BM25 corpus). The FE TS type silently
    # ignores the extra field.
    slug: str = Field(default="")
    # Prior completed turns, oldest first. Empty for a
    # single-shot (never-followed-up) answer.
    turns: list[QaTurn] = Field(default_factory=list)

QaCancelledEvent

Bases: BaseModel

Terminal QA event: answer generation was cancelled.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1720
1721
1722
1723
1724
class QaCancelledEvent(BaseModel):
    """Terminal QA event: answer generation was cancelled."""

    model_config = _CFG
    type: Literal["cancelled"]

QaCompleteEvent

Bases: BaseModel

Terminal QA event: answer generation succeeded.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1712
1713
1714
1715
1716
1717
class QaCompleteEvent(BaseModel):
    """Terminal QA event: answer generation succeeded."""

    model_config = _CFG
    type: Literal["complete"]
    total_blocks: int = Field(alias="totalBlocks")

QaErrorEvent

Bases: BaseModel

Terminal QA event: answer generation failed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1727
1728
1729
1730
1731
1732
class QaErrorEvent(BaseModel):
    """Terminal QA event: answer generation failed."""

    model_config = _CFG
    type: Literal["error"]
    error: WikiError

QaEventUnion

Bases: RootModel[_QaEventAnnotated]

Discriminated union of all Q&A SSE events.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1756
1757
class QaEventUnion(RootModel[_QaEventAnnotated]):
    """Discriminated union of all Q&A SSE events."""

QaHeartbeatEvent

Bases: BaseModel

QA keep-alive event; consumers must ignore it.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1735
1736
1737
1738
1739
class QaHeartbeatEvent(BaseModel):
    """QA keep-alive event; consumers must ignore it."""

    model_config = _CFG
    type: Literal["heartbeat"]

QaTurn

Bases: BaseModel

One completed question+answer round within a continued QA answer.

Snapshotted from QaAnswer's top-level fields when WikiQaSession.follow_up starts a new turn on the same answer/session — see QaAnswer.turns.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
class QaTurn(BaseModel):
    """One completed question+answer round within a continued QA answer.

    Snapshotted from ``QaAnswer``'s top-level fields when
    ``WikiQaSession.follow_up`` starts a new turn on the same answer/session
    — see ``QaAnswer.turns``.
    """

    model_config = _CFG

    question: str
    blocks: list[BlockUnion] = Field(default_factory=list)
    summary_sources: list[str] = Field(default_factory=list, alias="summarySources")
    accessed_sources: list[str] = Field(default_factory=list, alias="accessedSources")
    models_used: list[str] = Field(default_factory=list, alias="modelsUsed")
    status: QaStatus = Field(default="complete")

QueuedEvent

Bases: BaseModel

Emitted when a job is accepted into the queue.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1423
1424
1425
1426
1427
1428
1429
1430
class QueuedEvent(BaseModel):
    """Emitted when a job is accepted into the queue."""

    model_config = _CFG
    type: Literal["queued"]
    job_id: str = Field(alias="jobId")
    slug: str
    total_count: int = Field(alias="totalCount")

RefreshDecision

Bases: BaseModel

Which refresh path a project takes, and — when it is full — why.

The ONE answer to "scoped or full", computed once per refresh and then carried everywhere it is needed: the route's response body, the IndexingJob snapshot (so a console can say why a rebuild was full), and the resume branch that must re-drive a stranded scoped job the same way it ran the first time. One record rather than a mode flag plus a reason string plus a mismatch list, because those three can only ever disagree.

:meth:decide is PURE — it takes the already-probed inputs as arguments and touches no store, no clock and no subprocess, so the whole policy is testable without a repository on disk. Probing belongs to the callers at the edges.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
class RefreshDecision(BaseModel):
    """Which refresh path a project takes, and — when it is full — why.

    The ONE answer to "scoped or full", computed once per refresh and then
    carried everywhere it is needed: the route's response body, the
    ``IndexingJob`` snapshot (so a console can say why a rebuild was full), and
    the resume branch that must re-drive a stranded scoped job the same way it
    ran the first time. One record rather than a mode flag plus a reason string
    plus a mismatch list, because those three can only ever disagree.

    :meth:`decide` is PURE — it takes the already-probed inputs as arguments and
    touches no store, no clock and no subprocess, so the whole policy is
    testable without a repository on disk. Probing belongs to the callers at the
    edges.
    """

    model_config = ConfigDict(extra="forbid", populate_by_name=True, frozen=True)

    path: RefreshPath
    # ``None`` IFF ``path == "scoped"`` — enforced below rather than left to
    # each reader, the same "credential is None IFF source is anonymous"
    # discipline ``CredentialCandidate`` already uses: a consumer checks one
    # field, never two that could contradict.
    reason: RefreshFullReason | None = None
    # Populated only for ``fingerprint_mismatch``; every disagreeing field, so a
    # console renders each reason a rebuild was needed in one pass.
    mismatches: list[FingerprintMismatch] = Field(default_factory=list)

    @model_validator(mode="after")
    def _reason_iff_full(self) -> RefreshDecision:
        """A full path always names its reason; a scoped path never has one."""
        if (self.reason is None) is (self.path == "full"):
            raise ValueError(
                "RefreshDecision.reason must be set for path='full' and absent "
                f"for path='scoped' (got path={self.path!r}, reason={self.reason!r})"
            )
        return self

    @classmethod
    def decide(
        cls,
        *,
        mode: RefreshMode,
        project: Project,
        current: IndexFingerprint,
    ) -> RefreshDecision:
        """Choose the refresh path for *project* under *mode*.

        Ordered cheapest-and-most-decisive first, so the reason a reader is
        shown is the one that would still hold if everything after it were
        fixed. ``current`` is what THIS refresh would build with; it is compared
        against what the stored index was actually built with
        (``Project.fingerprint``) through the same three-state
        :class:`FingerprintDecision` the graph phase already stamps — "never
        compared" stays distinct from "compared and matched", so an index
        nothing fingerprinted rebuilds instead of silently reusing.
        """
        if mode == "full":
            return cls(path="full", reason="requested")
        # A graph-only project builds no manifest, so a content diff against it
        # would read every file as added on every run.
        if project.graph_only:
            return cls(path="full", reason="graph_only")
        # Nothing to compute a delta against, and nothing to attribute it to.
        if not project.commit_sha:
            return cls(path="full", reason="no_prior_index")
        verdict = FingerprintDecision.compute(current, project.fingerprint)
        if verdict.reason == "unknown":
            return cls(path="full", reason="fingerprint_unknown")
        if verdict.reason == "mismatch":
            return cls(
                path="full",
                reason="fingerprint_mismatch",
                mismatches=verdict.mismatches,
            )
        return cls(path="scoped")

decide(*, mode: RefreshMode, project: Project, current: IndexFingerprint) -> RefreshDecision classmethod

Choose the refresh path for project under mode.

Ordered cheapest-and-most-decisive first, so the reason a reader is shown is the one that would still hold if everything after it were fixed. current is what THIS refresh would build with; it is compared against what the stored index was actually built with (Project.fingerprint) through the same three-state :class:FingerprintDecision the graph phase already stamps — "never compared" stays distinct from "compared and matched", so an index nothing fingerprinted rebuilds instead of silently reusing.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
@classmethod
def decide(
    cls,
    *,
    mode: RefreshMode,
    project: Project,
    current: IndexFingerprint,
) -> RefreshDecision:
    """Choose the refresh path for *project* under *mode*.

    Ordered cheapest-and-most-decisive first, so the reason a reader is
    shown is the one that would still hold if everything after it were
    fixed. ``current`` is what THIS refresh would build with; it is compared
    against what the stored index was actually built with
    (``Project.fingerprint``) through the same three-state
    :class:`FingerprintDecision` the graph phase already stamps — "never
    compared" stays distinct from "compared and matched", so an index
    nothing fingerprinted rebuilds instead of silently reusing.
    """
    if mode == "full":
        return cls(path="full", reason="requested")
    # A graph-only project builds no manifest, so a content diff against it
    # would read every file as added on every run.
    if project.graph_only:
        return cls(path="full", reason="graph_only")
    # Nothing to compute a delta against, and nothing to attribute it to.
    if not project.commit_sha:
        return cls(path="full", reason="no_prior_index")
    verdict = FingerprintDecision.compute(current, project.fingerprint)
    if verdict.reason == "unknown":
        return cls(path="full", reason="fingerprint_unknown")
    if verdict.reason == "mismatch":
        return cls(
            path="full",
            reason="fingerprint_mismatch",
            mismatches=verdict.mismatches,
        )
    return cls(path="scoped")

RepoCredential

Bases: BaseModel

A persisted repository credential — a git token OR an SSH/deploy key.

Stored per-scope in the isolated credential store (CredentialScope — see mewbo_graph.wiki.credentials, which also owns the host-covers-repo sharing rule). The credential itself carries NO scope field: the scope is the store KEY, and it is stamped into the blob at save time — one binding, not two that can disagree. Plaintext-at-rest behind the store's _encode/_decode seam; ALWAYS redacted in-flight (SSE / transcript / logs).

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
class RepoCredential(BaseModel):
    """A persisted repository credential — a git token OR an SSH/deploy key.

    Stored per-scope in the isolated credential store (``CredentialScope`` —
    see ``mewbo_graph.wiki.credentials``, which also owns the host-covers-repo
    sharing rule). The credential itself carries NO scope field: the scope is
    the store KEY, and it is stamped into the blob at save time — one binding,
    not two that can disagree. Plaintext-at-rest behind the store's
    ``_encode``/``_decode`` seam; ALWAYS redacted in-flight (SSE / transcript /
    logs).
    """

    model_config = _CFG

    kind: Literal["token", "ssh_key"]
    value: str
    username: str | None = None
    updated_at: str | None = Field(default=None, alias="updatedAt")

    @field_validator("value", mode="before")
    @classmethod
    def _strip_value(cls, v: Any) -> Any:
        """Strip surrounding whitespace BEFORE validation — a real, silent footgun.

        A PAT pasted from a terminal or a file carries a trailing newline; git
        injects it verbatim into the URL/header and the remote answers 401 with
        no hint that an invisible character is the cause. Stripping is safe for
        an SSH key too: only the surrounding whitespace goes (internal newlines
        are preserved), and ``clone._ssh_env_for`` re-appends the terminating
        newline the key file requires.
        """
        return v.strip() if isinstance(v, str) else v

    @field_validator("value")
    @classmethod
    def _value_not_empty(cls, v: str) -> str:
        """Reject an empty credential — an empty token/key is never useful.

        Runs AFTER :meth:`_strip_value`, so a whitespace-only value is empty here.
        """
        if not v:
            raise ValueError("credential value must not be empty")
        return v

    @property
    def dedup_key(self) -> tuple[str, str, str | None]:
        """The identity two credentials must share to be considered the same one.

        ``username`` is part of it deliberately: a value shared by two usernames
        (a GitLab ``oauth2`` deploy token vs a PAT) authenticates DIFFERENTLY, so
        the resolution chain must try both rather than dedup the second away.
        """
        return (self.kind, self.value, self.username)

dedup_key: tuple[str, str, str | None] property

The identity two credentials must share to be considered the same one.

username is part of it deliberately: a value shared by two usernames (a GitLab oauth2 deploy token vs a PAT) authenticates DIFFERENTLY, so the resolution chain must try both rather than dedup the second away.

ScannedEvent

Bases: BaseModel

Emitted after a file has been analysed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1443
1444
1445
1446
1447
1448
1449
1450
class ScannedEvent(BaseModel):
    """Emitted after a file has been analysed."""

    model_config = _CFG
    type: Literal["scanned"]
    file: str
    index: int
    total_count: int = Field(alias="totalCount")

ScanningEvent

Bases: BaseModel

Emitted just before a file is analysed.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1433
1434
1435
1436
1437
1438
1439
1440
class ScanningEvent(BaseModel):
    """Emitted just before a file is analysed."""

    model_config = _CFG
    type: Literal["scanning"]
    file: str
    index: int
    total_count: int = Field(alias="totalCount")

ScopePreview

Bases: BaseModel

Counts describing what one scoped refresh actually touched.

Produced by RefreshReport.scope_preview() in mewbo_graph.wiki.refresh and carried on both the job's SSE stream and its snapshot. It lives HERE rather than beside its producer because IndexingJob persists it: a model the store round-trips is a wire type, and refresh.py already imports this module, so defining it there and importing it back would close a cycle.

Deliberately FLAT rather than mirroring RefreshReport's four-stage shape. The reader is a progress panel rendering a row of counts, and a nested payload would make it walk three levels to reach an integer it displays verbatim.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
class ScopePreview(BaseModel):
    """Counts describing what one scoped refresh actually touched.

    Produced by ``RefreshReport.scope_preview()`` in
    ``mewbo_graph.wiki.refresh`` and carried on both the job's SSE stream and
    its snapshot. It lives HERE rather than beside its producer because
    ``IndexingJob`` persists it: a model the store round-trips is a wire type,
    and ``refresh.py`` already imports this module, so defining it there and
    importing it back would close a cycle.

    Deliberately FLAT rather than mirroring ``RefreshReport``'s four-stage
    shape. The reader is a progress panel rendering a row of counts, and a
    nested payload would make it walk three levels to reach an integer it
    displays verbatim.
    """

    model_config = ConfigDict(extra="forbid", populate_by_name=True, frozen=True)

    files_added: int = Field(alias="filesAdded")
    files_modified: int = Field(alias="filesModified")
    files_deleted: int = Field(alias="filesDeleted")
    # Files re-parsed whose entity + edge signatures came back identical — the
    # Salsa early cutoff. High relative to ``files_modified`` means the diff was
    # mostly comments/formatting and cost nothing downstream.
    early_cutoff_files: int = Field(alias="earlyCutoffFiles")
    affected_entities: int = Field(alias="affectedEntities")
    memory_kept: int = Field(alias="memoryKept")
    memory_invalidated: int = Field(alias="memoryInvalidated")
    memory_revalidated: int = Field(alias="memoryRevalidated")
    pages_keep: int = Field(alias="pagesKeep")
    pages_edit: int = Field(alias="pagesEdit")
    pages_regenerate: int = Field(alias="pagesRegenerate")
    new_pages: int = Field(alias="newPages")
    # LLM calls the deterministic pass actually made — the memory reconciler's
    # drift band is the only stage that can reach one. Non-zero here is what
    # makes "Free" an honest word rather than an approximate one.
    llm_calls: int = Field(alias="llmCalls")

SourceRef

Bases: BaseModel

Source-file reference with optional line range.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
363
364
365
366
367
368
class SourceRef(BaseModel):
    """Source-file reference with optional line range."""

    model_config = _CFG
    path: str
    lines: str | None = None

SourcesBlock

Bases: BaseModel

Cited sources block.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
317
318
319
320
321
322
class SourcesBlock(BaseModel):
    """Cited sources block."""

    model_config = _CFG
    kind: Literal["sources"]
    items: list[str]

StepMeasurement

Bases: BaseModel

One completed step's observed elapsed time and final unit count.

The record belongs on :class:Project, the stable current-index snapshot, rather than a job that the next index has to search for. It stays bounded by the declared plan: one measurement per step, never per file, node, or page.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
class StepMeasurement(BaseModel):
    """One completed step's observed elapsed time and final unit count.

    The record belongs on :class:`Project`, the stable current-index snapshot,
    rather than a job that the next index has to search for. It stays bounded by
    the declared plan: one measurement per step, never per file, node, or page.
    """

    model_config = _CFG

    seconds: float = Field(ge=0)
    units: int | None = Field(default=None, ge=0)

    @classmethod
    def from_record(
        cls, record: StepRecord, now: datetime
    ) -> StepMeasurement | None:
        """Project one completed non-skipped ledger record into a measurement.

        Cost: ``O(1)``. The clock arrives from the caller for testability. A
        skipped resumed phase has no fresh cost, so it must not replace an earlier
        measurement with a near-zero duration.
        """
        if record.state != "done":
            return None
        seconds = record.elapsed_seconds(now)
        if seconds is None or seconds < 0:
            return None
        return cls(seconds=seconds, units=record.total)

    def blended_with(self, observed: StepMeasurement) -> StepMeasurement:
        """Blend a newer reading into this one. Cost: ``O(1)``.

        The previous reading keeps 75% of the result and the newest completed run
        supplies 25%, which damps one-off noise without making calibration stale.
        When both readings count units, blend their rates projected onto the new
        count — otherwise a doubled repository would inherit half its time.
        """
        if self.units and observed.units:
            prior_at_new_size = (self.seconds / self.units) * observed.units
            seconds = (prior_at_new_size * 0.75) + (observed.seconds * 0.25)
        else:
            seconds = (self.seconds * 0.75) + (observed.seconds * 0.25)
        return StepMeasurement(seconds=seconds, units=observed.units)

blended_with(observed: StepMeasurement) -> StepMeasurement

Blend a newer reading into this one. Cost: O(1).

The previous reading keeps 75% of the result and the newest completed run supplies 25%, which damps one-off noise without making calibration stale. When both readings count units, blend their rates projected onto the new count — otherwise a doubled repository would inherit half its time.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
78
79
80
81
82
83
84
85
86
87
88
89
90
91
def blended_with(self, observed: StepMeasurement) -> StepMeasurement:
    """Blend a newer reading into this one. Cost: ``O(1)``.

    The previous reading keeps 75% of the result and the newest completed run
    supplies 25%, which damps one-off noise without making calibration stale.
    When both readings count units, blend their rates projected onto the new
    count — otherwise a doubled repository would inherit half its time.
    """
    if self.units and observed.units:
        prior_at_new_size = (self.seconds / self.units) * observed.units
        seconds = (prior_at_new_size * 0.75) + (observed.seconds * 0.25)
    else:
        seconds = (self.seconds * 0.75) + (observed.seconds * 0.25)
    return StepMeasurement(seconds=seconds, units=observed.units)

from_record(record: StepRecord, now: datetime) -> StepMeasurement | None classmethod

Project one completed non-skipped ledger record into a measurement.

Cost: O(1). The clock arrives from the caller for testability. A skipped resumed phase has no fresh cost, so it must not replace an earlier measurement with a near-zero duration.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
@classmethod
def from_record(
    cls, record: StepRecord, now: datetime
) -> StepMeasurement | None:
    """Project one completed non-skipped ledger record into a measurement.

    Cost: ``O(1)``. The clock arrives from the caller for testability. A
    skipped resumed phase has no fresh cost, so it must not replace an earlier
    measurement with a near-zero duration.
    """
    if record.state != "done":
        return None
    seconds = record.elapsed_seconds(now)
    if seconds is None or seconds < 0:
        return None
    return cls(seconds=seconds, units=record.total)

SummaryReadyEvent

Bases: BaseModel

Emitted once summary sources are known.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1678
1679
1680
1681
1682
1683
class SummaryReadyEvent(BaseModel):
    """Emitted once summary sources are known."""

    model_config = _CFG
    type: Literal["summary_ready"]
    sources: list[str]

TableBlock

Bases: BaseModel

Table block.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
325
326
327
328
329
330
331
class TableBlock(BaseModel):
    """Table block."""

    model_config = _CFG
    kind: Literal["table"]
    head: list[str]
    rows: list[list[InlineNode]]

TocEntry

Bases: BaseModel

In-page table-of-contents entry.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
230
231
232
233
234
235
236
237
class TocEntry(BaseModel):
    """In-page table-of-contents entry."""

    model_config = _CFG

    id: str
    label: str
    lvl: Literal[1, 2, 3]

UlBlock

Bases: BaseModel

Unordered-list block.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
300
301
302
303
304
305
class UlBlock(BaseModel):
    """Unordered-list block."""

    model_config = _CFG
    kind: Literal["ul"]
    items: list[InlineNode]

WikiError

Bases: BaseModel

Typed error returned by wiki API endpoints and streamed events.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
class WikiError(BaseModel):
    """Typed error returned by wiki API endpoints and streamed events."""

    model_config = _CFG

    code: WikiErrorCode
    message: str
    hint: str | None = None
    fields: dict[str, str] | None = None
    retry_after: float | None = Field(default=None, alias="retryAfter")

WikiPage

Bases: BaseModel

Full wiki page including body, TOC, and sidebar nav.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
383
384
385
386
387
388
389
390
391
392
393
class WikiPage(BaseModel):
    """Full wiki page including body, TOC, and sidebar nav."""

    model_config = _CFG

    id: str
    title: str
    frontmatter: Frontmatter
    body: str
    toc: list[TocEntry]
    nav: list[NavEntry]

WizardSubmission

Bases: BaseModel

Wizard POST body for triggering a new indexing job.

repo_url is optional: a NON-git "catalog" workspace (programmatic document ingestion via CatalogIngestor / POST .../documents) has no clone URL. The git indexing pipeline still requires it — its own validation rejects a clone with an empty URL — but the model itself no longer forces one so the same submission shape carries a repo-less catalog project.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
class WizardSubmission(BaseModel):
    """Wizard POST body for triggering a new indexing job.

    ``repo_url`` is optional: a NON-git "catalog" workspace (programmatic
    document ingestion via ``CatalogIngestor`` / ``POST .../documents``) has no
    clone URL. The git indexing pipeline still requires it — its own validation
    rejects a clone with an empty URL — but the model itself no longer forces
    one so the same submission shape carries a repo-less catalog project.
    """

    model_config = _CFG

    repo_url: str | None = Field(default=None, alias="repoUrl")
    slug: str
    platform: PlatformId
    token: str | None = None
    depth: DepthMode
    language: str
    model: str
    filter_mode: FilterMode = Field(alias="filterMode")
    dirs: list[str]
    files: list[str]
    # Request a DETERMINISTIC, zero-LLM index: clone → scan → AST graph →
    # finalize, SKIPPING enrich/plan/pages. Produces a populated graph + zero
    # pages (developer mode). Honoured ONLY when ``runtime.developer_mode`` is
    # on — the index route forces it False otherwise, so an unprivileged caller
    # can never opt into the no-docs path. Optional + defaulted so ordinary
    # submissions are unchanged.
    graph_only: bool = Field(default=False, alias="graphOnly")
    # Optional branch/tag/sha to clone; null = the repo's default branch (the
    # behaviour when omitted is unchanged).
    ref: str | None = Field(default=None)
    # Ordered model ladder the indexer falls back through when its primary model
    # fails. ``None`` = inherit the configured fallback policy; a list
    # overrides it for this job only.
    fallback_models: list[str] | None = Field(default=None, alias="fallbackModels")
    # Embedding model this project's vectors are built and searched with.
    # ``None`` = inherit the deployment default (``wiki.embedding.model``), which
    # is what every project indexed before this field existed carries.
    #
    # It is per-PROJECT rather than per-job because the write side and the read
    # side have to agree: two embedding models rarely share a vector width, and a
    # store holding both returns wrong neighbours rather than erroring. Changing
    # it is therefore a full rebuild, which nothing here has to arrange —
    # ``IndexFingerprint.embedding_model`` already records the model a run
    # actually embedded with, and a mismatch against it forces the full path.
    embedding_model: str | None = Field(default=None, alias="embeddingModel")
    # Free-text operator guidance appended to the indexer's playbook.
    # ``None`` = no guidance.
    custom_instructions: str | None = Field(default=None, alias="customInstructions")
    # External MCP servers to attach for the duration of an index, in the
    # standard ``{name: {command|url, …}}`` MCP config shape. ``None`` = attach
    # nothing (today's behaviour). Operator-set at onboarding or in settings
    # ONLY — an MCP server entry names a process to spawn, so no per-run path
    # and nothing an agent can reach may write it.
    mcp_servers: dict[str, dict] | None = Field(default=None, alias="mcpServers")

    @field_validator("custom_instructions")
    @classmethod
    def check_custom_instructions(cls, v: str | None) -> str | None:
        """Strip, collapse blank to ``None``, and cap the length.

        THE rule for this field, delegated to by :class:`ProjectSettings` and by
        the api's ``ProjectSettingsPatch`` so the three hops cannot disagree
        about what a valid value is.

        The cap is why this is a validator rather than a bare field: the text is
        appended to the page-writer prompt of EVERY page in the fan-out, so it
        is paid once per page, not once per index — a repository with 200 pages
        pays for it 200 times.
        """
        if v is None:
            return None
        stripped = v.strip()
        if not stripped:
            return None
        if len(stripped) > MAX_CUSTOM_INSTRUCTIONS_CHARS:
            raise ValueError(
                f"customInstructions is {len(stripped)} characters, over the "
                f"{MAX_CUSTOM_INSTRUCTIONS_CHARS}-character limit. This text is "
                "appended to the prompt of every page the indexer writes, so it "
                "is paid once per page rather than once per index."
            )
        return stripped

    @field_validator("mcp_servers")
    @classmethod
    def check_mcp_servers(cls, v: dict[str, dict] | None) -> dict[str, dict] | None:
        """Validate the server map's shape; an empty map collapses to ``None``.

        THE rule for this field, shared by the same three hops as
        :meth:`check_custom_instructions`. Deliberately shallow: the VALUE is the
        standard MCP server-config shape that ``get_merged_mcp_config`` consumes
        verbatim, and re-modelling it here would be a second, drifting copy of a
        schema this package does not own. What is checked is what this layer
        genuinely owns — that the map is name→object, with usable names.
        """
        if v is None:
            return None
        if not v:
            return None
        for name, entry in v.items():
            if not isinstance(name, str) or not name.strip():
                raise ValueError("mcpServers keys must be non-empty server names")
            if not isinstance(entry, dict):
                raise ValueError(
                    f"mcpServers[{name!r}] must be an object describing one MCP server"
                )
        return {name.strip(): entry for name, entry in v.items()}

check_custom_instructions(v: str | None) -> str | None classmethod

Strip, collapse blank to None, and cap the length.

THE rule for this field, delegated to by :class:ProjectSettings and by the api's ProjectSettingsPatch so the three hops cannot disagree about what a valid value is.

The cap is why this is a validator rather than a bare field: the text is appended to the page-writer prompt of EVERY page in the fan-out, so it is paid once per page, not once per index — a repository with 200 pages pays for it 200 times.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
@field_validator("custom_instructions")
@classmethod
def check_custom_instructions(cls, v: str | None) -> str | None:
    """Strip, collapse blank to ``None``, and cap the length.

    THE rule for this field, delegated to by :class:`ProjectSettings` and by
    the api's ``ProjectSettingsPatch`` so the three hops cannot disagree
    about what a valid value is.

    The cap is why this is a validator rather than a bare field: the text is
    appended to the page-writer prompt of EVERY page in the fan-out, so it
    is paid once per page, not once per index — a repository with 200 pages
    pays for it 200 times.
    """
    if v is None:
        return None
    stripped = v.strip()
    if not stripped:
        return None
    if len(stripped) > MAX_CUSTOM_INSTRUCTIONS_CHARS:
        raise ValueError(
            f"customInstructions is {len(stripped)} characters, over the "
            f"{MAX_CUSTOM_INSTRUCTIONS_CHARS}-character limit. This text is "
            "appended to the prompt of every page the indexer writes, so it "
            "is paid once per page rather than once per index."
        )
    return stripped

check_mcp_servers(v: dict[str, dict] | None) -> dict[str, dict] | None classmethod

Validate the server map's shape; an empty map collapses to None.

THE rule for this field, shared by the same three hops as :meth:check_custom_instructions. Deliberately shallow: the VALUE is the standard MCP server-config shape that get_merged_mcp_config consumes verbatim, and re-modelling it here would be a second, drifting copy of a schema this package does not own. What is checked is what this layer genuinely owns — that the map is name→object, with usable names.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
@field_validator("mcp_servers")
@classmethod
def check_mcp_servers(cls, v: dict[str, dict] | None) -> dict[str, dict] | None:
    """Validate the server map's shape; an empty map collapses to ``None``.

    THE rule for this field, shared by the same three hops as
    :meth:`check_custom_instructions`. Deliberately shallow: the VALUE is the
    standard MCP server-config shape that ``get_merged_mcp_config`` consumes
    verbatim, and re-modelling it here would be a second, drifting copy of a
    schema this package does not own. What is checked is what this layer
    genuinely owns — that the map is name→object, with usable names.
    """
    if v is None:
        return None
    if not v:
        return None
    for name, entry in v.items():
        if not isinstance(name, str) or not name.strip():
            raise ValueError("mcpServers keys must be non-empty server names")
        if not isinstance(entry, dict):
            raise ValueError(
                f"mcpServers[{name!r}] must be an object describing one MCP server"
            )
    return {name.strip(): entry for name, entry in v.items()}

make_graph_node(**data: Any) -> GraphNode

Construct the per-kind GraphNode subclass for a dynamic type.

Thin factory over :data:_NODE_CLS_BY_TYPE for the (few) call sites whose type is only known at runtime — collapsing an otherwise-repeated per-kind branch (DRY). Static sites should use the per-kind class directly.

Source code in packages/mewbo_graph/src/mewbo_graph/wiki/types.py
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
def make_graph_node(**data: Any) -> GraphNode:
    """Construct the per-kind ``GraphNode`` subclass for a dynamic ``type``.

    Thin factory over :data:`_NODE_CLS_BY_TYPE` for the (few) call sites whose
    ``type`` is only known at runtime — collapsing an otherwise-repeated
    per-kind branch (DRY). Static sites should use the per-kind class directly.
    """
    kind = data.get("type")
    try:
        cls = _NODE_CLS_BY_TYPE[kind]  # type: ignore[index]
    except KeyError:
        raise ValueError(f"unknown graph node type: {kind!r}") from None
    # ``cls`` is a concrete union member, but the type checker only knows it as
    # ``type[GraphNodeBase]`` — narrow to the union at this single seam.
    return cast("GraphNode", cls(**data))

Source Capability Graph (mewbo_graph.scg)

mewbo_graph.scg.router

ScgRouter — the cheap query→route mechanism over the SCG (spec §6).

Routing is the graph's only query-time job: control routing. Given a natural -language query, the router embeds it, vector-searches seed nodes in the store, expands one hop along capability/route edges to assemble candidate :class:RouteRecipes, and ranks them with a deterministic, zero-LLM score (cosine(seed) + edge weight). The agentic traversal engine consumes the ranked recipes; spending sub-agents is a downstream concern.

This is the lightweight pre-rank, not the full hypothesis search. It mirrors HippoRAG2's "cheap structural pre-rank before spending agents" stance: route first, traverse second.

SCALE SEAM — Personalized PageRank

A query-seeded Personalized PageRank (PPR) over the SCG with hub damping is the documented upgrade for ranking quality at catalog scale (HippoRAG2 2502.14802, PathRAG 2502.14902). It lands behind this same route() signature — callers never change. It is deliberately NOT implemented now: the additive cosine + weight score is cheaper, fully deterministic, and sufficient for the small catalogs SCG ships with first.

ScgRouter

Cheap, deterministic query→route over the Source Capability Graph.

Dependency-injected with an :class:ScgStore and a query embedder (the wiki :class:~mewbo_graph.wiki.embedder.Embedder by default; tests inject a fake). An OPTIONAL :class:~mewbo_graph.scg.memory_bridge.ScgMemoryBridge makes routing memory-aware: when injected, the top-k learned connector notes for the query boost pathways already known to produce results and damp discovered dead ends — a retrieval-plus-arithmetic step, NO LLM, so the zero-LLM routing core is preserved. Omit it (None) and routing is memory-blind (the historical structure-only behaviour).

Holds no per-query state — all behaviour is expressed over the injected collaborators, so a single router instance is reusable across queries.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/router.py
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
class ScgRouter:
    """Cheap, deterministic query→route over the Source Capability Graph.

    Dependency-injected with an :class:`ScgStore` and a query embedder (the wiki
    :class:`~mewbo_graph.wiki.embedder.Embedder` by default; tests inject a fake).
    An OPTIONAL :class:`~mewbo_graph.scg.memory_bridge.ScgMemoryBridge` makes
    routing **memory-aware**: when injected, the top-k learned connector
    notes for the query boost pathways already known to produce results and damp
    discovered dead ends — a retrieval-plus-arithmetic step, NO LLM, so the
    zero-LLM routing core is preserved. Omit it (``None``) and routing is
    memory-blind (the historical structure-only behaviour).

    Holds no per-query state — all behaviour is expressed over the injected
    collaborators, so a single router instance is reusable across queries.
    """

    # Edges traversed when expanding a seed toward a recipe. Capability/route
    # edges connect a capability to what it PRODUCES/CONSUMES and how its
    # fields RESOLVE — the executable pathways the router proposes.
    _EXPAND_KINDS: frozenset[str] = frozenset(
        {"PRODUCES", "CONSUMES", "RESOLVES_TO", "SUPPORTS_QUERY"}
    )

    def __init__(
        self,
        *,
        store: ScgStore,
        embedder: _QueryEmbedder,
        memory_bridge: ScgMemoryBridge | None = None,
    ) -> None:
        """Bind the SCG store + query embedder (+ optional memory bridge)."""
        self.store = store
        self.embedder = embedder
        self.memory_bridge = memory_bridge

    def route(self, query: str, *, k: int = 5) -> list[RouteRecipe]:
        """Return up to *k* :class:`RouteRecipe`s best matching *query*.

        Thin wrapper over :meth:`route_with_memory` that drops the bias map — the
        stable, historical signature for callers that only need the ranked
        recipes (the memory bias still applies when a bridge is injected).
        """
        recipes, _bias = self.route_with_memory(query, k=k)
        return recipes

    def route_with_memory(
        self, query: str, *, k: int = 5
    ) -> tuple[list[RouteRecipe], ScgMemoryBias]:
        """Rank recipes AND return the learned-memory bias map that shaped them.

        Embed → vector-search seed nodes → expand one hop along capability/route
        edges → assemble candidate recipes → rank by ``cosine(seed) + edge
        weight + memory_boost`` (still zero-LLM — the memory term is a vector
        read + a polarity-weighted sum). Returns ``([], empty_bias)`` for an
        empty graph or no match.

        The returned :class:`ScgMemoryBias` carries the per-capability anchored
        HINTS too, so the ``scg_route`` plugin tool can surface "how to call this
        right" guidance on each recipe without re-reading memory.

        Honours the ambient :class:`ScgScope`: a candidate recipe whose
        steps reach an out-of-scope source is dropped, AND a memory note anchored
        to an out-of-scope source contributes no bias — routing and learning both
        stay inside the workspace over the otherwise GLOBAL shared graph.
        """
        qvec = self.embedder.embed_query(query)
        seeds = self.store.vector_search(qvec, k=k)
        if not seeds:
            return [], ScgMemoryBias.empty()

        # source_key → recipe, so a vector-search hit (mapped to its graph
        # anchor) can be matched against the recipes reachable from it. Filter to
        # the workspace scope here so an out-of-scope pathway is never a candidate.
        recipes = {
            r.source_key: r
            for r in self.store.list_recipes()
            if ScgScope.permits_recipe_steps(r.steps)
        }
        # The learned-memory bias for this query (empty when no bridge is
        # injected or no embedding backend is configured — degrades gracefully).
        bias = self._memory_bias(qvec)
        # Pre-index inbound capability/route edges by target ONCE per route()
        # (one full edge scan), not once per seed — the inbound expansion below
        # is then an O(1) dict lookup. (PPR is the documented scale upgrade
        # behind this same signature.)
        inbound = self._inbound_index()
        best: dict[SourceKey, float] = {}
        for node_id, sim in seeds:
            node = self.store.get_node(node_id)
            if node is None:
                continue
            for key, weight in self._candidate_keys(node.source_key, inbound):
                if key not in recipes:
                    continue
                # cosine(seed) + edge weight + learned-memory boost (the memory-bias term;
                # 0.0 when the pathway's steps carry no learned signal).
                score = sim + weight + bias.boost_for_steps(recipes[key].steps)
                if score > best.get(key, float("-inf")):
                    best[key] = score

        ranked = sorted(
            best.items(),
            # Score desc, then source_key asc — a deterministic tie-break.
            key=lambda kv: (-kv[1], kv[0]),
        )
        # Honour the documented bound: one-hop expansion can map the seeds to
        # more than *k* recipe candidates, so slice the ranked result to *k*.
        return [recipes[key] for key, _ in ranked[:k]], bias

    def _memory_bias(self, qvec: list[float]) -> ScgMemoryBias:
        """Build the learned-memory bias for *qvec* (empty when no bridge)."""
        if self.memory_bridge is None:
            return ScgMemoryBias.empty()
        return ScgMemoryBias.for_query(
            self.memory_bridge, _CONNECTOR_SLUG, qvec, k=10
        )

    def _candidate_keys(
        self, seed_key: SourceKey, inbound: dict[SourceKey, list[ScgEdge]]
    ) -> list[tuple[SourceKey, float]]:
        """``(recipe_key, edge_weight)`` candidates reachable from *seed_key*.

        The seed itself is a candidate at weight ``0.0`` (it may anchor a recipe
        directly). Each one-hop capability/route edge contributes its target as
        a candidate carrying the edge's weight, so a seed entity reachable from
        a capability's recipe still routes. ``inbound`` is the pre-built
        target→edges index (see :meth:`_inbound_index`) so this is allocation-
        and scan-free per seed.
        """
        out: list[tuple[SourceKey, float]] = [(seed_key, 0.0)]
        for edge in self.store.neighbors(seed_key):
            if edge.kind in self._EXPAND_KINDS:
                out.append((edge.target, edge.weight))
        # Also expand inbound: an entity seed reaches the capability that
        # PRODUCES it (the recipe lives on the capability, not the entity).
        for edge in inbound.get(seed_key, ()):
            out.append((edge.source, edge.weight))
        return out

    def _inbound_index(self) -> dict[SourceKey, list[ScgEdge]]:
        """Index capability/route edges by ``target`` — one full edge scan.

        Built once per :meth:`route` so the per-seed inbound expansion is an
        O(1) dict lookup instead of a fresh ``list_edges()`` scan per seed.
        """
        index: dict[SourceKey, list[ScgEdge]] = {}
        for edge in self.store.list_edges():
            if edge.kind in self._EXPAND_KINDS:
                index.setdefault(edge.target, []).append(edge)
        return index

__init__(*, store: ScgStore, embedder: _QueryEmbedder, memory_bridge: ScgMemoryBridge | None = None) -> None

Bind the SCG store + query embedder (+ optional memory bridge).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/router.py
72
73
74
75
76
77
78
79
80
81
82
def __init__(
    self,
    *,
    store: ScgStore,
    embedder: _QueryEmbedder,
    memory_bridge: ScgMemoryBridge | None = None,
) -> None:
    """Bind the SCG store + query embedder (+ optional memory bridge)."""
    self.store = store
    self.embedder = embedder
    self.memory_bridge = memory_bridge

route(query: str, *, k: int = 5) -> list[RouteRecipe]

Return up to k :class:RouteRecipes best matching query.

Thin wrapper over :meth:route_with_memory that drops the bias map — the stable, historical signature for callers that only need the ranked recipes (the memory bias still applies when a bridge is injected).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/router.py
84
85
86
87
88
89
90
91
92
def route(self, query: str, *, k: int = 5) -> list[RouteRecipe]:
    """Return up to *k* :class:`RouteRecipe`s best matching *query*.

    Thin wrapper over :meth:`route_with_memory` that drops the bias map — the
    stable, historical signature for callers that only need the ranked
    recipes (the memory bias still applies when a bridge is injected).
    """
    recipes, _bias = self.route_with_memory(query, k=k)
    return recipes

route_with_memory(query: str, *, k: int = 5) -> tuple[list[RouteRecipe], ScgMemoryBias]

Rank recipes AND return the learned-memory bias map that shaped them.

Embed → vector-search seed nodes → expand one hop along capability/route edges → assemble candidate recipes → rank by cosine(seed) + edge weight + memory_boost (still zero-LLM — the memory term is a vector read + a polarity-weighted sum). Returns ([], empty_bias) for an empty graph or no match.

The returned :class:ScgMemoryBias carries the per-capability anchored HINTS too, so the scg_route plugin tool can surface "how to call this right" guidance on each recipe without re-reading memory.

Honours the ambient :class:ScgScope: a candidate recipe whose steps reach an out-of-scope source is dropped, AND a memory note anchored to an out-of-scope source contributes no bias — routing and learning both stay inside the workspace over the otherwise GLOBAL shared graph.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/router.py
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
def route_with_memory(
    self, query: str, *, k: int = 5
) -> tuple[list[RouteRecipe], ScgMemoryBias]:
    """Rank recipes AND return the learned-memory bias map that shaped them.

    Embed → vector-search seed nodes → expand one hop along capability/route
    edges → assemble candidate recipes → rank by ``cosine(seed) + edge
    weight + memory_boost`` (still zero-LLM — the memory term is a vector
    read + a polarity-weighted sum). Returns ``([], empty_bias)`` for an
    empty graph or no match.

    The returned :class:`ScgMemoryBias` carries the per-capability anchored
    HINTS too, so the ``scg_route`` plugin tool can surface "how to call this
    right" guidance on each recipe without re-reading memory.

    Honours the ambient :class:`ScgScope`: a candidate recipe whose
    steps reach an out-of-scope source is dropped, AND a memory note anchored
    to an out-of-scope source contributes no bias — routing and learning both
    stay inside the workspace over the otherwise GLOBAL shared graph.
    """
    qvec = self.embedder.embed_query(query)
    seeds = self.store.vector_search(qvec, k=k)
    if not seeds:
        return [], ScgMemoryBias.empty()

    # source_key → recipe, so a vector-search hit (mapped to its graph
    # anchor) can be matched against the recipes reachable from it. Filter to
    # the workspace scope here so an out-of-scope pathway is never a candidate.
    recipes = {
        r.source_key: r
        for r in self.store.list_recipes()
        if ScgScope.permits_recipe_steps(r.steps)
    }
    # The learned-memory bias for this query (empty when no bridge is
    # injected or no embedding backend is configured — degrades gracefully).
    bias = self._memory_bias(qvec)
    # Pre-index inbound capability/route edges by target ONCE per route()
    # (one full edge scan), not once per seed — the inbound expansion below
    # is then an O(1) dict lookup. (PPR is the documented scale upgrade
    # behind this same signature.)
    inbound = self._inbound_index()
    best: dict[SourceKey, float] = {}
    for node_id, sim in seeds:
        node = self.store.get_node(node_id)
        if node is None:
            continue
        for key, weight in self._candidate_keys(node.source_key, inbound):
            if key not in recipes:
                continue
            # cosine(seed) + edge weight + learned-memory boost (the memory-bias term;
            # 0.0 when the pathway's steps carry no learned signal).
            score = sim + weight + bias.boost_for_steps(recipes[key].steps)
            if score > best.get(key, float("-inf")):
                best[key] = score

    ranked = sorted(
        best.items(),
        # Score desc, then source_key asc — a deterministic tie-break.
        key=lambda kv: (-kv[1], kv[0]),
    )
    # Honour the documented bound: one-hop expansion can map the seeds to
    # more than *k* recipe candidates, so slice the ranked result to *k*.
    return [recipes[key] for key, _ in ranked[:k]], bias

mewbo_graph.scg.parser

ScgParser — the registry that maps sources into the persisted SCG (spec §6).

This is the parser's control layer: the per-type :class:~mewbo_graph.scg.providers.base.SourceStructureProviders do the descriptor→graph parsing; ScgParser owns persistence, embedding, and the cross-capability / cross-source joins that no single provider can see:

  • :meth:parse_source — dispatch one descriptor to the provider for its source_type, clean re-map (store.delete_source first so a re-index never accumulates stale/duplicate nodes), persist nodes/edges/recipes + the descriptor, then embed every node and upsert an :class:ScgEmbedding.
  • :meth:link_sources — run the injected :class:TypeAligner to deposit RESOLVES_TO hypothesis edges across sources (no-op without an aligner).
  • :meth:compute_param_edges — the In-N-Out (2509.01560) producer→consumer join: match a capability's PRODUCES output field to another capability's input binding by field name and emit a CONSUMES edge carrying binds=(out_key, in_key) so traversal can chain ops into qualified paths.

The embedder is the wiki :class:~mewbo_graph.wiki.embedder.Embedder (constructed via make_embedder() by default; tests inject a fake). Embedding is best-effort — a missing/failed embedding backend degrades to a structure-only SCG, never a hard failure (mirrors the wiki's BM25-fallback stance).

Security invariant (spec §6): nodes carry only redacted descriptors; this class copies no token/credential/data — it persists exactly what the providers emit.

ScgParser

Maps mapped sources into the persisted Source Capability Graph.

Dependency-injected with the :class:ScgStore to persist into, the list of :class:SourceStructureProviders to dispatch over (built into an internal :class:StructureProviderRegistry), a node embedder (the wiki :class:Embedder by default), and an optional :class:TypeAligner for the cross-source RESOLVES_TO pass. Holds no per-source state — every method operates over the injected store, so one parser instance maps a whole catalog.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/parser.py
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
class ScgParser:
    """Maps mapped sources into the persisted Source Capability Graph.

    Dependency-injected with the :class:`ScgStore` to persist into, the list of
    :class:`SourceStructureProvider`s to dispatch over (built into an internal
    :class:`StructureProviderRegistry`), a node embedder (the wiki
    :class:`Embedder` by default), and an optional :class:`TypeAligner` for the
    cross-source ``RESOLVES_TO`` pass. Holds no per-source state — every method
    operates over the injected store, so one parser instance maps a whole
    catalog.
    """

    def __init__(
        self,
        *,
        store: ScgStore,
        providers: list[SourceStructureProvider],
        embedder: _NodeEmbedder | None = None,
        aligner: TypeAligner | None = None,
    ) -> None:
        """Bind the store, providers (→ registry), embedder, and aligner."""
        self._store = store
        self._registry = StructureProviderRegistry(providers)
        self._embedder = embedder if embedder is not None else make_embedder()
        self._aligner = aligner

    # -- map one source -----------------------------------------------------

    def parse_source(self, descriptor: SourceDescriptor) -> StructureGraph:
        """Map one source into the SCG and return its parsed structure graph.

        Clean re-map: every prior node/edge/recipe/embedding for this source is
        deleted first, so re-indexing replaces rather than accumulates. The
        descriptor is persisted (with its tool-list :class:`ManifestHash` stamped
        on ``schema_version``) so later ``link_sources`` / re-maps can find it and
        the workspace-save drift check can compare the live surface against the
        mapped one.
        """
        graph = self._registry.build(descriptor)
        graph.recipes.extend(self._default_recipes(graph))

        # Stamp the manifest fingerprint so a later live tool-list hash can detect
        # drift without re-introspecting the mapped graph. Idempotent: re-mapping
        # the SAME surface re-stamps the SAME hash (the descriptor is a value).
        stamped = descriptor.model_copy(
            update={"schema_version": ManifestHash.of_descriptor_raw(descriptor.raw)}
        )

        self._store.delete_source(descriptor.source_id)
        self._store.upsert_nodes(graph.nodes)
        self._store.upsert_edges(graph.edges)
        self._store.upsert_recipes(graph.recipes)
        self._store.upsert_source(stamped)
        self._embed_nodes(graph.nodes)
        return graph

    @staticmethod
    def _default_recipes(graph: StructureGraph) -> list[RouteRecipe]:
        """Single-step recipes for capabilities the provider left recipe-less.

        ``ScgRouter.route`` only ranks persisted recipes, so a capability
        without one is unroutable — a graph with zero recipes routes nothing.
        Every capability is trivially a one-step pathway to itself; providers
        that emit richer multi-capability recipes stay authoritative (their
        keys are skipped here).
        """
        covered = {r.source_key for r in graph.recipes}
        return [
            RouteRecipe(source_key=node.source_key, steps=[node.source_key])
            for node in graph.nodes
            if node.kind == "capability" and node.source_key not in covered
        ]

    # -- cross-source RESOLVES_TO ------------------------------------------

    def link_sources(self, source_ids: list[str]) -> list[ScgEdge]:
        """Run the injected aligner across *source_ids*, persisting RESOLVES_TO.

        Returns the emitted edges (already upserted by the aligner). Without an
        aligner injected this is a deterministic no-op (``[]``), never a raise.
        """
        if self._aligner is None:
            return []
        return self._aligner.align(source_ids)

    # -- In-N-Out producer → consumer --------------------------------------

    def compute_param_edges(self) -> list[ScgEdge]:
        """Wire ``CONSUMES`` edges from producing ops to consuming ops by field.

        The In-N-Out join (``2509.01560``): a capability's ``PRODUCES`` output
        field (``<cap>.<name>``) matched to *another* capability's input binding
        of the same field *name* yields a ``CONSUMES`` edge
        ``producer → consumer`` carrying ``binds=(out_key, in_key)`` — the seam
        the router chains into qualified multi-hop paths. Deterministic;
        self-edges are skipped. Returns (and persists) the edges emitted.
        """
        produced = self._produced_fields()
        consumed = self._consumed_fields()
        edges: list[ScgEdge] = []
        seen: set[tuple[SourceKey, SourceKey, SourceKey, SourceKey]] = set()
        for name, producers in produced.items():
            for in_cap, in_key in consumed.get(name, []):
                for out_cap, out_key in producers:
                    if out_cap == in_cap:
                        continue  # an op never consumes its own output
                    dedup = (out_cap, in_cap, out_key, in_key)
                    if dedup in seen:
                        continue
                    seen.add(dedup)
                    edges.append(
                        ScgEdge(
                            source=out_cap,
                            target=in_cap,
                            kind="CONSUMES",
                            binds=(out_key, in_key),
                            method="type_align",
                            evidence=[f"field_name={name}"],
                        )
                    )
        if edges:
            self._store.upsert_edges(edges)
        return edges

    # -- embedding ----------------------------------------------------------

    def _embed_nodes(self, nodes: list[ScgNode]) -> None:
        """Embed *nodes* and upsert one :class:`ScgEmbedding` each (best-effort).

        Failure is non-fatal: a missing/erroring embedding backend leaves a
        structure-only SCG (mirrors the wiki's BM25 fallback). The embed text
        blends the node name, doc, and example queries — the retrievable surface.
        """
        if not nodes:
            return
        items = [(n.node_id, self._embed_text(n)) for n in nodes]
        try:
            rows = self._embedder.embed_nodes(items)
        except Exception as exc:  # noqa: BLE001 — embedding is best-effort
            logging.warning("SCG node embedding skipped: {}", exc)
            return
        model = getattr(self._embedder, "model", "")
        self._store.upsert_embeddings(
            [
                ScgEmbedding(
                    node_id=row.node_id,
                    vector=list(row.vector),
                    model=model,
                    dim=row.dim,
                )
                for row in rows
            ]
        )

    @staticmethod
    def _embed_text(node: ScgNode) -> str:
        """The retrievable text for a node: name + doc + example queries."""
        parts = [node.name]
        if node.doc:
            parts.append(node.doc)
        parts.extend(node.example_queries)
        return "\n".join(parts)

    # -- field indexing (In-N-Out helpers) ---------------------------------

    def _produced_fields(self) -> dict[str, list[tuple[SourceKey, SourceKey]]]:
        """``field_name -> [(capability_key, produced_field_key)]`` from PRODUCES."""
        out: dict[str, list[tuple[SourceKey, SourceKey]]] = {}
        for edge in self._store.list_edges(kind="PRODUCES"):
            name = field_leaf(edge.target)
            out.setdefault(name, []).append((edge.source, edge.target))
        return out

    def _consumed_fields(self) -> dict[str, list[tuple[SourceKey, SourceKey]]]:
        """``field_name -> [(capability_key, input_field_key)]`` from bindings."""
        out: dict[str, list[tuple[SourceKey, SourceKey]]] = {}
        for cap in self._store.query_nodes(kind="capability"):
            for binding in cap.bindings:
                name = field_leaf(binding.field_key)
                out.setdefault(name, []).append((cap.source_key, binding.field_key))
        return out

__init__(*, store: ScgStore, providers: list[SourceStructureProvider], embedder: _NodeEmbedder | None = None, aligner: TypeAligner | None = None) -> None

Bind the store, providers (→ registry), embedder, and aligner.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/parser.py
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
def __init__(
    self,
    *,
    store: ScgStore,
    providers: list[SourceStructureProvider],
    embedder: _NodeEmbedder | None = None,
    aligner: TypeAligner | None = None,
) -> None:
    """Bind the store, providers (→ registry), embedder, and aligner."""
    self._store = store
    self._registry = StructureProviderRegistry(providers)
    self._embedder = embedder if embedder is not None else make_embedder()
    self._aligner = aligner

compute_param_edges() -> list[ScgEdge]

Wire CONSUMES edges from producing ops to consuming ops by field.

The In-N-Out join (2509.01560): a capability's PRODUCES output field (<cap>.<name>) matched to another capability's input binding of the same field name yields a CONSUMES edge producer → consumer carrying binds=(out_key, in_key) — the seam the router chains into qualified multi-hop paths. Deterministic; self-edges are skipped. Returns (and persists) the edges emitted.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/parser.py
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
def compute_param_edges(self) -> list[ScgEdge]:
    """Wire ``CONSUMES`` edges from producing ops to consuming ops by field.

    The In-N-Out join (``2509.01560``): a capability's ``PRODUCES`` output
    field (``<cap>.<name>``) matched to *another* capability's input binding
    of the same field *name* yields a ``CONSUMES`` edge
    ``producer → consumer`` carrying ``binds=(out_key, in_key)`` — the seam
    the router chains into qualified multi-hop paths. Deterministic;
    self-edges are skipped. Returns (and persists) the edges emitted.
    """
    produced = self._produced_fields()
    consumed = self._consumed_fields()
    edges: list[ScgEdge] = []
    seen: set[tuple[SourceKey, SourceKey, SourceKey, SourceKey]] = set()
    for name, producers in produced.items():
        for in_cap, in_key in consumed.get(name, []):
            for out_cap, out_key in producers:
                if out_cap == in_cap:
                    continue  # an op never consumes its own output
                dedup = (out_cap, in_cap, out_key, in_key)
                if dedup in seen:
                    continue
                seen.add(dedup)
                edges.append(
                    ScgEdge(
                        source=out_cap,
                        target=in_cap,
                        kind="CONSUMES",
                        binds=(out_key, in_key),
                        method="type_align",
                        evidence=[f"field_name={name}"],
                    )
                )
    if edges:
        self._store.upsert_edges(edges)
    return edges

Run the injected aligner across source_ids, persisting RESOLVES_TO.

Returns the emitted edges (already upserted by the aligner). Without an aligner injected this is a deterministic no-op ([]), never a raise.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/parser.py
154
155
156
157
158
159
160
161
162
def link_sources(self, source_ids: list[str]) -> list[ScgEdge]:
    """Run the injected aligner across *source_ids*, persisting RESOLVES_TO.

    Returns the emitted edges (already upserted by the aligner). Without an
    aligner injected this is a deterministic no-op (``[]``), never a raise.
    """
    if self._aligner is None:
        return []
    return self._aligner.align(source_ids)

parse_source(descriptor: SourceDescriptor) -> StructureGraph

Map one source into the SCG and return its parsed structure graph.

Clean re-map: every prior node/edge/recipe/embedding for this source is deleted first, so re-indexing replaces rather than accumulates. The descriptor is persisted (with its tool-list :class:ManifestHash stamped on schema_version) so later link_sources / re-maps can find it and the workspace-save drift check can compare the live surface against the mapped one.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/parser.py
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
def parse_source(self, descriptor: SourceDescriptor) -> StructureGraph:
    """Map one source into the SCG and return its parsed structure graph.

    Clean re-map: every prior node/edge/recipe/embedding for this source is
    deleted first, so re-indexing replaces rather than accumulates. The
    descriptor is persisted (with its tool-list :class:`ManifestHash` stamped
    on ``schema_version``) so later ``link_sources`` / re-maps can find it and
    the workspace-save drift check can compare the live surface against the
    mapped one.
    """
    graph = self._registry.build(descriptor)
    graph.recipes.extend(self._default_recipes(graph))

    # Stamp the manifest fingerprint so a later live tool-list hash can detect
    # drift without re-introspecting the mapped graph. Idempotent: re-mapping
    # the SAME surface re-stamps the SAME hash (the descriptor is a value).
    stamped = descriptor.model_copy(
        update={"schema_version": ManifestHash.of_descriptor_raw(descriptor.raw)}
    )

    self._store.delete_source(descriptor.source_id)
    self._store.upsert_nodes(graph.nodes)
    self._store.upsert_edges(graph.edges)
    self._store.upsert_recipes(graph.recipes)
    self._store.upsert_source(stamped)
    self._embed_nodes(graph.nodes)
    return graph

mewbo_graph.scg.entity_resolution

Type-level cross-source entity resolution — spec §6 (abstain-by-default).

:class:TypeAligner runs at map time: it compares entity_type nodes across the given sources and deposits durable RESOLVES_TO edges (method="type_align") for the type correspondences it is confident about — e.g. Jira.Issue <=> Linear.Ticket. The edge is a weighted, provenanced hypothesis (Graphiti validity window already on :class:ScgEdge), never an asserted truth: traversal weighs it, it is not a hard join.

Abstain by default. An edge is emitted only on positive evidence:

  • a name/field-overlap heuristic produces a similarity in [0, 1];
  • pairs at/above confident_threshold are emitted on the heuristic alone;
  • pairs in the ambiguous band (band_low..confident_threshold) are emitted only if an injected LLM affirms them (one call per band pair); with no LLM injected, band pairs abstain (NONE-default — mirrors the memory layer's dedup tier-3 stance);
  • everything below band_low abstains.

Cross-source only: same-source pairs and non-entity_type nodes are skipped.

Instance-level ER is explicitly NOT here. Resolving two concrete records ("is Jira issue #42 the same work item as Linear ticket ENG-7?") happens online, inside the probe agent, which keys-blocks and selects natively over live data. This class owns only the offline, type-level schema correspondence that scopes where the probe agent should even look — never the data behind it.

Security invariant (spec §6): operates purely over SCG structure nodes, which carry only redacted descriptors — no token, credential, or record value.

TypeAligner

Map-time, type-level cross-source entity resolver (abstain-by-default).

Dependency-injected: a :class:ScgStore to read entity_type nodes from and persist hypothesis edges into, plus an optional Callable[[str], str] that disambiguates only the ambiguous band (the confident and reject tiers never spend a token). Thresholds are read once from scg.entity_resolution config with calibrated code defaults, so the whole feature stays gated and tunable without editing this class.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/entity_resolution.py
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
class TypeAligner:
    """Map-time, type-level cross-source entity resolver (abstain-by-default).

    Dependency-injected: a :class:`ScgStore` to read ``entity_type`` nodes from
    and persist hypothesis edges into, plus an optional ``Callable[[str], str]``
    that disambiguates only the ambiguous *band* (the confident and reject tiers
    never spend a token). Thresholds are read once from ``scg.entity_resolution``
    config with calibrated code defaults, so the whole feature stays gated and
    tunable without editing this class.
    """

    _PROMPT = (
        "Two entity types from different data sources may describe the SAME "
        "real-world concept. Reply with exactly 'yes' or 'no'.\n"
        "A: {a_source}.{a_name} fields={a_fields}\n"
        "B: {b_source}.{b_name} fields={b_fields}\n"
        "Do A and B describe the same kind of entity? Answer yes or no."
    )

    def __init__(
        self, *, store: ScgStore, llm: Callable[[str], str] | None = None
    ) -> None:
        """Inject the store and (optionally) the band-disambiguation LLM."""
        self._store = store
        self._llm = llm
        # Per-``align`` memo of ``source_key -> field-name set``. ``_field_names``
        # depends only on the node, but is read 4-6× per cross-source pair (an
        # O(pairs) loop), and each miss scans the whole edge file on the JSON
        # backend. Compute once per node, reuse for every pair (spec: compute once).
        self._fields_cache: dict[SourceKey, set[str]] = {}
        self._confident = float(
            get_config_value(
                "scg",
                "entity_resolution",
                "confident_threshold",
                default=_DEFAULT_CONFIDENT_THRESHOLD,
            )
        )
        self._band_low = float(
            get_config_value(
                "scg", "entity_resolution", "band_low", default=_DEFAULT_BAND_LOW
            )
        )

    # -- public API ---------------------------------------------------------

    def align(self, source_ids: list[str]) -> list[ScgEdge]:
        """Emit durable ``RESOLVES_TO`` edges for confident type correspondences.

        Compares every cross-source pair of ``entity_type`` nodes drawn from
        *source_ids*, abstains by default, and upserts the surviving hypothesis
        edges into the store. Returns the edges emitted (empty when none clear
        the bar). Deterministic for a fixed store + injected LLM.
        """
        # Fresh per-align memo so a re-map sees the current graph, not a stale set.
        self._fields_cache = {}
        types_by_source = {sid: self._entity_types(sid) for sid in dict.fromkeys(source_ids)}
        edges: list[ScgEdge] = []
        sources = sorted(types_by_source)
        for i, left in enumerate(sources):
            for right in sources[i + 1 :]:
                for a in types_by_source[left]:
                    for b in types_by_source[right]:
                        edge = self._resolve_pair(a, b)
                        if edge is not None:
                            edges.append(edge)
        if edges:
            self._store.upsert_edges(edges)
        return edges

    # -- per-pair decision --------------------------------------------------

    def _resolve_pair(self, a: ScgNode, b: ScgNode) -> ScgEdge | None:
        """Decide one cross-source type pair; return an edge or None (abstain)."""
        score = self._similarity(a, b)
        if score >= self._confident:
            return self._edge(a, b, weight=score, evidence=self._evidence(a, b, score))
        if score >= self._band_low and self._llm_affirms(a, b):
            evidence = [*self._evidence(a, b, score), "llm: affirmed band pair"]
            # Damp the weight: an LLM-promoted band pair is a softer hypothesis
            # than a heuristic-confident one (calibrated, capped at confident).
            return self._edge(a, b, weight=min(score + 0.1, self._confident), evidence=evidence)
        return None

    def _llm_affirms(self, a: ScgNode, b: ScgNode) -> bool:
        """Ask the injected LLM once whether the band pair is the same type."""
        if self._llm is None:
            return False  # NONE-default: no LLM -> abstain on band pairs.
        prompt = self._PROMPT.format(
            a_source=a.source_id,
            a_name=a.name,
            a_fields=sorted(self._field_names(a)),
            b_source=b.source_id,
            b_name=b.name,
            b_fields=sorted(self._field_names(b)),
        )
        return self._llm(prompt).strip().lower().startswith("y")

    # -- heuristic ----------------------------------------------------------

    def _similarity(self, a: ScgNode, b: ScgNode) -> float:
        """Blend name + field-overlap into a calibrated ``[0, 1]`` similarity.

        Field overlap (Jaccard over field names) is the strong signal; an exact
        name match adds a bounded bonus. Names alone never clear the confident
        bar — schema overlap is what makes a type correspondence executable.
        """
        field_sim = self._jaccard(self._field_names(a), self._field_names(b))
        name_bonus = 0.2 if a.name.lower() == b.name.lower() else 0.0
        return min(field_sim + name_bonus, 1.0)

    @staticmethod
    def _jaccard(a: set[str], b: set[str]) -> float:
        """Jaccard similarity of two name sets (0.0 if both empty)."""
        if not a and not b:
            return 0.0
        union = a | b
        return len(a & b) / len(union) if union else 0.0

    def _field_names(self, node: ScgNode) -> set[str]:
        """The field *names* of an entity type (last ``.`` segment, lower-cased).

        Memoized per ``source_key`` for the duration of one ``align`` so the
        whole-edge-file scan below runs once per node, not once per pair.

        Freshly-parsed sources (OpenAPI / MCP) emit each field as a separate
        ``field`` node linked by a ``HAS_FIELD`` edge and leave
        ``entity_type.bindings`` empty — so the field set is derived from the
        entity type's ``HAS_FIELD`` neighbours in the graph. Falls back to
        ``bindings`` (the parser's In-N-Out producer/consumer pass populates
        them) when no ``HAS_FIELD`` edges are present, so ER fires on both the
        parsed-graph and the binding-bearing path.
        """
        cached = self._fields_cache.get(node.source_key)
        if cached is not None:
            return cached
        names = {
            field_leaf(edge.target)
            for edge in self._store.neighbors(node.source_key)
            if edge.kind == "HAS_FIELD"
        }
        if not names:
            names = {field_leaf(binding.field_key) for binding in node.bindings}
        self._fields_cache[node.source_key] = names
        return names

    def _evidence(self, a: ScgNode, b: ScgNode, score: float) -> list[str]:
        """Human-readable provenance strings for the emitted edge."""
        shared = sorted(self._field_names(a) & self._field_names(b))
        out = [f"field_overlap={score:.2f}"]
        if shared:
            out.append(f"shared_fields={shared}")
        if a.name.lower() == b.name.lower():
            out.append(f"name_match={a.name}")
        return out

    # -- node loading + edge construction -----------------------------------

    def _entity_types(self, source_id: str) -> list[ScgNode]:
        """Deterministically-ordered ``entity_type`` nodes for one source."""
        nodes = self._store.query_nodes(source_id=source_id, kind="entity_type")
        return sorted(nodes, key=lambda n: n.source_key)

    @staticmethod
    def _edge(
        a: ScgNode, b: ScgNode, *, weight: float, evidence: list[str]
    ) -> ScgEdge:
        """Build a canonically-oriented ``RESOLVES_TO`` hypothesis edge."""
        source, target = sorted((a.source_key, b.source_key))
        binds: tuple[SourceKey, SourceKey] = (source, target)
        return ScgEdge(
            source=source,
            target=target,
            kind="RESOLVES_TO",
            weight=weight,
            binds=binds,
            method="type_align",
            evidence=evidence,
        )

__init__(*, store: ScgStore, llm: Callable[[str], str] | None = None) -> None

Inject the store and (optionally) the band-disambiguation LLM.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/entity_resolution.py
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
def __init__(
    self, *, store: ScgStore, llm: Callable[[str], str] | None = None
) -> None:
    """Inject the store and (optionally) the band-disambiguation LLM."""
    self._store = store
    self._llm = llm
    # Per-``align`` memo of ``source_key -> field-name set``. ``_field_names``
    # depends only on the node, but is read 4-6× per cross-source pair (an
    # O(pairs) loop), and each miss scans the whole edge file on the JSON
    # backend. Compute once per node, reuse for every pair (spec: compute once).
    self._fields_cache: dict[SourceKey, set[str]] = {}
    self._confident = float(
        get_config_value(
            "scg",
            "entity_resolution",
            "confident_threshold",
            default=_DEFAULT_CONFIDENT_THRESHOLD,
        )
    )
    self._band_low = float(
        get_config_value(
            "scg", "entity_resolution", "band_low", default=_DEFAULT_BAND_LOW
        )
    )

align(source_ids: list[str]) -> list[ScgEdge]

Emit durable RESOLVES_TO edges for confident type correspondences.

Compares every cross-source pair of entity_type nodes drawn from source_ids, abstains by default, and upserts the surviving hypothesis edges into the store. Returns the edges emitted (empty when none clear the bar). Deterministic for a fixed store + injected LLM.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/entity_resolution.py
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
def align(self, source_ids: list[str]) -> list[ScgEdge]:
    """Emit durable ``RESOLVES_TO`` edges for confident type correspondences.

    Compares every cross-source pair of ``entity_type`` nodes drawn from
    *source_ids*, abstains by default, and upserts the surviving hypothesis
    edges into the store. Returns the edges emitted (empty when none clear
    the bar). Deterministic for a fixed store + injected LLM.
    """
    # Fresh per-align memo so a re-map sees the current graph, not a stale set.
    self._fields_cache = {}
    types_by_source = {sid: self._entity_types(sid) for sid in dict.fromkeys(source_ids)}
    edges: list[ScgEdge] = []
    sources = sorted(types_by_source)
    for i, left in enumerate(sources):
        for right in sources[i + 1 :]:
            for a in types_by_source[left]:
                for b in types_by_source[right]:
                    edge = self._resolve_pair(a, b)
                    if edge is not None:
                        edges.append(edge)
    if edges:
        self._store.upsert_edges(edges)
    return edges

mewbo_graph.scg.memory_bridge

ScgMemoryBridge — the learned-layer flywheel over the SCG.

The SCG structure (schemas + pathways) is search-owned; the learned layer is shared with the wiki layer's memory substrate — there is ZERO re-implementation of the atomic-note / anchor / dedup machinery here. This module is a thin seam that:

  • lets the wiki layer's :class:~mewbo_graph.wiki.memory.InsightIngestor resolve connector anchors against the SCG instead of the code graph (:class:ScgAnchorResolver), and
  • deposits / retrieves connector insights under corpus="connector" (:class:ScgMemoryBridge).

Why the resolver is correctness-critical: memory_vector_search defaults to exclude_invalidated=True, which only returns notes that have a live ANCHORS edge. The ingestor creates that edge only for anchors its StructureProvider can resolve. The default CodeStructureProvider resolves file#Name code keys — it can never resolve a connector source_key, so a connector insight would be written but then silently dropped on read. ScgAnchorResolver resolves source_key → :class:ScgNode, so the edge is created and the insight surfaces.

Retrieval goes straight through the store's memory_vector_search ANN seam (NOT MultiplexExpander): the expander's code-graph neighbour expansion no-ops for connectors — they have no tree-sitter CALLS/IMPORTS edges to walk.

ScgAnchorResolver

StructureProvider backed by the SCG store (source_key → node).

Implements the wiki layer's StructureProvider Protocol (resolve / resolve_many / entity_key_of) so the shared InsightIngestor can resolve connector source_key anchors instead of dropping them. Stateless beyond the injected store — a re-map mutates the graph, so caching would go stale (mirrors CodeStructureProvider).

The multiplex entity_key for a connector is simply its source_key (<source_id>#<Qualified.Name>); SCG nodes already carry no byte offsets, so anchors survive a re-index.

A source_key is resolved KIND-AGNOSTICALLY across _ANCHORABLE_KINDS because node_id is content-addressed per (source_key, kind): an MCP tool-list source maps each tool to a capability node while an OpenAPI source exposes entity_type nodes, so a fixed-kind probe drops every anchor of the other shape — which is exactly the bug that left connector insights edge-less (and thus invisible to memory_vector_search).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
class ScgAnchorResolver:
    """``StructureProvider`` backed by the SCG store (``source_key`` → node).

    Implements the wiki layer's ``StructureProvider`` Protocol (``resolve`` /
    ``resolve_many`` / ``entity_key_of``) so the shared ``InsightIngestor`` can
    resolve connector ``source_key`` anchors instead of dropping them.
    Stateless beyond the injected store — a re-map mutates the
    graph, so caching would go stale (mirrors ``CodeStructureProvider``).

    The multiplex ``entity_key`` for a connector is simply its ``source_key``
    (``<source_id>#<Qualified.Name>``); SCG nodes already carry no byte offsets,
    so anchors survive a re-index.

    A ``source_key`` is resolved KIND-AGNOSTICALLY across ``_ANCHORABLE_KINDS``
    because ``node_id`` is content-addressed per ``(source_key, kind)``: an MCP
    tool-list source maps each tool to a ``capability`` node while an OpenAPI
    source exposes ``entity_type`` nodes, so a fixed-kind probe drops every
    anchor of the other shape — which is exactly the bug that left connector
    insights edge-less (and thus invisible to ``memory_vector_search``).
    """

    def __init__(self, store: ScgStore) -> None:
        """Compose over an SCG store (dependency-injected)."""
        self._store = store

    def resolve(self, slug: str, entity_key: EntityKey) -> ScgNode | None:
        """Return the SCG node addressed by ``entity_key`` (a ``source_key``).

        Probes each anchorable kind in priority order (``capability`` first, the
        MCP-tool-list shape; then ``entity_type``, the OpenAPI shape) and returns
        the first live node — so a connector ``source_key`` resolves regardless
        of which structure the source's provider emitted.
        """
        return self._lookup(entity_key)

    def resolve_many(
        self, slug: str, entity_keys: list[EntityKey]
    ) -> dict[EntityKey, ScgNode]:
        """Resolve a batch by ``source_key``; misses are omitted from the result."""
        out: dict[EntityKey, ScgNode] = {}
        for key in entity_keys:
            if key in out:
                continue
            node = self.resolve(slug, key)
            if node is not None:
                out[key] = node
        return out

    def entity_key_of(self, slug: str, node_id: str) -> EntityKey | None:
        """Return the ``source_key`` (== ``entity_key``) for an SCG ``node_id``."""
        node = self._store.get_node(node_id)
        return node.source_key if node is not None else None

    def _lookup(self, source_key: EntityKey) -> ScgNode | None:
        """Return the live SCG node for ``source_key`` across anchorable kinds.

        ``node_id`` is content-addressed per ``(source_key, kind)``, so this
        probes ``_ANCHORABLE_KINDS`` in priority order and returns the first
        node that exists — a ``capability`` (MCP tool list) OR ``entity_type``
        (OpenAPI) source resolves identically. ``None`` only when no anchorable
        node carries this ``source_key`` (a genuinely unresolvable anchor).
        """
        from .types import ScgNode as _ScgNode

        for kind in _ANCHORABLE_KINDS:
            node = self._store.get_node(_ScgNode.make_id(source_key, kind))
            if node is not None:
                return node
        return None

__init__(store: ScgStore) -> None

Compose over an SCG store (dependency-injected).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
115
116
117
def __init__(self, store: ScgStore) -> None:
    """Compose over an SCG store (dependency-injected)."""
    self._store = store

entity_key_of(slug: str, node_id: str) -> EntityKey | None

Return the source_key (== entity_key) for an SCG node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
142
143
144
145
def entity_key_of(self, slug: str, node_id: str) -> EntityKey | None:
    """Return the ``source_key`` (== ``entity_key``) for an SCG ``node_id``."""
    node = self._store.get_node(node_id)
    return node.source_key if node is not None else None

resolve(slug: str, entity_key: EntityKey) -> ScgNode | None

Return the SCG node addressed by entity_key (a source_key).

Probes each anchorable kind in priority order (capability first, the MCP-tool-list shape; then entity_type, the OpenAPI shape) and returns the first live node — so a connector source_key resolves regardless of which structure the source's provider emitted.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
119
120
121
122
123
124
125
126
127
def resolve(self, slug: str, entity_key: EntityKey) -> ScgNode | None:
    """Return the SCG node addressed by ``entity_key`` (a ``source_key``).

    Probes each anchorable kind in priority order (``capability`` first, the
    MCP-tool-list shape; then ``entity_type``, the OpenAPI shape) and returns
    the first live node — so a connector ``source_key`` resolves regardless
    of which structure the source's provider emitted.
    """
    return self._lookup(entity_key)

resolve_many(slug: str, entity_keys: list[EntityKey]) -> dict[EntityKey, ScgNode]

Resolve a batch by source_key; misses are omitted from the result.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
129
130
131
132
133
134
135
136
137
138
139
140
def resolve_many(
    self, slug: str, entity_keys: list[EntityKey]
) -> dict[EntityKey, ScgNode]:
    """Resolve a batch by ``source_key``; misses are omitted from the result."""
    out: dict[EntityKey, ScgNode] = {}
    for key in entity_keys:
        if key in out:
            continue
        node = self.resolve(slug, key)
        if node is not None:
            out[key] = node
    return out

ScgMemoryBridge

Deposit / retrieve connector insights over the wiki layer's memory substrate.

The learned-layer flywheel for Agentic Search: data-location wins, failure constraints, resolved bindings and learned edge weights are written as atomic connector notes and read back to bias traversal. All atomic-note / dedup / anchor work is the shared InsightIngestor — this class only pins corpus="connector" and swaps in the SCG-backed anchor resolver.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
class ScgMemoryBridge:
    """Deposit / retrieve connector insights over the wiki layer's memory substrate.

    The learned-layer flywheel for Agentic Search: data-location wins, failure
    constraints, resolved bindings and learned edge weights are written as
    atomic connector notes and read back to bias traversal. All atomic-note /
    dedup / anchor work is the shared ``InsightIngestor`` — this class only
    pins ``corpus="connector"`` and swaps in the SCG-backed anchor resolver.
    """

    def __init__(
        self,
        *,
        wiki_store: WikiStoreBase,
        embedder: object,
        llm: object | None = None,
    ) -> None:
        """Wire collaborators (all injected); ``llm`` is opt-in (dedup tier-3)."""
        self._store = wiki_store
        self._embedder = embedder
        self._llm = llm
        # The resolver is overridable so a caller can point it at a specific
        # SCG store; default constructs against the process-wide singleton lazily.
        self._resolver: ScgAnchorResolver | None = None

    @property
    def resolver(self) -> ScgAnchorResolver:
        """The anchor resolver; lazily bound to the process-wide SCG store."""
        if self._resolver is None:
            from .store import get_scg_store

            self._resolver = ScgAnchorResolver(get_scg_store())
        return self._resolver

    @resolver.setter
    def resolver(self, resolver: ScgAnchorResolver) -> None:
        """Override the anchor resolver (e.g. to target a specific SCG store)."""
        self._resolver = resolver

    def write_insight(
        self,
        slug: str,
        content: str,
        *,
        source_keys: list[str],
        kind: MemoryKind = "propositional",
        labels: list[str] | None = None,
        polarity: Polarity = _DEFAULT_POLARITY,
        workspace: str | None = None,
    ) -> IngestResult:
        """Deposit one connector insight anchored to ``source_keys``.

        Routes through the shared ``InsightIngestor`` with ``corpus="connector"``
        and the SCG-backed anchor resolver, so resolvable anchors create the live
        ``ANCHORS`` edge that retrieval requires. The resolver is passed at
        construction (``provider=``) — connector source_keys resolve instead of
        being dropped, with no post-construction mutation of the ingestor.

        ``polarity`` records whether the fact is positive evidence or a dead end;
        it rides a reserved ``scg:<polarity>`` label so memory-aware routing can
        boost / damp the anchored capability. ``workspace`` (if given —
        ambient from :class:`ScgScope` at the call site) rides a reserved
        ``ws:<id>`` label for **attribution only**, NEVER a partition: the shared
        graph cross-pollinates, so a cross-workspace read still surfaces the note.
        """
        tags = list(labels or [])
        tags.append(polarity_label(polarity))
        if workspace:
            tags.append(f"ws:{workspace}")
        ingestor = InsightIngestor.from_store(
            self._store,
            embedder=self._embedder,
            llm=self._llm,
            provider=self.resolver,
        )
        return ingestor.ingest(
            slug,
            content,
            anchors=list(source_keys),
            corpus=CONNECTOR_CORPUS,
            kind=kind,
            labels=tags,
        )

    def read_insights(
        self, slug: str, query_vec: list[float], *, k: int = 10
    ) -> list[MemoryNode]:
        """Return the top-``k`` connector insights for ``query_vec``.

        Reads through the store's ``memory_vector_search`` ANN seam filtered to
        ``corpus="connector"`` (NOT ``MultiplexExpander`` — its code-graph
        neighbour expansion no-ops for connectors). Embeddings are resolved back
        to their nodes in rank order; the corpus filter already excludes other
        corpora, so the node lookup never returns a non-connector note.
        """
        filt = MemoryFilter(corpus=CONNECTOR_CORPUS)
        out: list[MemoryNode] = []
        for emb in self._store.memory_vector_search(slug, query_vec, k, filt=filt):
            node = self._store.get_memory_node(slug, emb.node_id)
            if node is not None:
                out.append(node)
        return out

    def read_anchored_insights(
        self, slug: str, query_vec: list[float], *, k: int = 10
    ) -> list[tuple[MemoryNode, float, list[SourceKey]]]:
        """Top-``k`` connector insights, each with its cosine score + live anchors.

        The retrieval surface memory-aware routing consumes: a note is
        useless to routing without knowing WHICH capability ``source_key`` it
        hangs off, so this returns ``(note, score, anchored_source_keys)`` in one
        pass — the live ``ANCHORS`` edge targets per note are read off the store's
        edge index (``list_memory_edges`` already excludes invalidated edges).
        The score is the same brute-force cosine the router uses, so the bias is
        commensurate with the seed similarity it blends into.
        """
        filt = MemoryFilter(corpus=CONNECTOR_CORPUS)
        out: list[tuple[MemoryNode, float, list[SourceKey]]] = []
        for emb in self._store.memory_vector_search(slug, query_vec, k, filt=filt):
            node = self._store.get_memory_node(slug, emb.node_id)
            if node is None:
                continue
            anchors = [
                e.target
                for e in self._store.list_memory_edges(slug, node_id=node.node_id)
                if e.type == "ANCHORS"
            ]
            out.append((node, Embedder.cosine(query_vec, emb.vector), anchors))
        return out

resolver: ScgAnchorResolver property writable

The anchor resolver; lazily bound to the process-wide SCG store.

__init__(*, wiki_store: WikiStoreBase, embedder: object, llm: object | None = None) -> None

Wire collaborators (all injected); llm is opt-in (dedup tier-3).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
175
176
177
178
179
180
181
182
183
184
185
186
187
188
def __init__(
    self,
    *,
    wiki_store: WikiStoreBase,
    embedder: object,
    llm: object | None = None,
) -> None:
    """Wire collaborators (all injected); ``llm`` is opt-in (dedup tier-3)."""
    self._store = wiki_store
    self._embedder = embedder
    self._llm = llm
    # The resolver is overridable so a caller can point it at a specific
    # SCG store; default constructs against the process-wide singleton lazily.
    self._resolver: ScgAnchorResolver | None = None

read_anchored_insights(slug: str, query_vec: list[float], *, k: int = 10) -> list[tuple[MemoryNode, float, list[SourceKey]]]

Top-k connector insights, each with its cosine score + live anchors.

The retrieval surface memory-aware routing consumes: a note is useless to routing without knowing WHICH capability source_key it hangs off, so this returns (note, score, anchored_source_keys) in one pass — the live ANCHORS edge targets per note are read off the store's edge index (list_memory_edges already excludes invalidated edges). The score is the same brute-force cosine the router uses, so the bias is commensurate with the seed similarity it blends into.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
def read_anchored_insights(
    self, slug: str, query_vec: list[float], *, k: int = 10
) -> list[tuple[MemoryNode, float, list[SourceKey]]]:
    """Top-``k`` connector insights, each with its cosine score + live anchors.

    The retrieval surface memory-aware routing consumes: a note is
    useless to routing without knowing WHICH capability ``source_key`` it
    hangs off, so this returns ``(note, score, anchored_source_keys)`` in one
    pass — the live ``ANCHORS`` edge targets per note are read off the store's
    edge index (``list_memory_edges`` already excludes invalidated edges).
    The score is the same brute-force cosine the router uses, so the bias is
    commensurate with the seed similarity it blends into.
    """
    filt = MemoryFilter(corpus=CONNECTOR_CORPUS)
    out: list[tuple[MemoryNode, float, list[SourceKey]]] = []
    for emb in self._store.memory_vector_search(slug, query_vec, k, filt=filt):
        node = self._store.get_memory_node(slug, emb.node_id)
        if node is None:
            continue
        anchors = [
            e.target
            for e in self._store.list_memory_edges(slug, node_id=node.node_id)
            if e.type == "ANCHORS"
        ]
        out.append((node, Embedder.cosine(query_vec, emb.vector), anchors))
    return out

read_insights(slug: str, query_vec: list[float], *, k: int = 10) -> list[MemoryNode]

Return the top-k connector insights for query_vec.

Reads through the store's memory_vector_search ANN seam filtered to corpus="connector" (NOT MultiplexExpander — its code-graph neighbour expansion no-ops for connectors). Embeddings are resolved back to their nodes in rank order; the corpus filter already excludes other corpora, so the node lookup never returns a non-connector note.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
def read_insights(
    self, slug: str, query_vec: list[float], *, k: int = 10
) -> list[MemoryNode]:
    """Return the top-``k`` connector insights for ``query_vec``.

    Reads through the store's ``memory_vector_search`` ANN seam filtered to
    ``corpus="connector"`` (NOT ``MultiplexExpander`` — its code-graph
    neighbour expansion no-ops for connectors). Embeddings are resolved back
    to their nodes in rank order; the corpus filter already excludes other
    corpora, so the node lookup never returns a non-connector note.
    """
    filt = MemoryFilter(corpus=CONNECTOR_CORPUS)
    out: list[MemoryNode] = []
    for emb in self._store.memory_vector_search(slug, query_vec, k, filt=filt):
        node = self._store.get_memory_node(slug, emb.node_id)
        if node is not None:
            out.append(node)
    return out

write_insight(slug: str, content: str, *, source_keys: list[str], kind: MemoryKind = 'propositional', labels: list[str] | None = None, polarity: Polarity = _DEFAULT_POLARITY, workspace: str | None = None) -> IngestResult

Deposit one connector insight anchored to source_keys.

Routes through the shared InsightIngestor with corpus="connector" and the SCG-backed anchor resolver, so resolvable anchors create the live ANCHORS edge that retrieval requires. The resolver is passed at construction (provider=) — connector source_keys resolve instead of being dropped, with no post-construction mutation of the ingestor.

polarity records whether the fact is positive evidence or a dead end; it rides a reserved scg:<polarity> label so memory-aware routing can boost / damp the anchored capability. workspace (if given — ambient from :class:ScgScope at the call site) rides a reserved ws:<id> label for attribution only, NEVER a partition: the shared graph cross-pollinates, so a cross-workspace read still surfaces the note.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
def write_insight(
    self,
    slug: str,
    content: str,
    *,
    source_keys: list[str],
    kind: MemoryKind = "propositional",
    labels: list[str] | None = None,
    polarity: Polarity = _DEFAULT_POLARITY,
    workspace: str | None = None,
) -> IngestResult:
    """Deposit one connector insight anchored to ``source_keys``.

    Routes through the shared ``InsightIngestor`` with ``corpus="connector"``
    and the SCG-backed anchor resolver, so resolvable anchors create the live
    ``ANCHORS`` edge that retrieval requires. The resolver is passed at
    construction (``provider=``) — connector source_keys resolve instead of
    being dropped, with no post-construction mutation of the ingestor.

    ``polarity`` records whether the fact is positive evidence or a dead end;
    it rides a reserved ``scg:<polarity>`` label so memory-aware routing can
    boost / damp the anchored capability. ``workspace`` (if given —
    ambient from :class:`ScgScope` at the call site) rides a reserved
    ``ws:<id>`` label for **attribution only**, NEVER a partition: the shared
    graph cross-pollinates, so a cross-workspace read still surfaces the note.
    """
    tags = list(labels or [])
    tags.append(polarity_label(polarity))
    if workspace:
        tags.append(f"ws:{workspace}")
    ingestor = InsightIngestor.from_store(
        self._store,
        embedder=self._embedder,
        llm=self._llm,
        provider=self.resolver,
    )
    return ingestor.ingest(
        slug,
        content,
        anchors=list(source_keys),
        corpus=CONNECTOR_CORPUS,
        kind=kind,
        labels=tags,
    )

polarity_label(polarity: Polarity) -> str

The reserved label encoding polarity (the one canonical mapping).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
65
66
67
def polarity_label(polarity: Polarity) -> str:
    """The reserved label encoding *polarity* (the one canonical mapping)."""
    return f"{_POLARITY_PREFIX}{polarity}"

polarity_of(node: MemoryNode) -> Polarity

Read a note's polarity off its labels; default positive if unlabelled.

A dead-end label damps routing; anything else, an untagged note included, reads as positive evidence — the conservative default, so an untagged corpus keeps biasing toward known-good pathways with no backfill.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/memory_bridge.py
70
71
72
73
74
75
76
77
78
79
def polarity_of(node: MemoryNode) -> Polarity:
    """Read a note's polarity off its labels; default ``positive`` if unlabelled.

    A dead-end label damps routing; anything else, an untagged note included,
    reads as positive evidence — the conservative default, so an untagged corpus
    keeps biasing toward known-good pathways with no backfill.
    """
    if polarity_label("dead_end") in node.labels:
        return "dead_end"
    return "positive"

mewbo_graph.scg.store

Persistence for the Source Capability Graph (SCG) — JSON or MongoDB.

The SCG structure store is search-owned and deliberately SEPARATE from the run store (agentic_search_runs): a re-map of a source rewrites graph nodes without touching any in-flight run. It mirrors the project/wiki/run dual-backend pattern: an abstract base + a filesystem driver + a Mongo driver + a config-driven factory + a process-wide singleton.

Five entity families, each in its own storage namespace:

  • nodes — :class:ScgNode, keyed on the derived node_id.
  • edges — :class:ScgEdge, keyed on the (source, target, kind) triple.
  • recipes — :class:RouteRecipe, keyed on source_key.
  • embeddings — :class:ScgEmbedding, keyed on node_id.
  • sources — :class:SourceDescriptor, keyed on source_id.

JSON layout under <cache_dir>/agentic_search/scg/ — one file per collection holding a {key: doc} map (small graphs; whole-file rewrite under a lock)::

nodes.json
edges.json
recipes.json
embeddings.json
sources.json

Mongo collections: agentic_search_scg_nodes, agentic_search_scg_edges, agentic_search_scg_recipes, agentic_search_scg_embeddings, agentic_search_scg_sources.

Per-source mappings are GLOBAL and content-addressed (node_id = sha1(source_key|kind)[:16]) — the SCG is "a tenant of the same three-layer multiplex graph that powers the Agentic Wiki" and the layers cross-pollinate without explicit wiring (docs/features-search.md). That scope therefore does NOT hard-partition this store by workspace; a workspace is a scoped VIEW over the shared graph — see :mod:mewbo_graph.scg.scope (the source-id allowlist :class:ScgRouter honours at query time) — so a re-map in one workspace stays a cheap idempotent upsert that every workspace mapping that source benefits from.

Security invariant (spec §6): SCG nodes carry only a redacted auth_scope descriptor — this store never sees or persists a token/credential.

JsonScgStore

Bases: ScgStore

Filesystem-backed SCG store under <cache_dir>/agentic_search/scg/.

Each collection is one JSON file holding a {natural_key: doc} map. All mutations take _lock and rewrite the whole file — SCGs are small enough that this is simpler and safer than partial writes (single-instance use; the Mongo driver is the multi-worker path).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
class JsonScgStore(ScgStore):
    """Filesystem-backed SCG store under ``<cache_dir>/agentic_search/scg/``.

    Each collection is one JSON file holding a ``{natural_key: doc}`` map. All
    mutations take ``_lock`` and rewrite the whole file — SCGs are small enough
    that this is simpler and safer than partial writes (single-instance use;
    the Mongo driver is the multi-worker path).
    """

    _NODES = "nodes"
    _EDGES = "edges"
    _RECIPES = "recipes"
    _EMBEDDINGS = "embeddings"
    _SOURCES = "sources"

    def __init__(self, root_dir: str | Path | None = None) -> None:
        """Initialise + create the directory tree."""
        super().__init__()
        if root_dir is None:
            home = get_config_value("runtime", "cache_dir", default="") or ".mewbo"
            root_dir = Path(home) / "agentic_search" / "scg"
        self.root_dir = Path(root_dir)
        self.root_dir.mkdir(parents=True, exist_ok=True)
        self._lock = threading.Lock()

    # -- helpers ------------------------------------------------------------

    def _path(self, collection: str) -> Path:
        return self.root_dir / f"{collection}.json"

    def _load(self, collection: str) -> dict[str, dict[str, object]]:
        path = self._path(collection)
        if not path.exists():
            return {}
        try:
            data = json.loads(path.read_text(encoding="utf-8"))
        except (json.JSONDecodeError, OSError):
            logging.warning("Skipping malformed SCG collection at {}", path)
            return {}
        return data if isinstance(data, dict) else {}

    def _save(self, collection: str, data: dict[str, dict[str, object]]) -> None:
        self._path(collection).write_text(
            json.dumps(data, indent=2), encoding="utf-8"
        )

    def _upsert(
        self, collection: str, items: list[tuple[str, dict[str, object]]]
    ) -> None:
        if not items:
            return
        with self._lock:
            data = self._load(collection)
            for key, doc in items:
                data[key] = doc
            self._save(collection, data)

    # -- Writes -------------------------------------------------------------

    def upsert_nodes(self, nodes: list[ScgNode]) -> None:
        """Upsert nodes, keyed on ``node_id``."""
        self._upsert(
            self._NODES, [(n.node_id, n.model_dump(mode="json")) for n in nodes]
        )
        self._invalidate_nodes()

    def upsert_edges(self, edges: list[ScgEdge]) -> None:
        """Upsert edges, keyed on the ``(source, target, kind)`` triple."""
        self._upsert(
            self._EDGES, [(_edge_key(e), e.model_dump(mode="json")) for e in edges]
        )

    def upsert_recipes(self, recipes: list[RouteRecipe]) -> None:
        """Upsert route recipes, keyed on ``source_key``."""
        self._upsert(
            self._RECIPES,
            [(r.source_key, r.model_dump(mode="json")) for r in recipes],
        )

    def upsert_embeddings(self, embeddings: list[ScgEmbedding]) -> None:
        """Upsert embeddings, keyed on ``node_id``."""
        self._upsert(
            self._EMBEDDINGS,
            [(e.node_id, e.model_dump(mode="json")) for e in embeddings],
        )

    def upsert_source(self, descriptor: SourceDescriptor) -> None:
        """Upsert a source descriptor, keyed on ``source_id``."""
        self._upsert(
            self._SOURCES,
            [(descriptor.source_id, descriptor.model_dump(mode="json"))],
        )

    # -- Reads --------------------------------------------------------------

    def get_node(self, node_id: str) -> ScgNode | None:
        """Return one node by id, or None if absent."""
        with self._lock:
            doc = self._load(self._NODES).get(node_id)
        return ScgNode.model_validate(doc) if doc is not None else None

    def _query_nodes_uncached(
        self,
        *,
        source_id: str | None = None,
        kind: str | None = None,
        name_contains: str | None = None,
    ) -> list[ScgNode]:
        """Scan ``nodes.json`` for the filter (AND-composed); no caching."""
        with self._lock:
            docs = list(self._load(self._NODES).values())
        needle = name_contains.lower() if name_contains else None
        out: list[ScgNode] = []
        for doc in docs:
            node = ScgNode.model_validate(doc)
            if source_id is not None and node.source_id != source_id:
                continue
            if kind is not None and node.kind != kind:
                continue
            if needle is not None and needle not in node.name.lower():
                continue
            out.append(node)
        return out

    def list_edges(
        self, *, source: SourceKey | None = None, kind: str | None = None
    ) -> list[ScgEdge]:
        """Return edges matching every supplied filter (AND-composed)."""
        with self._lock:
            docs = list(self._load(self._EDGES).values())
        out: list[ScgEdge] = []
        for doc in docs:
            edge = ScgEdge.model_validate(doc)
            if source is not None and edge.source != source:
                continue
            if kind is not None and edge.kind != kind:
                continue
            out.append(edge)
        return out

    def neighbors(self, source_key: SourceKey) -> list[ScgEdge]:
        """Return the outgoing edges whose ``source`` is *source_key*."""
        return self.list_edges(source=source_key)

    def list_recipes(self, *, source_id: str | None = None) -> list[RouteRecipe]:
        """Return route recipes, optionally scoped to one *source_id*."""
        with self._lock:
            docs = list(self._load(self._RECIPES).values())
        prefix = f"{source_id}#" if source_id is not None else None
        out: list[RouteRecipe] = []
        for doc in docs:
            recipe = RouteRecipe.model_validate(doc)
            if prefix is not None and not recipe.source_key.startswith(prefix):
                continue
            out.append(recipe)
        return out

    def list_embeddings(self) -> list[ScgEmbedding]:
        """Return all embeddings."""
        with self._lock:
            docs = list(self._load(self._EMBEDDINGS).values())
        return [ScgEmbedding.model_validate(d) for d in docs]

    def list_sources(self) -> list[SourceDescriptor]:
        """Return all source descriptors."""
        with self._lock:
            docs = list(self._load(self._SOURCES).values())
        return [SourceDescriptor.model_validate(d) for d in docs]

    # -- Scoped delete ------------------------------------------------------

    def delete_source(self, source_id: str) -> int:
        """Delete every entity for *source_id*; return the count removed.

        NOTE — non-atomic: the scoped delete is a sequence of per-collection file
        rewrites under one lock, so a mid-sequence crash can leave dangling edges
        pointing at an already-removed node. Acceptable for the single-instance
        dev (JSON) path; the multi-worker path is the Mongo backend. No
        transaction is layered on here (YAGNI for JSON v1).
        """
        prefix = f"{source_id}#"
        removed = 0
        with self._lock:
            # Nodes whose node_id we'll need to evict from embeddings too.
            nodes = self._load(self._NODES)
            evicted_ids = {
                nid
                for nid, doc in nodes.items()
                if doc.get("source_id") == source_id
            }
            kept_nodes = {
                nid: doc for nid, doc in nodes.items() if nid not in evicted_ids
            }
            removed += len(nodes) - len(kept_nodes)
            self._save(self._NODES, kept_nodes)

            removed += self._delete_where(
                self._EDGES,
                lambda d: str(d.get("source", "")).startswith(prefix)
                or str(d.get("target", "")).startswith(prefix),
            )
            removed += self._delete_where(
                self._RECIPES,
                lambda d: str(d.get("source_key", "")).startswith(prefix),
            )
            removed += self._delete_where(
                self._EMBEDDINGS,
                lambda d: d.get("node_id") in evicted_ids,
            )
            removed += self._delete_where(
                self._SOURCES,
                lambda d: d.get("source_id") == source_id,
            )
        self._invalidate_nodes()
        return removed

    def _delete_where(self, collection: str, predicate: _DocPredicate) -> int:
        """Drop docs matching *predicate* from *collection*; return count removed.

        Caller already holds ``_lock``.
        """
        data = self._load(collection)
        kept = {k: v for k, v in data.items() if not predicate(v)}
        removed = len(data) - len(kept)
        if removed:
            self._save(collection, kept)
        return removed

__init__(root_dir: str | Path | None = None) -> None

Initialise + create the directory tree.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
288
289
290
291
292
293
294
295
296
def __init__(self, root_dir: str | Path | None = None) -> None:
    """Initialise + create the directory tree."""
    super().__init__()
    if root_dir is None:
        home = get_config_value("runtime", "cache_dir", default="") or ".mewbo"
        root_dir = Path(home) / "agentic_search" / "scg"
    self.root_dir = Path(root_dir)
    self.root_dir.mkdir(parents=True, exist_ok=True)
    self._lock = threading.Lock()

delete_source(source_id: str) -> int

Delete every entity for source_id; return the count removed.

NOTE — non-atomic: the scoped delete is a sequence of per-collection file rewrites under one lock, so a mid-sequence crash can leave dangling edges pointing at an already-removed node. Acceptable for the single-instance dev (JSON) path; the multi-worker path is the Mongo backend. No transaction is layered on here (YAGNI for JSON v1).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
def delete_source(self, source_id: str) -> int:
    """Delete every entity for *source_id*; return the count removed.

    NOTE — non-atomic: the scoped delete is a sequence of per-collection file
    rewrites under one lock, so a mid-sequence crash can leave dangling edges
    pointing at an already-removed node. Acceptable for the single-instance
    dev (JSON) path; the multi-worker path is the Mongo backend. No
    transaction is layered on here (YAGNI for JSON v1).
    """
    prefix = f"{source_id}#"
    removed = 0
    with self._lock:
        # Nodes whose node_id we'll need to evict from embeddings too.
        nodes = self._load(self._NODES)
        evicted_ids = {
            nid
            for nid, doc in nodes.items()
            if doc.get("source_id") == source_id
        }
        kept_nodes = {
            nid: doc for nid, doc in nodes.items() if nid not in evicted_ids
        }
        removed += len(nodes) - len(kept_nodes)
        self._save(self._NODES, kept_nodes)

        removed += self._delete_where(
            self._EDGES,
            lambda d: str(d.get("source", "")).startswith(prefix)
            or str(d.get("target", "")).startswith(prefix),
        )
        removed += self._delete_where(
            self._RECIPES,
            lambda d: str(d.get("source_key", "")).startswith(prefix),
        )
        removed += self._delete_where(
            self._EMBEDDINGS,
            lambda d: d.get("node_id") in evicted_ids,
        )
        removed += self._delete_where(
            self._SOURCES,
            lambda d: d.get("source_id") == source_id,
        )
    self._invalidate_nodes()
    return removed

get_node(node_id: str) -> ScgNode | None

Return one node by id, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
368
369
370
371
372
def get_node(self, node_id: str) -> ScgNode | None:
    """Return one node by id, or None if absent."""
    with self._lock:
        doc = self._load(self._NODES).get(node_id)
    return ScgNode.model_validate(doc) if doc is not None else None

list_edges(*, source: SourceKey | None = None, kind: str | None = None) -> list[ScgEdge]

Return edges matching every supplied filter (AND-composed).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
def list_edges(
    self, *, source: SourceKey | None = None, kind: str | None = None
) -> list[ScgEdge]:
    """Return edges matching every supplied filter (AND-composed)."""
    with self._lock:
        docs = list(self._load(self._EDGES).values())
    out: list[ScgEdge] = []
    for doc in docs:
        edge = ScgEdge.model_validate(doc)
        if source is not None and edge.source != source:
            continue
        if kind is not None and edge.kind != kind:
            continue
        out.append(edge)
    return out

list_embeddings() -> list[ScgEmbedding]

Return all embeddings.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
430
431
432
433
434
def list_embeddings(self) -> list[ScgEmbedding]:
    """Return all embeddings."""
    with self._lock:
        docs = list(self._load(self._EMBEDDINGS).values())
    return [ScgEmbedding.model_validate(d) for d in docs]

list_recipes(*, source_id: str | None = None) -> list[RouteRecipe]

Return route recipes, optionally scoped to one source_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
417
418
419
420
421
422
423
424
425
426
427
428
def list_recipes(self, *, source_id: str | None = None) -> list[RouteRecipe]:
    """Return route recipes, optionally scoped to one *source_id*."""
    with self._lock:
        docs = list(self._load(self._RECIPES).values())
    prefix = f"{source_id}#" if source_id is not None else None
    out: list[RouteRecipe] = []
    for doc in docs:
        recipe = RouteRecipe.model_validate(doc)
        if prefix is not None and not recipe.source_key.startswith(prefix):
            continue
        out.append(recipe)
    return out

list_sources() -> list[SourceDescriptor]

Return all source descriptors.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
436
437
438
439
440
def list_sources(self) -> list[SourceDescriptor]:
    """Return all source descriptors."""
    with self._lock:
        docs = list(self._load(self._SOURCES).values())
    return [SourceDescriptor.model_validate(d) for d in docs]

neighbors(source_key: SourceKey) -> list[ScgEdge]

Return the outgoing edges whose source is source_key.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
413
414
415
def neighbors(self, source_key: SourceKey) -> list[ScgEdge]:
    """Return the outgoing edges whose ``source`` is *source_key*."""
    return self.list_edges(source=source_key)

upsert_edges(edges: list[ScgEdge]) -> None

Upsert edges, keyed on the (source, target, kind) triple.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
339
340
341
342
343
def upsert_edges(self, edges: list[ScgEdge]) -> None:
    """Upsert edges, keyed on the ``(source, target, kind)`` triple."""
    self._upsert(
        self._EDGES, [(_edge_key(e), e.model_dump(mode="json")) for e in edges]
    )

upsert_embeddings(embeddings: list[ScgEmbedding]) -> None

Upsert embeddings, keyed on node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
352
353
354
355
356
357
def upsert_embeddings(self, embeddings: list[ScgEmbedding]) -> None:
    """Upsert embeddings, keyed on ``node_id``."""
    self._upsert(
        self._EMBEDDINGS,
        [(e.node_id, e.model_dump(mode="json")) for e in embeddings],
    )

upsert_nodes(nodes: list[ScgNode]) -> None

Upsert nodes, keyed on node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
332
333
334
335
336
337
def upsert_nodes(self, nodes: list[ScgNode]) -> None:
    """Upsert nodes, keyed on ``node_id``."""
    self._upsert(
        self._NODES, [(n.node_id, n.model_dump(mode="json")) for n in nodes]
    )
    self._invalidate_nodes()

upsert_recipes(recipes: list[RouteRecipe]) -> None

Upsert route recipes, keyed on source_key.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
345
346
347
348
349
350
def upsert_recipes(self, recipes: list[RouteRecipe]) -> None:
    """Upsert route recipes, keyed on ``source_key``."""
    self._upsert(
        self._RECIPES,
        [(r.source_key, r.model_dump(mode="json")) for r in recipes],
    )

upsert_source(descriptor: SourceDescriptor) -> None

Upsert a source descriptor, keyed on source_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
359
360
361
362
363
364
def upsert_source(self, descriptor: SourceDescriptor) -> None:
    """Upsert a source descriptor, keyed on ``source_id``."""
    self._upsert(
        self._SOURCES,
        [(descriptor.source_id, descriptor.model_dump(mode="json"))],
    )

MongoScgStore

Bases: ScgStore

MongoDB-backed SCG store.

Collections (one per family): agentic_search_scg_nodes (node_id PK), agentic_search_scg_edges ((source, target, kind) PK), agentic_search_scg_recipes (source_key PK), agentic_search_scg_embeddings (node_id PK), agentic_search_scg_sources (source_id PK).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
class MongoScgStore(ScgStore):
    """MongoDB-backed SCG store.

    Collections (one per family): ``agentic_search_scg_nodes`` (``node_id`` PK),
    ``agentic_search_scg_edges`` (``(source, target, kind)`` PK),
    ``agentic_search_scg_recipes`` (``source_key`` PK),
    ``agentic_search_scg_embeddings`` (``node_id`` PK),
    ``agentic_search_scg_sources`` (``source_id`` PK).
    """

    NODES = "agentic_search_scg_nodes"
    EDGES = "agentic_search_scg_edges"
    RECIPES = "agentic_search_scg_recipes"
    EMBEDDINGS = "agentic_search_scg_embeddings"
    SOURCES = "agentic_search_scg_sources"

    def __init__(
        self,
        *,
        client: _MongoClient | None = None,
        uri: str | None = None,
        database: str | None = None,
    ) -> None:
        """Connect + ensure unique indexes for each upsert key."""
        super().__init__()
        if client is None:
            from pymongo import MongoClient

            _uri = uri or get_config_value(
                "storage", "mongodb", "uri", default="mongodb://localhost:27017"
            )
            client = MongoClient(_uri, serverSelectionTimeoutMS=5000)
            client.admin.command("ping")
        if database is None:
            database = get_config_value(
                "storage", "mongodb", "database", default="mewbo"
            )
        self._client = client
        self._db = client[database]
        self._ensure_indexes()

    def _col(self, name: str) -> _Collection:
        return self._db[name]

    def _ensure_indexes(self) -> None:
        from pymongo import ASCENDING

        self._col(self.NODES).create_index(
            [("node_id", ASCENDING)], name="ix_scg_node_id", unique=True, background=True
        )
        self._col(self.NODES).create_index(
            [("source_id", ASCENDING)], name="ix_scg_node_source", background=True
        )
        self._col(self.EDGES).create_index(
            [("source", ASCENDING), ("target", ASCENDING), ("kind", ASCENDING)],
            name="ix_scg_edge_triple",
            unique=True,
            background=True,
        )
        self._col(self.RECIPES).create_index(
            [("source_key", ASCENDING)],
            name="ix_scg_recipe_key",
            unique=True,
            background=True,
        )
        self._col(self.EMBEDDINGS).create_index(
            [("node_id", ASCENDING)],
            name="ix_scg_emb_node",
            unique=True,
            background=True,
        )
        self._col(self.SOURCES).create_index(
            [("source_id", ASCENDING)],
            name="ix_scg_source_id",
            unique=True,
            background=True,
        )

    # -- Writes -------------------------------------------------------------

    def upsert_nodes(self, nodes: list[ScgNode]) -> None:
        """Upsert nodes, keyed on ``node_id``."""
        for n in nodes:
            self._col(self.NODES).replace_one(
                {"node_id": n.node_id}, n.model_dump(mode="json"), upsert=True
            )
        self._invalidate_nodes()

    def upsert_edges(self, edges: list[ScgEdge]) -> None:
        """Upsert edges, keyed on the ``(source, target, kind)`` triple."""
        for e in edges:
            self._col(self.EDGES).replace_one(
                {"source": e.source, "target": e.target, "kind": e.kind},
                e.model_dump(mode="json"),
                upsert=True,
            )

    def upsert_recipes(self, recipes: list[RouteRecipe]) -> None:
        """Upsert route recipes, keyed on ``source_key``."""
        for r in recipes:
            self._col(self.RECIPES).replace_one(
                {"source_key": r.source_key}, r.model_dump(mode="json"), upsert=True
            )

    def upsert_embeddings(self, embeddings: list[ScgEmbedding]) -> None:
        """Upsert embeddings, keyed on ``node_id``."""
        for emb in embeddings:
            self._col(self.EMBEDDINGS).replace_one(
                {"node_id": emb.node_id}, emb.model_dump(mode="json"), upsert=True
            )

    def upsert_source(self, descriptor: SourceDescriptor) -> None:
        """Upsert a source descriptor, keyed on ``source_id``."""
        self._col(self.SOURCES).replace_one(
            {"source_id": descriptor.source_id},
            descriptor.model_dump(mode="json"),
            upsert=True,
        )

    # -- Reads --------------------------------------------------------------

    def get_node(self, node_id: str) -> ScgNode | None:
        """Return one node by id, or None if absent."""
        doc = self._col(self.NODES).find_one({"node_id": node_id}, {"_id": 0})
        return ScgNode.model_validate(doc) if doc else None

    def _query_nodes_uncached(
        self,
        *,
        source_id: str | None = None,
        kind: str | None = None,
        name_contains: str | None = None,
    ) -> list[ScgNode]:
        """Query the nodes collection for the filter (AND-composed); no caching."""
        query: dict[str, object] = {}
        if source_id is not None:
            query["source_id"] = source_id
        if kind is not None:
            query["kind"] = kind
        if name_contains is not None:
            # Case-insensitive substring (literal — escape regex metachars).
            query["name"] = {"$regex": re.escape(name_contains), "$options": "i"}
        cursor = self._col(self.NODES).find(query, {"_id": 0})
        return [ScgNode.model_validate(d) for d in cursor]

    def list_edges(
        self, *, source: SourceKey | None = None, kind: str | None = None
    ) -> list[ScgEdge]:
        """Return edges matching every supplied filter (AND-composed)."""
        query: dict[str, object] = {}
        if source is not None:
            query["source"] = source
        if kind is not None:
            query["kind"] = kind
        cursor = self._col(self.EDGES).find(query, {"_id": 0})
        return [ScgEdge.model_validate(d) for d in cursor]

    def neighbors(self, source_key: SourceKey) -> list[ScgEdge]:
        """Return the outgoing edges whose ``source`` is *source_key*."""
        return self.list_edges(source=source_key)

    def list_recipes(self, *, source_id: str | None = None) -> list[RouteRecipe]:
        """Return route recipes, optionally scoped to one *source_id*."""
        query: dict[str, object] = {}
        if source_id is not None:
            query["source_key"] = {"$regex": f"^{re.escape(source_id)}#"}
        cursor = self._col(self.RECIPES).find(query, {"_id": 0})
        return [RouteRecipe.model_validate(d) for d in cursor]

    def list_embeddings(self) -> list[ScgEmbedding]:
        """Return all embeddings."""
        cursor = self._col(self.EMBEDDINGS).find({}, {"_id": 0})
        return [ScgEmbedding.model_validate(d) for d in cursor]

    def list_sources(self) -> list[SourceDescriptor]:
        """Return all source descriptors."""
        cursor = self._col(self.SOURCES).find({}, {"_id": 0})
        return [SourceDescriptor.model_validate(d) for d in cursor]

    # -- Scoped delete ------------------------------------------------------

    def delete_source(self, source_id: str) -> int:
        """Delete every entity for *source_id*; return the count removed."""
        prefix = f"{source_id}#"
        starts = {"$regex": f"^{re.escape(prefix)}"}
        evicted_ids = [
            d["node_id"]
            for d in self._col(self.NODES).find({"source_id": source_id}, {"node_id": 1})
        ]
        removed = 0
        removed += self._col(self.NODES).delete_many(
            {"source_id": source_id}
        ).deleted_count
        removed += self._col(self.EDGES).delete_many(
            {"$or": [{"source": starts}, {"target": starts}]}
        ).deleted_count
        removed += self._col(self.RECIPES).delete_many(
            {"source_key": starts}
        ).deleted_count
        if evicted_ids:
            removed += self._col(self.EMBEDDINGS).delete_many(
                {"node_id": {"$in": evicted_ids}}
            ).deleted_count
        removed += self._col(self.SOURCES).delete_many(
            {"source_id": source_id}
        ).deleted_count
        self._invalidate_nodes()
        return removed

__init__(*, client: _MongoClient | None = None, uri: str | None = None, database: str | None = None) -> None

Connect + ensure unique indexes for each upsert key.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
def __init__(
    self,
    *,
    client: _MongoClient | None = None,
    uri: str | None = None,
    database: str | None = None,
) -> None:
    """Connect + ensure unique indexes for each upsert key."""
    super().__init__()
    if client is None:
        from pymongo import MongoClient

        _uri = uri or get_config_value(
            "storage", "mongodb", "uri", default="mongodb://localhost:27017"
        )
        client = MongoClient(_uri, serverSelectionTimeoutMS=5000)
        client.admin.command("ping")
    if database is None:
        database = get_config_value(
            "storage", "mongodb", "database", default="mewbo"
        )
    self._client = client
    self._db = client[database]
    self._ensure_indexes()

delete_source(source_id: str) -> int

Delete every entity for source_id; return the count removed.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
def delete_source(self, source_id: str) -> int:
    """Delete every entity for *source_id*; return the count removed."""
    prefix = f"{source_id}#"
    starts = {"$regex": f"^{re.escape(prefix)}"}
    evicted_ids = [
        d["node_id"]
        for d in self._col(self.NODES).find({"source_id": source_id}, {"node_id": 1})
    ]
    removed = 0
    removed += self._col(self.NODES).delete_many(
        {"source_id": source_id}
    ).deleted_count
    removed += self._col(self.EDGES).delete_many(
        {"$or": [{"source": starts}, {"target": starts}]}
    ).deleted_count
    removed += self._col(self.RECIPES).delete_many(
        {"source_key": starts}
    ).deleted_count
    if evicted_ids:
        removed += self._col(self.EMBEDDINGS).delete_many(
            {"node_id": {"$in": evicted_ids}}
        ).deleted_count
    removed += self._col(self.SOURCES).delete_many(
        {"source_id": source_id}
    ).deleted_count
    self._invalidate_nodes()
    return removed

get_node(node_id: str) -> ScgNode | None

Return one node by id, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
628
629
630
631
def get_node(self, node_id: str) -> ScgNode | None:
    """Return one node by id, or None if absent."""
    doc = self._col(self.NODES).find_one({"node_id": node_id}, {"_id": 0})
    return ScgNode.model_validate(doc) if doc else None

list_edges(*, source: SourceKey | None = None, kind: str | None = None) -> list[ScgEdge]

Return edges matching every supplied filter (AND-composed).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
652
653
654
655
656
657
658
659
660
661
662
def list_edges(
    self, *, source: SourceKey | None = None, kind: str | None = None
) -> list[ScgEdge]:
    """Return edges matching every supplied filter (AND-composed)."""
    query: dict[str, object] = {}
    if source is not None:
        query["source"] = source
    if kind is not None:
        query["kind"] = kind
    cursor = self._col(self.EDGES).find(query, {"_id": 0})
    return [ScgEdge.model_validate(d) for d in cursor]

list_embeddings() -> list[ScgEmbedding]

Return all embeddings.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
676
677
678
679
def list_embeddings(self) -> list[ScgEmbedding]:
    """Return all embeddings."""
    cursor = self._col(self.EMBEDDINGS).find({}, {"_id": 0})
    return [ScgEmbedding.model_validate(d) for d in cursor]

list_recipes(*, source_id: str | None = None) -> list[RouteRecipe]

Return route recipes, optionally scoped to one source_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
668
669
670
671
672
673
674
def list_recipes(self, *, source_id: str | None = None) -> list[RouteRecipe]:
    """Return route recipes, optionally scoped to one *source_id*."""
    query: dict[str, object] = {}
    if source_id is not None:
        query["source_key"] = {"$regex": f"^{re.escape(source_id)}#"}
    cursor = self._col(self.RECIPES).find(query, {"_id": 0})
    return [RouteRecipe.model_validate(d) for d in cursor]

list_sources() -> list[SourceDescriptor]

Return all source descriptors.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
681
682
683
684
def list_sources(self) -> list[SourceDescriptor]:
    """Return all source descriptors."""
    cursor = self._col(self.SOURCES).find({}, {"_id": 0})
    return [SourceDescriptor.model_validate(d) for d in cursor]

neighbors(source_key: SourceKey) -> list[ScgEdge]

Return the outgoing edges whose source is source_key.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
664
665
666
def neighbors(self, source_key: SourceKey) -> list[ScgEdge]:
    """Return the outgoing edges whose ``source`` is *source_key*."""
    return self.list_edges(source=source_key)

upsert_edges(edges: list[ScgEdge]) -> None

Upsert edges, keyed on the (source, target, kind) triple.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
595
596
597
598
599
600
601
602
def upsert_edges(self, edges: list[ScgEdge]) -> None:
    """Upsert edges, keyed on the ``(source, target, kind)`` triple."""
    for e in edges:
        self._col(self.EDGES).replace_one(
            {"source": e.source, "target": e.target, "kind": e.kind},
            e.model_dump(mode="json"),
            upsert=True,
        )

upsert_embeddings(embeddings: list[ScgEmbedding]) -> None

Upsert embeddings, keyed on node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
611
612
613
614
615
616
def upsert_embeddings(self, embeddings: list[ScgEmbedding]) -> None:
    """Upsert embeddings, keyed on ``node_id``."""
    for emb in embeddings:
        self._col(self.EMBEDDINGS).replace_one(
            {"node_id": emb.node_id}, emb.model_dump(mode="json"), upsert=True
        )

upsert_nodes(nodes: list[ScgNode]) -> None

Upsert nodes, keyed on node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
587
588
589
590
591
592
593
def upsert_nodes(self, nodes: list[ScgNode]) -> None:
    """Upsert nodes, keyed on ``node_id``."""
    for n in nodes:
        self._col(self.NODES).replace_one(
            {"node_id": n.node_id}, n.model_dump(mode="json"), upsert=True
        )
    self._invalidate_nodes()

upsert_recipes(recipes: list[RouteRecipe]) -> None

Upsert route recipes, keyed on source_key.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
604
605
606
607
608
609
def upsert_recipes(self, recipes: list[RouteRecipe]) -> None:
    """Upsert route recipes, keyed on ``source_key``."""
    for r in recipes:
        self._col(self.RECIPES).replace_one(
            {"source_key": r.source_key}, r.model_dump(mode="json"), upsert=True
        )

upsert_source(descriptor: SourceDescriptor) -> None

Upsert a source descriptor, keyed on source_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
618
619
620
621
622
623
624
def upsert_source(self, descriptor: SourceDescriptor) -> None:
    """Upsert a source descriptor, keyed on ``source_id``."""
    self._col(self.SOURCES).replace_one(
        {"source_id": descriptor.source_id},
        descriptor.model_dump(mode="json"),
        upsert=True,
    )

ScgStore

Bases: ABC

Abstract base for SCG structure-persistence backends.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
class ScgStore(abc.ABC):
    """Abstract base for SCG structure-persistence backends."""

    def __init__(self) -> None:
        """Initialise the shared node-query cache (drivers must ``super().__init__``)."""
        self._node_cache = _NodeQueryCache()

    def _invalidate_nodes(self) -> None:
        """Drop cached ``query_nodes`` results after a node-collection write."""
        self._node_cache.clear()

    # -- Writes -------------------------------------------------------------

    @abc.abstractmethod
    def upsert_nodes(self, nodes: list[ScgNode]) -> None:
        """Upsert nodes, keyed on ``node_id``."""

    @abc.abstractmethod
    def upsert_edges(self, edges: list[ScgEdge]) -> None:
        """Upsert edges, keyed on the ``(source, target, kind)`` triple."""

    @abc.abstractmethod
    def upsert_recipes(self, recipes: list[RouteRecipe]) -> None:
        """Upsert route recipes, keyed on ``source_key``."""

    @abc.abstractmethod
    def upsert_embeddings(self, embeddings: list[ScgEmbedding]) -> None:
        """Upsert embeddings, keyed on ``node_id``."""

    @abc.abstractmethod
    def upsert_source(self, descriptor: SourceDescriptor) -> None:
        """Upsert a source descriptor, keyed on ``source_id``."""

    # -- Reads --------------------------------------------------------------

    @abc.abstractmethod
    def get_node(self, node_id: str) -> ScgNode | None:
        """Return one node by id, or None if absent."""

    def query_nodes(
        self,
        *,
        source_id: str | None = None,
        kind: str | None = None,
        name_contains: str | None = None,
    ) -> list[ScgNode]:
        """Return nodes matching every supplied filter (AND-composed).

        Memoized by the filter triple (see :class:`_NodeQueryCache`); node writes
        invalidate the cache. Backends implement the raw scan in
        :meth:`_query_nodes_uncached`.
        """
        key = (source_id, kind, name_contains)
        cached = self._node_cache.get(key)
        if cached is not None:
            return cached
        nodes = self._query_nodes_uncached(
            source_id=source_id, kind=kind, name_contains=name_contains
        )
        self._node_cache.put(key, nodes)
        return nodes

    @abc.abstractmethod
    def _query_nodes_uncached(
        self,
        *,
        source_id: str | None = None,
        kind: str | None = None,
        name_contains: str | None = None,
    ) -> list[ScgNode]:
        """Scan the node collection for the filter, with no caching."""

    @abc.abstractmethod
    def list_edges(
        self, *, source: SourceKey | None = None, kind: str | None = None
    ) -> list[ScgEdge]:
        """Return edges matching every supplied filter (AND-composed)."""

    @abc.abstractmethod
    def neighbors(self, source_key: SourceKey) -> list[ScgEdge]:
        """Return the outgoing edges whose ``source`` is *source_key*."""

    @abc.abstractmethod
    def list_recipes(self, *, source_id: str | None = None) -> list[RouteRecipe]:
        """Return route recipes, optionally scoped to one *source_id*."""

    @abc.abstractmethod
    def list_embeddings(self) -> list[ScgEmbedding]:
        """Return all embeddings."""

    @abc.abstractmethod
    def list_sources(self) -> list[SourceDescriptor]:
        """Return all source descriptors."""

    # -- Vector search + scoped delete -------------------------------------

    def vector_search(self, qvec: list[float], k: int) -> list[tuple[str, float]]:
        """Return ``(node_id, cosine_score)`` for the top-*k* embeddings.

        Brute-force cosine over every stored vector — the documented scale seam
        (mirrors the wiki ``vector_search``): an ANN index lands behind this
        signature without changing callers.
        """
        embeddings = self.list_embeddings()
        if not embeddings:
            return []
        scored = [
            (e.node_id, Embedder.cosine(qvec, e.vector)) for e in embeddings
        ]
        scored.sort(key=lambda t: t[1], reverse=True)
        return scored[:k]

    @abc.abstractmethod
    def delete_source(self, source_id: str) -> int:
        """Delete every node/edge/recipe/embedding/source for *source_id*.

        Scoped wipe for a clean re-map; returns the total document count removed.
        """

__init__() -> None

Initialise the shared node-query cache (drivers must super().__init__).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
151
152
153
def __init__(self) -> None:
    """Initialise the shared node-query cache (drivers must ``super().__init__``)."""
    self._node_cache = _NodeQueryCache()

delete_source(source_id: str) -> int abstractmethod

Delete every node/edge/recipe/embedding/source for source_id.

Scoped wipe for a clean re-map; returns the total document count removed.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
260
261
262
263
264
265
@abc.abstractmethod
def delete_source(self, source_id: str) -> int:
    """Delete every node/edge/recipe/embedding/source for *source_id*.

    Scoped wipe for a clean re-map; returns the total document count removed.
    """

get_node(node_id: str) -> ScgNode | None abstractmethod

Return one node by id, or None if absent.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
183
184
185
@abc.abstractmethod
def get_node(self, node_id: str) -> ScgNode | None:
    """Return one node by id, or None if absent."""

list_edges(*, source: SourceKey | None = None, kind: str | None = None) -> list[ScgEdge] abstractmethod

Return edges matching every supplied filter (AND-composed).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
220
221
222
223
224
@abc.abstractmethod
def list_edges(
    self, *, source: SourceKey | None = None, kind: str | None = None
) -> list[ScgEdge]:
    """Return edges matching every supplied filter (AND-composed)."""

list_embeddings() -> list[ScgEmbedding] abstractmethod

Return all embeddings.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
234
235
236
@abc.abstractmethod
def list_embeddings(self) -> list[ScgEmbedding]:
    """Return all embeddings."""

list_recipes(*, source_id: str | None = None) -> list[RouteRecipe] abstractmethod

Return route recipes, optionally scoped to one source_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
230
231
232
@abc.abstractmethod
def list_recipes(self, *, source_id: str | None = None) -> list[RouteRecipe]:
    """Return route recipes, optionally scoped to one *source_id*."""

list_sources() -> list[SourceDescriptor] abstractmethod

Return all source descriptors.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
238
239
240
@abc.abstractmethod
def list_sources(self) -> list[SourceDescriptor]:
    """Return all source descriptors."""

neighbors(source_key: SourceKey) -> list[ScgEdge] abstractmethod

Return the outgoing edges whose source is source_key.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
226
227
228
@abc.abstractmethod
def neighbors(self, source_key: SourceKey) -> list[ScgEdge]:
    """Return the outgoing edges whose ``source`` is *source_key*."""

query_nodes(*, source_id: str | None = None, kind: str | None = None, name_contains: str | None = None) -> list[ScgNode]

Return nodes matching every supplied filter (AND-composed).

Memoized by the filter triple (see :class:_NodeQueryCache); node writes invalidate the cache. Backends implement the raw scan in :meth:_query_nodes_uncached.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
def query_nodes(
    self,
    *,
    source_id: str | None = None,
    kind: str | None = None,
    name_contains: str | None = None,
) -> list[ScgNode]:
    """Return nodes matching every supplied filter (AND-composed).

    Memoized by the filter triple (see :class:`_NodeQueryCache`); node writes
    invalidate the cache. Backends implement the raw scan in
    :meth:`_query_nodes_uncached`.
    """
    key = (source_id, kind, name_contains)
    cached = self._node_cache.get(key)
    if cached is not None:
        return cached
    nodes = self._query_nodes_uncached(
        source_id=source_id, kind=kind, name_contains=name_contains
    )
    self._node_cache.put(key, nodes)
    return nodes

upsert_edges(edges: list[ScgEdge]) -> None abstractmethod

Upsert edges, keyed on the (source, target, kind) triple.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
165
166
167
@abc.abstractmethod
def upsert_edges(self, edges: list[ScgEdge]) -> None:
    """Upsert edges, keyed on the ``(source, target, kind)`` triple."""

upsert_embeddings(embeddings: list[ScgEmbedding]) -> None abstractmethod

Upsert embeddings, keyed on node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
173
174
175
@abc.abstractmethod
def upsert_embeddings(self, embeddings: list[ScgEmbedding]) -> None:
    """Upsert embeddings, keyed on ``node_id``."""

upsert_nodes(nodes: list[ScgNode]) -> None abstractmethod

Upsert nodes, keyed on node_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
161
162
163
@abc.abstractmethod
def upsert_nodes(self, nodes: list[ScgNode]) -> None:
    """Upsert nodes, keyed on ``node_id``."""

upsert_recipes(recipes: list[RouteRecipe]) -> None abstractmethod

Upsert route recipes, keyed on source_key.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
169
170
171
@abc.abstractmethod
def upsert_recipes(self, recipes: list[RouteRecipe]) -> None:
    """Upsert route recipes, keyed on ``source_key``."""

upsert_source(descriptor: SourceDescriptor) -> None abstractmethod

Upsert a source descriptor, keyed on source_id.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
177
178
179
@abc.abstractmethod
def upsert_source(self, descriptor: SourceDescriptor) -> None:
    """Upsert a source descriptor, keyed on ``source_id``."""

Return (node_id, cosine_score) for the top-k embeddings.

Brute-force cosine over every stored vector — the documented scale seam (mirrors the wiki vector_search): an ANN index lands behind this signature without changing callers.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
def vector_search(self, qvec: list[float], k: int) -> list[tuple[str, float]]:
    """Return ``(node_id, cosine_score)`` for the top-*k* embeddings.

    Brute-force cosine over every stored vector — the documented scale seam
    (mirrors the wiki ``vector_search``): an ANN index lands behind this
    signature without changing callers.
    """
    embeddings = self.list_embeddings()
    if not embeddings:
        return []
    scored = [
        (e.node_id, Embedder.cosine(qvec, e.vector)) for e in embeddings
    ]
    scored.sort(key=lambda t: t[1], reverse=True)
    return scored[:k]

create_scg_store() -> ScgStore

Return the configured SCG store driver (storage.driver; default JSON).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
722
723
724
725
726
727
def create_scg_store() -> ScgStore:
    """Return the configured SCG store driver (``storage.driver``; default JSON)."""
    driver = get_config_value("storage", "driver", default="json")
    if driver == "mongodb":
        return MongoScgStore()
    return JsonScgStore()

get_scg_store() -> ScgStore

Return the process-wide SCG store, creating it on first use.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
734
735
736
737
738
739
740
def get_scg_store() -> ScgStore:
    """Return the process-wide SCG store, creating it on first use."""
    global _store_singleton
    with _singleton_lock:
        if _store_singleton is None:
            _store_singleton = create_scg_store()
        return _store_singleton

reset_for_tests() -> None

Swap in a fresh, empty JSON store under a throwaway temp dir.

Keeps unit tests isolated from real data while still exercising the JSON backend end-to-end (mirrors the run store's reset_for_tests).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
750
751
752
753
754
755
756
757
def reset_for_tests() -> None:
    """Swap in a fresh, empty JSON store under a throwaway temp dir.

    Keeps unit tests isolated from real data while still exercising the JSON
    backend end-to-end (mirrors the run store's ``reset_for_tests``).
    """
    tmp = Path(tempfile.mkdtemp(prefix="mewbo-scg-"))
    set_scg_store(JsonScgStore(root_dir=tmp))

set_scg_store(store: ScgStore | None) -> None

Override the process-wide SCG store (used by tests).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/store.py
743
744
745
746
747
def set_scg_store(store: ScgStore | None) -> None:
    """Override the process-wide SCG store (used by tests)."""
    global _store_singleton
    with _singleton_lock:
        _store_singleton = store

mewbo_graph.scg.types

Typed contracts for the Source Capability Graph (SCG) — spec §6.

The SCG indexes reachability — the schemas and qualified pathways a source exposes, never the data behind them. These models are the shared surface the structure providers, parser, router, and traversal engine all build against; everything else in the scg package references them.

Conventions mirror :mod:mewbo_api.agentic_search.schemas and the wiki types:

  • Every model subclasses :class:_Wire (extra="forbid", populate_by_name=True) so unknown keys are rejected at the boundary.
  • node_id is a deterministic sha1(source_key|kind)[:16] derived by :meth:ScgNode.make_id and overwritten on every validate — that derivation is the single source of node identity (mirrors the wiki graph's _stable_id and the memory layer's content-addressed node ids).

Security invariant (spec §6): SCG nodes carry only a redacted auth_scope descriptor string — never persist tokens, credentials, or any secret.

CapabilityBinding

Bases: _Wire

One field a capability binds, with its access mode + allowed operators.

Binding patterns keep traversal honest: a capability "queryable by service_id, not free-text" emits mode="bound" so the router only proposes executable plans (Florescu/Vassalos SIGMOD'99).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/types.py
65
66
67
68
69
70
71
72
73
74
75
class CapabilityBinding(_Wire):
    """One field a capability binds, with its access mode + allowed operators.

    Binding patterns keep traversal honest: a capability "queryable by
    ``service_id``, not free-text" emits ``mode="bound"`` so the router only
    proposes *executable* plans (Florescu/Vassalos SIGMOD'99).
    """

    field_key: SourceKey
    mode: BindMode
    operators: list[str] = Field(default_factory=list)

RouteRecipe

Bases: _Wire

A precomputed qualified path (ordered SourceKey steps) over the SCG.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/types.py
140
141
142
143
144
145
class RouteRecipe(_Wire):
    """A precomputed qualified path (ordered ``SourceKey`` steps) over the SCG."""

    source_key: SourceKey
    steps: list[SourceKey]
    cost_estimate: float = 0.0

ScgEdge

Bases: _Wire

A directed, weighted, provenanced edge between two SourceKey nodes.

binds records the (source-field, target-field) pair an edge aligns on; method is the parser's evidence kind. valid_at / invalid_at carry the invalidate-don't-delete validity window (Graphiti) so learned edges can be retired without losing provenance.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/types.py
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
class ScgEdge(_Wire):
    """A directed, weighted, provenanced edge between two ``SourceKey`` nodes.

    ``binds`` records the (source-field, target-field) pair an edge aligns on;
    ``method`` is the parser's evidence kind. ``valid_at`` / ``invalid_at``
    carry the invalidate-don't-delete validity window (Graphiti) so learned
    edges can be retired without losing provenance.
    """

    source: SourceKey
    target: SourceKey
    kind: EdgeKind
    weight: float = 1.0
    binds: tuple[SourceKey, SourceKey] | None = None
    method: EdgeMethod | None = None
    evidence: list[str] = Field(default_factory=list)
    valid_at: str | None = None
    invalid_at: str | None = None

ScgEmbedding

Bases: _Wire

A dense embedding vector for an SCG node (parallels the wiki Embedding).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/types.py
151
152
153
154
155
156
157
class ScgEmbedding(_Wire):
    """A dense embedding vector for an SCG node (parallels the wiki Embedding)."""

    node_id: str
    vector: list[float]
    model: str
    dim: int

ScgNode

Bases: _Wire

A node in the Source Capability Graph.

node_id is always the canonical sha1(source_key|kind)[:16] — any supplied value is overwritten on validate so identity stays content-addressed and stable across re-indexes.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/types.py
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
class ScgNode(_Wire):
    """A node in the Source Capability Graph.

    ``node_id`` is *always* the canonical ``sha1(source_key|kind)[:16]`` — any
    supplied value is overwritten on validate so identity stays content-addressed
    and stable across re-indexes.
    """

    source_key: SourceKey
    node_id: str = ""
    kind: NodeKind
    source_id: str
    name: str
    doc: str = ""
    example_queries: list[str] = Field(default_factory=list)
    bindings: list[CapabilityBinding] = Field(default_factory=list)
    # Redacted auth descriptor ONLY — never a token/credential (spec §6).
    auth_scope: str | None = None

    @staticmethod
    def make_id(source_key: SourceKey, kind: NodeKind) -> str:
        """Deterministic node id over ``(source_key, kind)`` — sha1[:16]."""
        return hashlib.sha1(f"{source_key}|{kind}".encode()).hexdigest()[:16]

    @model_validator(mode="after")
    def _derive_node_id(self) -> ScgNode:
        """Force ``node_id`` to the canonical derivation (overwrites any input)."""
        canonical = self.make_id(self.source_key, self.kind)
        if self.node_id != canonical:
            object.__setattr__(self, "node_id", canonical)
        return self

make_id(source_key: SourceKey, kind: NodeKind) -> str staticmethod

Deterministic node id over (source_key, kind) — sha1[:16].

Source code in packages/mewbo_graph/src/mewbo_graph/scg/types.py
100
101
102
103
@staticmethod
def make_id(source_key: SourceKey, kind: NodeKind) -> str:
    """Deterministic node id over ``(source_key, kind)`` — sha1[:16]."""
    return hashlib.sha1(f"{source_key}|{kind}".encode()).hexdigest()[:16]

SourceDescriptor

Bases: _Wire

The raw, source-type-specific descriptor a structure provider parses.

raw is the opaque provider payload (OpenAPI doc, MCP tool list, GraphQL SDL, SQL schema…). Carries no secrets — auth lives in the connector config, not here.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/types.py
163
164
165
166
167
168
169
170
171
172
173
174
class SourceDescriptor(_Wire):
    """The raw, source-type-specific descriptor a structure provider parses.

    ``raw`` is the opaque provider payload (OpenAPI doc, MCP tool list, GraphQL
    SDL, SQL schema…). Carries no secrets — auth lives in the connector config,
    not here.
    """

    source_id: str
    source_type: str
    raw: dict[str, object]
    schema_version: str | None = None

StructureGraph

Bases: _Wire

The normalized provider output: nodes + edges + recipes for one source.

A provider returns one source's subgraph; ScgParser.parse_source upserts it into the persisted whole-catalog SCG directly (one source at a time).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/types.py
180
181
182
183
184
185
186
187
188
189
class StructureGraph(_Wire):
    """The normalized provider output: nodes + edges + recipes for one source.

    A provider returns one source's subgraph; ``ScgParser.parse_source`` upserts
    it into the persisted whole-catalog SCG directly (one source at a time).
    """

    nodes: list[ScgNode] = Field(default_factory=list)
    edges: list[ScgEdge] = Field(default_factory=list)
    recipes: list[RouteRecipe] = Field(default_factory=list)

field_leaf(field_key: SourceKey) -> str

Return the trailing .-segment of a <cap>.<name> field key (lower).

The one canonical home for the "trailing field-name segment, lower-cased" idiom shared by the parser's field indexing and the type aligner's field-overlap heuristic (DRY).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/types.py
36
37
38
39
40
41
42
43
def field_leaf(field_key: SourceKey) -> str:
    """Return the trailing ``.``-segment of a ``<cap>.<name>`` field key (lower).

    The one canonical home for the "trailing field-name segment, lower-cased"
    idiom shared by the parser's field indexing and the type aligner's
    field-overlap heuristic (DRY).
    """
    return field_key.rsplit(".", 1)[-1].lower()

mewbo_graph.scg.providers

SCG structure providers — the information→graph parser seam.

One :class:SourceStructureProvider per source type (RML declarative shell): OpenAPI, MCP tool list, and an LLM fallback for schemaless sources. A :class:StructureProviderRegistry dispatches a :class:~mewbo_graph.scg.types.SourceDescriptor to the matching provider by source_type. New source type = one class + one register call.

LlmStructureProvider

Coarse single-capability provider for schemaless sources (DI LLM).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/llm_fallback.py
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
class LlmStructureProvider:
    """Coarse single-capability provider for schemaless sources (DI LLM)."""

    source_type = "text"

    # Slugify a free-text label into a capability name token.
    _SLUG_RE = re.compile(r"[^a-z0-9]+")

    def __init__(self, *, llm: Callable[[str], str] | None = None) -> None:
        """Inject the text LLM callable; ``None`` (default) raises on use."""
        self._llm = llm

    def build_structure(self, descriptor: SourceDescriptor) -> StructureGraph:
        """Emit a source node + one coarse capability node from a description."""
        if self._llm is None:
            raise RuntimeError(
                "LlmStructureProvider requires an injected `llm` callable; "
                f"cannot parse schemaless source {descriptor.source_id!r}."
            )
        source_id = descriptor.source_id
        label = self._slug(self._llm(self._prompt(descriptor)))
        cap_key = f"{source_id}#{label}"

        return StructureGraph(
            nodes=[
                ScgNode(
                    source_key=source_id,
                    kind="source",
                    source_id=source_id,
                    name=source_id,
                    doc=self._description(descriptor),
                ),
                ScgNode(
                    source_key=cap_key,
                    kind="capability",
                    source_id=source_id,
                    name=label,
                    doc=self._description(descriptor),
                ),
            ],
            edges=[ScgEdge(source=source_id, target=cap_key, kind="HAS_ENTITY")],
        )

    # -- helpers ------------------------------------------------------------

    @classmethod
    def _slug(cls, label: str) -> str:
        """Normalize an LLM label to a capability name token; fallback ``search``."""
        slug = cls._SLUG_RE.sub("_", label.strip().lower()).strip("_")
        return slug or "search"

    @staticmethod
    def _description(descriptor: SourceDescriptor) -> str:
        """Pull a human description from the raw descriptor, or ''."""
        for key in ("description", "desc", "summary"):
            value = descriptor.raw.get(key)
            if isinstance(value, str) and value:
                return value
        return ""

    @classmethod
    def _prompt(cls, descriptor: SourceDescriptor) -> str:
        """Build the capability-labeling prompt for the injected LLM."""
        return (
            "Name the single primary capability of this data source as a short "
            "snake_case verb phrase (e.g. search_crm). Respond with only the "
            f"label.\n\nSource id: {descriptor.source_id}\n"
            f"Description: {cls._description(descriptor) or '(none)'}"
        )

__init__(*, llm: Callable[[str], str] | None = None) -> None

Inject the text LLM callable; None (default) raises on use.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/llm_fallback.py
36
37
38
def __init__(self, *, llm: Callable[[str], str] | None = None) -> None:
    """Inject the text LLM callable; ``None`` (default) raises on use."""
    self._llm = llm

build_structure(descriptor: SourceDescriptor) -> StructureGraph

Emit a source node + one coarse capability node from a description.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/llm_fallback.py
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
def build_structure(self, descriptor: SourceDescriptor) -> StructureGraph:
    """Emit a source node + one coarse capability node from a description."""
    if self._llm is None:
        raise RuntimeError(
            "LlmStructureProvider requires an injected `llm` callable; "
            f"cannot parse schemaless source {descriptor.source_id!r}."
        )
    source_id = descriptor.source_id
    label = self._slug(self._llm(self._prompt(descriptor)))
    cap_key = f"{source_id}#{label}"

    return StructureGraph(
        nodes=[
            ScgNode(
                source_key=source_id,
                kind="source",
                source_id=source_id,
                name=source_id,
                doc=self._description(descriptor),
            ),
            ScgNode(
                source_key=cap_key,
                kind="capability",
                source_id=source_id,
                name=label,
                doc=self._description(descriptor),
            ),
        ],
        edges=[ScgEdge(source=source_id, target=cap_key, kind="HAS_ENTITY")],
    )

McpToolListStructureProvider

Parse an MCP tool-list descriptor.raw into a StructureGraph.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/mcp_tool_list.py
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
class McpToolListStructureProvider:
    """Parse an MCP tool-list ``descriptor.raw`` into a StructureGraph."""

    source_type = "mcp_tool_list"

    def build_structure(self, descriptor: SourceDescriptor) -> StructureGraph:
        """Build a capability-per-tool subgraph for one MCP server source."""
        source_id = descriptor.source_id
        nodes: list[ScgNode] = [
            ScgNode(
                source_key=source_id,
                kind="source",
                source_id=source_id,
                name=source_id,
            )
        ]
        edges: list[ScgEdge] = []

        for tool in self._tools(descriptor.raw):
            cap_node, cap_edges = self._capability(source_id, tool)
            if cap_node is None:
                continue
            nodes.append(cap_node)
            edges.extend(cap_edges)

        return StructureGraph(nodes=nodes, edges=edges)

    # -- capability ---------------------------------------------------------

    @classmethod
    def _capability(
        cls, source_id: str, tool: dict[str, object]
    ) -> tuple[ScgNode | None, list[ScgEdge]]:
        """Build (capability node, edges) for one tool, or (None, []) if unnamed."""
        name = tool.get("name")
        if not isinstance(name, str) or not name:
            return None, []
        cap_key = f"{source_id}#{name}"

        bindings: list[CapabilityBinding] = []
        edges: list[ScgEdge] = [
            ScgEdge(source=source_id, target=cap_key, kind="HAS_ENTITY")
        ]

        input_props, required = cls._schema_props(tool.get("inputSchema"))
        for field_name in input_props:
            field_key = f"{cap_key}.{field_name}"
            bindings.append(
                CapabilityBinding(
                    field_key=field_key,
                    mode="bound" if field_name in required else "optional",
                    operators=list(_TOOL_OPERATORS),
                )
            )
            edges.append(
                ScgEdge(source=cap_key, target=field_key, kind="SUPPORTS_QUERY")
            )

        output_props, _ = cls._schema_props(tool.get("outputSchema"))
        for field_name in output_props:
            edges.append(
                ScgEdge(
                    source=cap_key,
                    target=f"{cap_key}.{field_name}",
                    kind="PRODUCES",
                )
            )

        return (
            ScgNode(
                source_key=cap_key,
                kind="capability",
                source_id=source_id,
                name=name,
                doc=cls._doc(tool),
                bindings=bindings,
            ),
            edges,
        )

    # -- raw-dict accessors -------------------------------------------------

    @staticmethod
    def _tools(raw: dict[str, object]) -> list[dict[str, object]]:
        """Return the tool dicts under ``raw["tools"]`` (skips malformed entries)."""
        tools = raw.get("tools")
        if not isinstance(tools, list):
            return []
        return [t for t in tools if isinstance(t, dict)]

    @staticmethod
    def _schema_props(schema: object) -> tuple[list[str], set[str]]:
        """Return (property names in declared order, required-name set)."""
        if not isinstance(schema, dict):
            return [], set()
        props = schema.get("properties")
        names = [str(k) for k in props] if isinstance(props, dict) else []
        req = schema.get("required")
        required = {str(r) for r in req} if isinstance(req, list) else set()
        return names, required

    @staticmethod
    def _doc(tool: dict[str, object]) -> str:
        """Return the tool's description string, or ''."""
        desc = tool.get("description")
        return desc if isinstance(desc, str) else ""

build_structure(descriptor: SourceDescriptor) -> StructureGraph

Build a capability-per-tool subgraph for one MCP server source.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/mcp_tool_list.py
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
def build_structure(self, descriptor: SourceDescriptor) -> StructureGraph:
    """Build a capability-per-tool subgraph for one MCP server source."""
    source_id = descriptor.source_id
    nodes: list[ScgNode] = [
        ScgNode(
            source_key=source_id,
            kind="source",
            source_id=source_id,
            name=source_id,
        )
    ]
    edges: list[ScgEdge] = []

    for tool in self._tools(descriptor.raw):
        cap_node, cap_edges = self._capability(source_id, tool)
        if cap_node is None:
            continue
        nodes.append(cap_node)
        edges.extend(cap_edges)

    return StructureGraph(nodes=nodes, edges=edges)

OpenApiStructureProvider

Parse an OpenAPI/Swagger descriptor.raw dict into a StructureGraph.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/openapi.py
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
class OpenApiStructureProvider:
    """Parse an OpenAPI/Swagger ``descriptor.raw`` dict into a StructureGraph."""

    source_type = "openapi"

    def build_structure(self, descriptor: SourceDescriptor) -> StructureGraph:
        """Build the source/entity/capability subgraph for one OpenAPI source."""
        source_id = descriptor.source_id
        raw = descriptor.raw

        nodes: list[ScgNode] = [self._source_node(descriptor)]
        edges: list[ScgEdge] = []

        for entity_nodes, entity_edges in self._iter_entities(source_id, raw):
            nodes.extend(entity_nodes)
            edges.extend(entity_edges)

        for cap_node, cap_edges in self._iter_capabilities(source_id, raw):
            nodes.append(cap_node)
            edges.extend(cap_edges)

        return StructureGraph(nodes=nodes, edges=edges)

    # -- nodes --------------------------------------------------------------

    @staticmethod
    def _source_node(descriptor: SourceDescriptor) -> ScgNode:
        """The root ``source`` node — carries an auth_scope descriptor, no secret."""
        info = descriptor.raw.get("info")
        doc = ""
        if isinstance(info, dict):
            title = info.get("title")
            if isinstance(title, str):
                doc = title
        return ScgNode(
            source_key=descriptor.source_id,
            kind="source",
            source_id=descriptor.source_id,
            name=descriptor.source_id,
            doc=doc,
        )

    @classmethod
    def _iter_entities(
        cls, source_id: str, raw: dict[str, object]
    ) -> list[tuple[list[ScgNode], list[ScgEdge]]]:
        """Yield (nodes, edges) per ``components.schemas`` entry."""
        out: list[tuple[list[ScgNode], list[ScgEdge]]] = []
        for name, schema in cls._schemas(raw).items():
            entity_key = f"{source_id}#{name}"
            nodes: list[ScgNode] = [
                ScgNode(
                    source_key=entity_key,
                    kind="entity_type",
                    source_id=source_id,
                    name=name,
                )
            ]
            edges: list[ScgEdge] = [
                ScgEdge(source=source_id, target=entity_key, kind="HAS_ENTITY")
            ]
            for field_name in cls._properties(schema):
                field_key = f"{entity_key}.{field_name}"
                nodes.append(
                    ScgNode(
                        source_key=field_key,
                        kind="field",
                        source_id=source_id,
                        name=field_name,
                    )
                )
                edges.append(
                    ScgEdge(source=entity_key, target=field_key, kind="HAS_FIELD")
                )
            out.append((nodes, edges))
        return out

    @classmethod
    def _iter_capabilities(
        cls, source_id: str, raw: dict[str, object]
    ) -> list[tuple[ScgNode, list[ScgEdge]]]:
        """Yield (capability node, edges) per operation across all paths."""
        out: list[tuple[ScgNode, list[ScgEdge]]] = []
        for op_name, operation in cls._operations(raw):
            cap_key = f"{source_id}#{op_name}"
            bindings: list[CapabilityBinding] = []
            edges: list[ScgEdge] = []
            for param in cls._parameters(operation):
                binding = cls._binding(cap_key, param)
                if binding is None:
                    continue
                bindings.append(binding)
                edges.append(
                    ScgEdge(
                        source=cap_key, target=binding.field_key, kind="SUPPORTS_QUERY"
                    )
                )
            out.append(
                (
                    ScgNode(
                        source_key=cap_key,
                        kind="capability",
                        source_id=source_id,
                        name=op_name,
                        doc=cls._operation_doc(operation),
                        bindings=bindings,
                    ),
                    edges,
                )
            )
        return out

    # -- binding ------------------------------------------------------------

    @staticmethod
    def _binding(
        cap_key: str, param: dict[str, object]
    ) -> CapabilityBinding | None:
        """Build a binding from one OpenAPI parameter, or None if unnamed."""
        name = param.get("name")
        if not isinstance(name, str) or not name:
            return None
        required = bool(param.get("required", False))
        location = param.get("in")
        operators = (
            list(_QUERY_OPERATORS) if location == "query" else []
        )
        return CapabilityBinding(
            field_key=f"{cap_key}.{name}",
            mode="bound" if required else "optional",
            operators=operators,
        )

    # -- raw-dict accessors (defensive: descriptors are untrusted shape) ----

    @staticmethod
    def _schemas(raw: dict[str, object]) -> dict[str, dict[str, object]]:
        """Return ``components.schemas`` (OpenAPI 3) merged with ``definitions`` (2.0)."""
        out: dict[str, dict[str, object]] = {}
        components = raw.get("components")
        if isinstance(components, dict):
            schemas = components.get("schemas")
            if isinstance(schemas, dict):
                for k, v in schemas.items():
                    if isinstance(v, dict):
                        out[str(k)] = v
        definitions = raw.get("definitions")  # Swagger 2.0
        if isinstance(definitions, dict):
            for k, v in definitions.items():
                if isinstance(v, dict):
                    out.setdefault(str(k), v)
        return out

    @staticmethod
    def _properties(schema: dict[str, object]) -> list[str]:
        """Return the property names of an object schema, in declared order."""
        props = schema.get("properties")
        if isinstance(props, dict):
            return [str(k) for k in props]
        return []

    @classmethod
    def _operations(
        cls, raw: dict[str, object]
    ) -> list[tuple[str, dict[str, object]]]:
        """Return (operation_id, operation) for every HTTP method under paths."""
        methods = ("get", "put", "post", "delete", "patch", "head", "options", "trace")
        out: list[tuple[str, dict[str, object]]] = []
        paths = raw.get("paths")
        if not isinstance(paths, dict):
            return out
        for path, item in paths.items():
            if not isinstance(item, dict):
                continue
            for method in methods:
                operation = item.get(method)
                if not isinstance(operation, dict):
                    continue
                op_id = operation.get("operationId")
                name = (
                    str(op_id)
                    if isinstance(op_id, str) and op_id
                    else f"{method}_{path}"
                )
                out.append((name, operation))
        return out

    @staticmethod
    def _parameters(operation: dict[str, object]) -> list[dict[str, object]]:
        """Return the operation's parameter dicts (skips malformed entries)."""
        params = operation.get("parameters")
        if not isinstance(params, list):
            return []
        return [p for p in params if isinstance(p, dict)]

    @staticmethod
    def _operation_doc(operation: dict[str, object]) -> str:
        """Prefer ``summary`` then ``description`` for the capability doc."""
        for key in ("summary", "description"):
            value = operation.get(key)
            if isinstance(value, str) and value:
                return value
        return ""

build_structure(descriptor: SourceDescriptor) -> StructureGraph

Build the source/entity/capability subgraph for one OpenAPI source.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/openapi.py
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
def build_structure(self, descriptor: SourceDescriptor) -> StructureGraph:
    """Build the source/entity/capability subgraph for one OpenAPI source."""
    source_id = descriptor.source_id
    raw = descriptor.raw

    nodes: list[ScgNode] = [self._source_node(descriptor)]
    edges: list[ScgEdge] = []

    for entity_nodes, entity_edges in self._iter_entities(source_id, raw):
        nodes.extend(entity_nodes)
        edges.extend(entity_edges)

    for cap_node, cap_edges in self._iter_capabilities(source_id, raw):
        nodes.append(cap_node)
        edges.extend(cap_edges)

    return StructureGraph(nodes=nodes, edges=edges)

SourceStructureProvider

Bases: Protocol

Parse one source type's descriptor into a normalized structure graph.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/base.py
27
28
29
30
31
32
33
34
35
36
@runtime_checkable
class SourceStructureProvider(Protocol):
    """Parse one source *type*'s descriptor into a normalized structure graph."""

    #: The ``SourceDescriptor.source_type`` this provider handles (dispatch key).
    source_type: str

    def build_structure(self, descriptor: SourceDescriptor) -> StructureGraph:
        """Parse ``descriptor.raw`` into a :class:`StructureGraph`. No network."""
        ...

build_structure(descriptor: SourceDescriptor) -> StructureGraph

Parse descriptor.raw into a :class:StructureGraph. No network.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/base.py
34
35
36
def build_structure(self, descriptor: SourceDescriptor) -> StructureGraph:
    """Parse ``descriptor.raw`` into a :class:`StructureGraph`. No network."""
    ...

StructureProviderRegistry

Dispatches a :class:SourceDescriptor to the provider for its type.

The declarative shell of the RML pattern: providers register by source_type and the registry routes build(descriptor) to the match. Construct via :meth:with_defaults for the built-in OpenAPI + MCP set.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/base.py
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
class StructureProviderRegistry:
    """Dispatches a :class:`SourceDescriptor` to the provider for its type.

    The declarative shell of the RML pattern: providers register by
    ``source_type`` and the registry routes ``build(descriptor)`` to the match.
    Construct via :meth:`with_defaults` for the built-in OpenAPI + MCP set.
    """

    def __init__(
        self, providers: list[SourceStructureProvider] | None = None
    ) -> None:
        """Build a registry, optionally seeded with *providers*."""
        self._providers: dict[str, SourceStructureProvider] = {}
        for provider in providers or []:
            self.register(provider)

    @classmethod
    def with_defaults(cls) -> StructureProviderRegistry:
        """Registry seeded with the schema-bearing built-in providers.

        Schemaless sources (``LlmStructureProvider``) are *not* auto-registered:
        they need an injected ``llm`` callable, so the caller wires them
        explicitly via :meth:`register`.
        """
        return cls([OpenApiStructureProvider(), McpToolListStructureProvider()])

    def register(self, provider: SourceStructureProvider) -> None:
        """Add or replace the provider for its ``source_type``."""
        self._providers[provider.source_type] = provider

    def providers(self) -> list[SourceStructureProvider]:
        """Return the registered providers (the parser's seam — public accessor)."""
        return list(self._providers.values())

    def for_type(self, source_type: str) -> SourceStructureProvider:
        """Return the provider for *source_type*, or raise ``KeyError``."""
        try:
            return self._providers[source_type]
        except KeyError:
            raise KeyError(f"No SCG structure provider for source_type={source_type!r}")

    def build(self, descriptor: SourceDescriptor) -> StructureGraph:
        """Dispatch *descriptor* to its provider and return the parsed graph."""
        return self.for_type(descriptor.source_type).build_structure(descriptor)

__init__(providers: list[SourceStructureProvider] | None = None) -> None

Build a registry, optionally seeded with providers.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/base.py
47
48
49
50
51
52
53
def __init__(
    self, providers: list[SourceStructureProvider] | None = None
) -> None:
    """Build a registry, optionally seeded with *providers*."""
    self._providers: dict[str, SourceStructureProvider] = {}
    for provider in providers or []:
        self.register(provider)

build(descriptor: SourceDescriptor) -> StructureGraph

Dispatch descriptor to its provider and return the parsed graph.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/base.py
80
81
82
def build(self, descriptor: SourceDescriptor) -> StructureGraph:
    """Dispatch *descriptor* to its provider and return the parsed graph."""
    return self.for_type(descriptor.source_type).build_structure(descriptor)

for_type(source_type: str) -> SourceStructureProvider

Return the provider for source_type, or raise KeyError.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/base.py
73
74
75
76
77
78
def for_type(self, source_type: str) -> SourceStructureProvider:
    """Return the provider for *source_type*, or raise ``KeyError``."""
    try:
        return self._providers[source_type]
    except KeyError:
        raise KeyError(f"No SCG structure provider for source_type={source_type!r}")

providers() -> list[SourceStructureProvider]

Return the registered providers (the parser's seam — public accessor).

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/base.py
69
70
71
def providers(self) -> list[SourceStructureProvider]:
    """Return the registered providers (the parser's seam — public accessor)."""
    return list(self._providers.values())

register(provider: SourceStructureProvider) -> None

Add or replace the provider for its source_type.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/base.py
65
66
67
def register(self, provider: SourceStructureProvider) -> None:
    """Add or replace the provider for its ``source_type``."""
    self._providers[provider.source_type] = provider

with_defaults() -> StructureProviderRegistry classmethod

Registry seeded with the schema-bearing built-in providers.

Schemaless sources (LlmStructureProvider) are not auto-registered: they need an injected llm callable, so the caller wires them explicitly via :meth:register.

Source code in packages/mewbo_graph/src/mewbo_graph/scg/providers/base.py
55
56
57
58
59
60
61
62
63
@classmethod
def with_defaults(cls) -> StructureProviderRegistry:
    """Registry seeded with the schema-bearing built-in providers.

    Schemaless sources (``LlmStructureProvider``) are *not* auto-registered:
    they need an injected ``llm`` callable, so the caller wires them
    explicitly via :meth:`register`.
    """
    return cls([OpenApiStructureProvider(), McpToolListStructureProvider()])

Clients (apps/)

  • API entry point: apps/mewbo_api/src/mewbo_api/backend.py
  • Console: apps/mewbo_console/ (React + Vite, connects via REST API)
  • CLI entry point: apps/mewbo_cli/src/mewbo_cli/cli_master.py

Home Assistant integration (mewbo_ha_conversation)

mewbo_ha_conversation.api

Mewbo API client.

MewboApiClient

Mewbo API Client.

Source code in apps/mewbo_ha_conversation/api.py
 32
 33
 34
 35
 36
 37
 38
 39
 40
 41
 42
 43
 44
 45
 46
 47
 48
 49
 50
 51
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
class MewboApiClient:
    """Mewbo API Client."""

    def __init__(
        self,
        base_url: str,
        api_key: str,
        timeout: int,
        session: aiohttp.ClientSession,
    ) -> None:
        """Initialize the API client.

        Args:
            base_url: Base URL for the Mewbo API.
            api_key: Key sent as the ``X-API-KEY`` header on every request; must
                match a token the server accepts (``api.master_token`` or a key
                minted via ``POST /api/keys``).
            timeout: Request timeout in seconds.
            session: Shared aiohttp client session.
        """
        self._base_url = base_url.rstrip("/")
        self._api_key = api_key
        self.timeout = timeout
        self._session = session

    async def async_get_heartbeat(self) -> bool:
        """Get heartbeat from the API.

        Returns:
            True when the service is considered healthy.
        """
        # TODO: Implement a heartbeat check
        return True

    async def async_get_models(self) -> str:
        """Get models from the API.

        Returns:
            JSON-serialized model list.
        """
        # TODO: This is monkey-patched for now
        response_data: ModelsResponse = {
            "models": [
                {
                    "name": "mewbo",
                    "modified_at": "2023-11-01T00:00:00.000000000-04:00",
                    "size": 0,
                    "digest": None,
                }
            ]
        }
        return json.dumps(response_data)

    async def async_generate(self, data: dict[str, Any] | None = None) -> MewboQueryResponse:
        """Generate a completion from the API.

        Args:
            data: Request payload including prompt and optional session ID.

        Returns:
            Parsed query response payload.

        Raises:
            ValueError: If prompt data is missing.
            ApiJsonError: If the API returns unexpected data.
        """
        if not data or "prompt" not in data:
            raise ValueError("Missing prompt in request data.")
        url_query = f"{self._base_url}/api/query"
        data_custom = {
            "query": str(data["prompt"]).strip(),
        }
        session_id = data.get("session_id") if isinstance(data, dict) else None
        if session_id:
            data_custom["session_id"] = session_id
        # Pass headers as None to use the default headers
        result = await self._mewbo_api_wrapper(
            method="post",
            url=url_query,
            data=data_custom,
            headers=None,
        )
        if isinstance(result, str):
            raise ApiJsonError("Unexpected text response from Mewbo API.")
        return result

    async def _mewbo_api_wrapper(
        self,
        method: str,
        url: str,
        data: dict[str, Any] | None = None,
        headers: dict[str, str] | None = None,
        decode_json: bool = True,
    ) -> MewboQueryResponse | str:
        """Perform an HTTP request to the Mewbo API.

        Args:
            method: HTTP method to use.
            url: Fully qualified request URL.
            data: Optional JSON payload to send.
            headers: Optional HTTP headers override.
            decode_json: Whether to parse JSON responses.

        Returns:
            Parsed response payload or raw text depending on decode_json.

        Raises:
            ApiJsonError: If the API returns an error payload.
            aiohttp.ClientResponseError: For non-2xx responses.
        """
        if headers is None:
            headers = {
                "accept": "application/json",
                "X-API-KEY": self._api_key,
                "X-Mewbo-Surface": "home-assistant",
                "Content-Type": "application/json",
            }
        async with async_timeout.timeout(self.timeout):
            response = await self._session.request(
                method=method,
                url=url,
                headers=headers,
                json=data,
            )
            response.raise_for_status()

            if decode_json:
                raw_data: dict[str, Any] = await response.json()
                if response.status == 404:
                    raise ApiJsonError(raw_data.get("error", "Unknown error"))
                task_result = str(raw_data.get("task_result", ""))
                response_data: MewboQueryResponse = {
                    "task_result": task_result,
                    "response": str(raw_data.get("response", task_result)),
                    "context": str(raw_data.get("context", task_result)),
                    "session_id": raw_data.get("session_id"),
                }
                LOGGER.debug("Response data: %s", response_data)
                return response_data
            else:
                LOGGER.debug("Fallback to text response")
                return await response.text()

__init__(base_url: str, api_key: str, timeout: int, session: aiohttp.ClientSession) -> None

Initialize the API client.

Parameters:

Name Type Description Default
base_url str

Base URL for the Mewbo API.

required
api_key str

Key sent as the X-API-KEY header on every request; must match a token the server accepts (api.master_token or a key minted via POST /api/keys).

required
timeout int

Request timeout in seconds.

required
session ClientSession

Shared aiohttp client session.

required
Source code in apps/mewbo_ha_conversation/api.py
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
def __init__(
    self,
    base_url: str,
    api_key: str,
    timeout: int,
    session: aiohttp.ClientSession,
) -> None:
    """Initialize the API client.

    Args:
        base_url: Base URL for the Mewbo API.
        api_key: Key sent as the ``X-API-KEY`` header on every request; must
            match a token the server accepts (``api.master_token`` or a key
            minted via ``POST /api/keys``).
        timeout: Request timeout in seconds.
        session: Shared aiohttp client session.
    """
    self._base_url = base_url.rstrip("/")
    self._api_key = api_key
    self.timeout = timeout
    self._session = session

async_generate(data: dict[str, Any] | None = None) -> MewboQueryResponse async

Generate a completion from the API.

Parameters:

Name Type Description Default
data dict[str, Any] | None

Request payload including prompt and optional session ID.

None

Returns:

Type Description
MewboQueryResponse

Parsed query response payload.

Raises:

Type Description
ValueError

If prompt data is missing.

ApiJsonError

If the API returns unexpected data.

Source code in apps/mewbo_ha_conversation/api.py
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
async def async_generate(self, data: dict[str, Any] | None = None) -> MewboQueryResponse:
    """Generate a completion from the API.

    Args:
        data: Request payload including prompt and optional session ID.

    Returns:
        Parsed query response payload.

    Raises:
        ValueError: If prompt data is missing.
        ApiJsonError: If the API returns unexpected data.
    """
    if not data or "prompt" not in data:
        raise ValueError("Missing prompt in request data.")
    url_query = f"{self._base_url}/api/query"
    data_custom = {
        "query": str(data["prompt"]).strip(),
    }
    session_id = data.get("session_id") if isinstance(data, dict) else None
    if session_id:
        data_custom["session_id"] = session_id
    # Pass headers as None to use the default headers
    result = await self._mewbo_api_wrapper(
        method="post",
        url=url_query,
        data=data_custom,
        headers=None,
    )
    if isinstance(result, str):
        raise ApiJsonError("Unexpected text response from Mewbo API.")
    return result

async_get_heartbeat() -> bool async

Get heartbeat from the API.

Returns:

Type Description
bool

True when the service is considered healthy.

Source code in apps/mewbo_ha_conversation/api.py
57
58
59
60
61
62
63
64
async def async_get_heartbeat(self) -> bool:
    """Get heartbeat from the API.

    Returns:
        True when the service is considered healthy.
    """
    # TODO: Implement a heartbeat check
    return True

async_get_models() -> str async

Get models from the API.

Returns:

Type Description
str

JSON-serialized model list.

Source code in apps/mewbo_ha_conversation/api.py
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
async def async_get_models(self) -> str:
    """Get models from the API.

    Returns:
        JSON-serialized model list.
    """
    # TODO: This is monkey-patched for now
    response_data: ModelsResponse = {
        "models": [
            {
                "name": "mewbo",
                "modified_at": "2023-11-01T00:00:00.000000000-04:00",
                "size": 0,
                "digest": None,
            }
        ]
    }
    return json.dumps(response_data)

MewboQueryResponse

Bases: TypedDict

Schema for the main query response.

Source code in apps/mewbo_ha_conversation/api.py
23
24
25
26
27
28
29
class MewboQueryResponse(TypedDict):
    """Schema for the main query response."""

    task_result: str
    response: str
    context: str
    session_id: str | None

ModelsResponse

Bases: TypedDict

Schema for the models list endpoint response.

Source code in apps/mewbo_ha_conversation/api.py
17
18
19
20
class ModelsResponse(TypedDict):
    """Schema for the models list endpoint response."""

    models: list[dict[str, Any]]

mewbo_ha_conversation.config_flow

Adds config flow for Mewbo.

MewboConfigFlow

Bases: ConfigFlow

Handle a config flow for Mewbo Conversation. Handles UI wizard.

Source code in apps/mewbo_ha_conversation/config_flow.py
 52
 53
 54
 55
 56
 57
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
class MewboConfigFlow(config_entries.ConfigFlow, domain=DOMAIN):  # type: ignore[call-arg]
    """Handle a config flow for Mewbo Conversation. Handles UI wizard."""

    VERSION = 1
    client: MewboApiClient

    async def async_step_user(self, user_input: dict[str, Any] | None = None) -> FlowResult:
        """Handle the initial config flow step.

        Args:
            user_input: Submitted form data, if available.

        Returns:
            FlowResult for the configuration step.
        """
        if user_input is None:
            return self.async_show_form(step_id="user", data_schema=STEP_USER_DATA_SCHEMA)

        # Search for duplicates with the same CONF_BASE_URL value.
        for existing_entry in self._async_current_entries(include_ignore=False):
            if existing_entry.data.get(CONF_BASE_URL) == user_input[CONF_BASE_URL]:
                return self.async_abort(reason="already_configured")

        errors: dict[str, str] = {}
        try:
            self.client = MewboApiClient(
                base_url=cv.url_no_path(user_input[CONF_BASE_URL]),
                api_key=user_input[CONF_API_KEY],
                timeout=user_input[CONF_TIMEOUT],
                session=async_create_clientsession(self.hass),
            )
            response = await self.client.async_get_heartbeat()
            if not response:
                raise vol.Invalid("Invalid Mewbo server")
        # except vol.Invalid:
        #     errors["base"] = "invalid_url"
        # except ApiTimeoutError:
        #     errors["base"] = "timeout_connect"
        # except ApiCommError:
        #     errors["base"] = "cannot_connect"
        # except ApiClientError as exception:
        #     LOGGER.exception("Unexpected exception: %s", exception)
        #     errors["base"] = "unknown"
        except Exception as exception:
            LOGGER.exception("Unexpected exception: %s", exception)
            errors["base"] = "unknown"
        else:
            return self.async_create_entry(
                title=f"Mewbo - {user_input[CONF_BASE_URL]}",
                data={
                    CONF_BASE_URL: user_input[CONF_BASE_URL],
                    CONF_API_KEY: user_input[CONF_API_KEY],
                },
                options={CONF_TIMEOUT: user_input[CONF_TIMEOUT]},
            )

        return self.async_show_form(
            step_id="user", data_schema=STEP_USER_DATA_SCHEMA, errors=errors
        )

    @staticmethod
    def async_get_options_flow(
        config_entry: config_entries.ConfigEntry,
    ) -> config_entries.OptionsFlow:
        """Create the options flow.

        Args:
            config_entry: Existing config entry to edit.

        Returns:
            Options flow handler.
        """
        return MewboOptionsFlow(config_entry)

async_get_options_flow(config_entry: config_entries.ConfigEntry) -> config_entries.OptionsFlow staticmethod

Create the options flow.

Parameters:

Name Type Description Default
config_entry ConfigEntry

Existing config entry to edit.

required

Returns:

Type Description
OptionsFlow

Options flow handler.

Source code in apps/mewbo_ha_conversation/config_flow.py
112
113
114
115
116
117
118
119
120
121
122
123
124
@staticmethod
def async_get_options_flow(
    config_entry: config_entries.ConfigEntry,
) -> config_entries.OptionsFlow:
    """Create the options flow.

    Args:
        config_entry: Existing config entry to edit.

    Returns:
        Options flow handler.
    """
    return MewboOptionsFlow(config_entry)

async_step_user(user_input: dict[str, Any] | None = None) -> FlowResult async

Handle the initial config flow step.

Parameters:

Name Type Description Default
user_input dict[str, Any] | None

Submitted form data, if available.

None

Returns:

Type Description
FlowResult

FlowResult for the configuration step.

Source code in apps/mewbo_ha_conversation/config_flow.py
 58
 59
 60
 61
 62
 63
 64
 65
 66
 67
 68
 69
 70
 71
 72
 73
 74
 75
 76
 77
 78
 79
 80
 81
 82
 83
 84
 85
 86
 87
 88
 89
 90
 91
 92
 93
 94
 95
 96
 97
 98
 99
100
101
102
103
104
105
106
107
108
109
110
async def async_step_user(self, user_input: dict[str, Any] | None = None) -> FlowResult:
    """Handle the initial config flow step.

    Args:
        user_input: Submitted form data, if available.

    Returns:
        FlowResult for the configuration step.
    """
    if user_input is None:
        return self.async_show_form(step_id="user", data_schema=STEP_USER_DATA_SCHEMA)

    # Search for duplicates with the same CONF_BASE_URL value.
    for existing_entry in self._async_current_entries(include_ignore=False):
        if existing_entry.data.get(CONF_BASE_URL) == user_input[CONF_BASE_URL]:
            return self.async_abort(reason="already_configured")

    errors: dict[str, str] = {}
    try:
        self.client = MewboApiClient(
            base_url=cv.url_no_path(user_input[CONF_BASE_URL]),
            api_key=user_input[CONF_API_KEY],
            timeout=user_input[CONF_TIMEOUT],
            session=async_create_clientsession(self.hass),
        )
        response = await self.client.async_get_heartbeat()
        if not response:
            raise vol.Invalid("Invalid Mewbo server")
    # except vol.Invalid:
    #     errors["base"] = "invalid_url"
    # except ApiTimeoutError:
    #     errors["base"] = "timeout_connect"
    # except ApiCommError:
    #     errors["base"] = "cannot_connect"
    # except ApiClientError as exception:
    #     LOGGER.exception("Unexpected exception: %s", exception)
    #     errors["base"] = "unknown"
    except Exception as exception:
        LOGGER.exception("Unexpected exception: %s", exception)
        errors["base"] = "unknown"
    else:
        return self.async_create_entry(
            title=f"Mewbo - {user_input[CONF_BASE_URL]}",
            data={
                CONF_BASE_URL: user_input[CONF_BASE_URL],
                CONF_API_KEY: user_input[CONF_API_KEY],
            },
            options={CONF_TIMEOUT: user_input[CONF_TIMEOUT]},
        )

    return self.async_show_form(
        step_id="user", data_schema=STEP_USER_DATA_SCHEMA, errors=errors
    )

MewboOptionsFlow

Bases: OptionsFlow

Mewbo config flow options handler.

Source code in apps/mewbo_ha_conversation/config_flow.py
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
class MewboOptionsFlow(config_entries.OptionsFlow):
    """Mewbo config flow options handler."""

    def __init__(self, config_entry: config_entries.ConfigEntry) -> None:
        """Initialize options flow.

        Args:
            config_entry: Config entry to manage.
        """
        self.config_entry = config_entry
        self.options = dict(config_entry.options)

    async def async_step_init(self, user_input: dict[str, Any] | None = None) -> FlowResult:
        """Show the options menu.

        Args:
            user_input: Submitted form data, if available.

        Returns:
            FlowResult for the options menu.
        """
        return self.async_show_menu(step_id="init", menu_options=MENU_OPTIONS)

    async def async_step_all_set(self, user_input: dict[str, Any] | None = None) -> FlowResult:
        """Handle the "all_set" options step.

        Args:
            user_input: Submitted form data, if available.

        Returns:
            FlowResult for the options menu.
        """
        return self.async_show_menu(step_id="init", menu_options=MENU_OPTIONS)

    async def async_step_general_config(
        self, user_input: dict[str, Any] | None = None
    ) -> FlowResult:
        """Handle the general configuration step.

        Args:
            user_input: Submitted form data, if available.

        Returns:
            FlowResult for the options menu.
        """
        return self.async_show_menu(step_id="init", menu_options=MENU_OPTIONS)

    async def async_step_prompt_system(
        self, user_input: dict[str, Any] | None = None
    ) -> FlowResult:
        """Handle the prompt system configuration step.

        Args:
            user_input: Submitted form data, if available.

        Returns:
            FlowResult for the options menu.
        """
        return self.async_show_menu(step_id="init", menu_options=MENU_OPTIONS)

    async def async_step_model_config(self, user_input: dict[str, Any] | None = None) -> FlowResult:
        """Handle the model configuration step.

        Args:
            user_input: Submitted form data, if available.

        Returns:
            FlowResult for the options menu.
        """
        return self.async_show_menu(step_id="init", menu_options=MENU_OPTIONS)

__init__(config_entry: config_entries.ConfigEntry) -> None

Initialize options flow.

Parameters:

Name Type Description Default
config_entry ConfigEntry

Config entry to manage.

required
Source code in apps/mewbo_ha_conversation/config_flow.py
130
131
132
133
134
135
136
137
def __init__(self, config_entry: config_entries.ConfigEntry) -> None:
    """Initialize options flow.

    Args:
        config_entry: Config entry to manage.
    """
    self.config_entry = config_entry
    self.options = dict(config_entry.options)

async_step_all_set(user_input: dict[str, Any] | None = None) -> FlowResult async

Handle the "all_set" options step.

Parameters:

Name Type Description Default
user_input dict[str, Any] | None

Submitted form data, if available.

None

Returns:

Type Description
FlowResult

FlowResult for the options menu.

Source code in apps/mewbo_ha_conversation/config_flow.py
150
151
152
153
154
155
156
157
158
159
async def async_step_all_set(self, user_input: dict[str, Any] | None = None) -> FlowResult:
    """Handle the "all_set" options step.

    Args:
        user_input: Submitted form data, if available.

    Returns:
        FlowResult for the options menu.
    """
    return self.async_show_menu(step_id="init", menu_options=MENU_OPTIONS)

async_step_general_config(user_input: dict[str, Any] | None = None) -> FlowResult async

Handle the general configuration step.

Parameters:

Name Type Description Default
user_input dict[str, Any] | None

Submitted form data, if available.

None

Returns:

Type Description
FlowResult

FlowResult for the options menu.

Source code in apps/mewbo_ha_conversation/config_flow.py
161
162
163
164
165
166
167
168
169
170
171
172
async def async_step_general_config(
    self, user_input: dict[str, Any] | None = None
) -> FlowResult:
    """Handle the general configuration step.

    Args:
        user_input: Submitted form data, if available.

    Returns:
        FlowResult for the options menu.
    """
    return self.async_show_menu(step_id="init", menu_options=MENU_OPTIONS)

async_step_init(user_input: dict[str, Any] | None = None) -> FlowResult async

Show the options menu.

Parameters:

Name Type Description Default
user_input dict[str, Any] | None

Submitted form data, if available.

None

Returns:

Type Description
FlowResult

FlowResult for the options menu.

Source code in apps/mewbo_ha_conversation/config_flow.py
139
140
141
142
143
144
145
146
147
148
async def async_step_init(self, user_input: dict[str, Any] | None = None) -> FlowResult:
    """Show the options menu.

    Args:
        user_input: Submitted form data, if available.

    Returns:
        FlowResult for the options menu.
    """
    return self.async_show_menu(step_id="init", menu_options=MENU_OPTIONS)

async_step_model_config(user_input: dict[str, Any] | None = None) -> FlowResult async

Handle the model configuration step.

Parameters:

Name Type Description Default
user_input dict[str, Any] | None

Submitted form data, if available.

None

Returns:

Type Description
FlowResult

FlowResult for the options menu.

Source code in apps/mewbo_ha_conversation/config_flow.py
187
188
189
190
191
192
193
194
195
196
async def async_step_model_config(self, user_input: dict[str, Any] | None = None) -> FlowResult:
    """Handle the model configuration step.

    Args:
        user_input: Submitted form data, if available.

    Returns:
        FlowResult for the options menu.
    """
    return self.async_show_menu(step_id="init", menu_options=MENU_OPTIONS)

async_step_prompt_system(user_input: dict[str, Any] | None = None) -> FlowResult async

Handle the prompt system configuration step.

Parameters:

Name Type Description Default
user_input dict[str, Any] | None

Submitted form data, if available.

None

Returns:

Type Description
FlowResult

FlowResult for the options menu.

Source code in apps/mewbo_ha_conversation/config_flow.py
174
175
176
177
178
179
180
181
182
183
184
185
async def async_step_prompt_system(
    self, user_input: dict[str, Any] | None = None
) -> FlowResult:
    """Handle the prompt system configuration step.

    Args:
        user_input: Submitted form data, if available.

    Returns:
        FlowResult for the options menu.
    """
    return self.async_show_menu(step_id="init", menu_options=MENU_OPTIONS)

mewbo_ha_conversation.const

Constants for mewbo_conversation.

mewbo_ha_conversation.coordinator

DataUpdateCoordinator for mewbo_conversation.

MewboDataUpdateCoordinator

Bases: DataUpdateCoordinator

Class to manage fetching data from the API.

Source code in apps/mewbo_ha_conversation/coordinator.py
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
class MewboDataUpdateCoordinator(DataUpdateCoordinator):
    """Class to manage fetching data from the API."""

    config_entry: ConfigEntry

    def __init__(
        self,
        hass: HomeAssistant,
        client: MewboApiClient,
    ) -> None:
        """Initialize the coordinator.

        Args:
            hass: Home Assistant core instance.
            client: API client for Mewbo.
        """
        self.client = client
        super().__init__(
            hass=hass,
            logger=LOGGER,
            name=DOMAIN,
            update_interval=timedelta(minutes=5),
        )

    async def _async_update_data(self) -> bool:
        """Update data via library.

        Returns:
            True when the heartbeat check succeeds.

        Raises:
            UpdateFailed: If the API heartbeat fails.
        """
        try:
            return await self.client.async_get_heartbeat()
        except ApiClientError as exception:
            raise UpdateFailed(exception) from exception

__init__(hass: HomeAssistant, client: MewboApiClient) -> None

Initialize the coordinator.

Parameters:

Name Type Description Default
hass HomeAssistant

Home Assistant core instance.

required
client MewboApiClient

API client for Mewbo.

required
Source code in apps/mewbo_ha_conversation/coordinator.py
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
def __init__(
    self,
    hass: HomeAssistant,
    client: MewboApiClient,
) -> None:
    """Initialize the coordinator.

    Args:
        hass: Home Assistant core instance.
        client: API client for Mewbo.
    """
    self.client = client
    super().__init__(
        hass=hass,
        logger=LOGGER,
        name=DOMAIN,
        update_interval=timedelta(minutes=5),
    )

mewbo_ha_conversation.exceptions

The exceptions used by Extended OpenAI Conversation.

ApiClientError

Bases: HomeAssistantError

Exception to indicate a general API error.

Source code in apps/mewbo_ha_conversation/exceptions.py
6
7
class ApiClientError(HomeAssistantError):
    """Exception to indicate a general API error."""

ApiCommError

Bases: ApiClientError

Exception to indicate a communication error.

Source code in apps/mewbo_ha_conversation/exceptions.py
10
11
class ApiCommError(ApiClientError):
    """Exception to indicate a communication error."""

ApiJsonError

Bases: ApiClientError

Exception to indicate an error with json response.

Source code in apps/mewbo_ha_conversation/exceptions.py
14
15
class ApiJsonError(ApiClientError):
    """Exception to indicate an error with json response."""

ApiTimeoutError

Bases: ApiClientError

Exception to indicate a timeout error.

Source code in apps/mewbo_ha_conversation/exceptions.py
18
19
class ApiTimeoutError(ApiClientError):
    """Exception to indicate a timeout error."""

mewbo_ha_conversation.helpers

Helper functions for Mewbo.

ExposedEntity

Bases: TypedDict

Typed representation of a Home Assistant entity exposed to conversation.

Source code in apps/mewbo_ha_conversation/helpers.py
13
14
15
16
17
18
19
class ExposedEntity(TypedDict):
    """Typed representation of a Home Assistant entity exposed to conversation."""

    entity_id: str
    name: str
    state: str
    aliases: list[str]

get_exposed_entities(hass: HomeAssistant) -> list[ExposedEntity]

Return exposed entities.

Parameters:

Name Type Description Default
hass HomeAssistant

Home Assistant core instance.

required

Returns:

Type Description
list[ExposedEntity]

List of exposed entities and their metadata.

Source code in apps/mewbo_ha_conversation/helpers.py
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
def get_exposed_entities(hass: HomeAssistant) -> list[ExposedEntity]:
    """Return exposed entities.

    Args:
        hass: Home Assistant core instance.

    Returns:
        List of exposed entities and their metadata.
    """
    hass_entity = entity_registry.async_get(hass)
    exposed_entities: list[ExposedEntity] = []

    for state in hass.states.async_all():
        if async_should_expose(hass, CONVERSATION_DOMAIN, state.entity_id):
            entity = hass_entity.async_get(state.entity_id)
            exposed_entities.append(
                {
                    "entity_id": state.entity_id,
                    "name": state.name,
                    "state": state.state,
                    "aliases": entity.aliases if entity else [],
                }
            )

    return exposed_entities