Skip to content

fix bugs - #934

Merged
alcholiclg merged 15 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness
Aug 12, 2026
Merged

fix bugs#934
alcholiclg merged 15 commits into
modelscope:mainfrom
alcholiclg:fix/runtime-robustness

Conversation

@alcholiclg

@alcholiclg alcholiclg commented Aug 10, 2026

Copy link
Copy Markdown
Collaborator

Change Summary

1.Improved vector memory reliability and usability: fixed provider/protocol resolution, added project-scoped model configuration, moved ingestion off the chat critical path, and surfaced ingestion health, configuration errors, and safe rebuild actions in WebUI.

2.Fixed parallel tool authorization and status reporting: WebUI can now display concurrent permission requests, while each tool reports completion and duration independently without waiting for the entire batch.

3.Fixed custom provider routing to preserve explicitly configured service names, credentials, and endpoints instead of incorrectly inferring a built-in provider from the model name.

Related issue number

Checklist

  • The pull request title is a good summary of the changes - it will be used in the changelog
  • Unit tests for the changes exist
  • Run pre-commit install and pre-commit run --all-files before git commit, and passed lint check.
  • Documentation reflects the changes where applicable

The orchestrator now owns the write discipline around a backend:
- schedule_add() runs the extraction-LLM + embedding cost (seconds) in a
  background task; flush_pending() is the teardown barrier so the last
  write is never dropped, and an inline fallback keeps writes when no
  loop is running.
- retrieval/ingestion/flush serialize on one per-store asyncio lock
  (embedded qdrant underneath is lock-free single-client code).
- a content-hash delta ledger (<base_dir>/ingest_state.json) makes each
  ingest send only messages the store has not seen; hashes are recorded
  only after a confirmed write, so a failed ingest retries naturally.
- ingest_status reports the last outcome (state/count/error/pending) so
  a UI can show memory working instead of silence.

Mem0Backend: per-turn retrieval cache (rounds 2..N of a tool-calling
turn reuse round 1's search instead of paying an embedding round-trip
each), on_messages returns the event count and propagates failures --
the orchestrator is the swallow-and-report layer now and needs the
exception to keep failed messages un-marked for retry.
…-side close

- add_memory(add_after_step) now fires only when a round closes the turn
  (assistant reply with no tool calls) and dispatches through the
  backend's schedule_add when available: tool rounds are intermediate
  state, and ingesting every round cost O(rounds x history) extraction
  calls where the closing ingest covers the whole turn.
- an interrupted round advances the ingest ledger WITHOUT ingesting
  (mark_ingested): a half-finished answer is not durable conversational
  truth and must not be swept into the next turn's delta.
- cleanup_tools drains scheduled ingestion (flush only -- memory
  instances are shared across agents of one store, so closing here would
  yank the store from a sibling agent); the new
  SharedMemoryManager.close_matching(base_dir) is the owner-of-last-
  resort that actually closes instances and releases the embedded
  store's exclusive file lock.
The number of recalled memories injected per turn was hardcoded twice
(search default 20, then a [:10] formatting slice). MemoryConfig gains
recall_top_k (default 10, read from the unified_memory node) and the
mem0 adapter threads it through search and formatting — consumers can
now size recall to their context budget.
@alcholiclg
alcholiclg merged commit 2a263f9 into modelscope:main Aug 12, 2026
1 check failed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants