MCPcopy Create free account
hub / github.com/Qinbf/groundmap / build_outline_data

Function build_outline_data

scripts/postprocess.py:219–278  ·  view source on GitHub ↗

基于已加锚点的 markdown 文本构建 outline.json 数据结构。 输入约定:text_with_anchors 应当已经被 add_anchors 处理过(每个块末尾 带 ` ^h-/^p-/...` 锚点)。本函数只切一次块,从行末回填 anchor,构建 嵌套 section 树。 若提供 previous_outline,会按 anchor(失配再按 title)把旧 agent_summary 合并进 新 outline——见 _restore_summaries。 **谁会传 previous_outli

(
    text_with_anchors: str,
    doc_path: str,
    *,
    previous_outline: dict | None = None,
)

Source from the content-addressed store, hash-verified

217
218
219def build_outline_data(
220 text_with_anchors: str,
221 doc_path: str,
222 *,
223 previous_outline: dict | None = None,
224) -> dict:
225 """
226 基于已加锚点的 markdown 文本构建 outline.json 数据结构。
227
228 输入约定:text_with_anchors 应当已经被 add_anchors 处理过(每个块末尾
229 带 ` ^h-/^p-/...` 锚点)。本函数只切一次块,从行末回填 anchor,构建
230 嵌套 section 树。
231
232 若提供 previous_outline,会按 anchor(失配再按 title)把旧 agent_summary 合并进
233 新 outline——见 _restore_summaries。
234
235 **谁会传 previous_outline(口径说明)**:只有 convert.py 的 raw re-convert 路径
236 传它,处理对象是**不可变的 raw 源**(CLAUDE.md 视 raw 为事实来源,仅当源文件真
237 改才重转,且增量转换跳过未变文件,故 re-convert 罕见)。k.py 的 wiki 读路径
238 (load_or_build_outline)**不传** previous_outline——wiki 页可被 agent/人直接
239 频繁编辑、整页换主题的风险高,过期即整体丢弃不合并。两条路径口径不同是**按对象
240 特性的有意取舍**,非疏忽。
241 已知残留:heading 锚点 hash 仅基于标题文本(见 :121),故"同标题换正文"在
242 re-convert 路径仍可能复活旧摘要;因 raw 不可变 + 重转罕见,该残留风险低、接受。
243
244 注:早期版本曾要求传入 add_anchors 返回的 blocks 列表,但实测 blocks
245 的 char_start/char_end 是基于"加锚前 body"的偏移,加锚后会漂移;保留
246 blocks 参数反而误导。本函数现在只接受 text_with_anchors,单次切块即可。
247 """
248 fm, body = strip_frontmatter(text_with_anchors)
249 fm_offset = len(fm)
250
251 # 切块 + 从行末回填 anchor(一次切块,一次扫描,无重复工作)
252 parsed_blocks = split_blocks(body)
253 for blk in parsed_blocks:
254 if blk.kind == "hr":
255 continue
256 m = _BLOCK_ANCHOR_TAIL_RE.search(blk.text)
257 if m:
258 blk.anchor = m.group(1)
259 # body 偏移 → 完整文件偏移
260 blk.char_start += fm_offset
261 blk.char_end += fm_offset
262
263 total_chars = len(text_with_anchors)
264 outline = build_outline(parsed_blocks, total_chars)
265 sections_dicts = [section_to_dict(s) for s in outline["sections"]]
266
267 if previous_outline:
268 old_summaries = _collect_summaries(previous_outline.get("sections", []))
269 if old_summaries:
270 _restore_summaries(sections_dicts, old_summaries)
271
272 return {
273 "doc_path": doc_path,
274 "doc_chars": total_chars,
275 "doc_paragraphs": outline["paragraphs_count"],
276 "generated_at": date.today().isoformat(),

Calls 6

strip_frontmatterFunction · 0.90
split_blocksFunction · 0.90
build_outlineFunction · 0.90
section_to_dictFunction · 0.90
_collect_summariesFunction · 0.85
_restore_summariesFunction · 0.85