Skip to content

[Summary] GraphGen Roadmap #49

Description

@ChenZiHong-Gavin

Backgraound

To establish GraphGen as an essential tool for training and evaluation data synthesis, its development roadmap focuses on two core pillars: implementing a robust, multi-dimensional data quality assessment and filtering system to ensure the reliability of generated knowledge graphs, and expanding its architecture to support multi-modal and multi-omics data inputs.

If you'd like to work on one of these tasks, please comment below to claim it and create an issue for the feature you'll be implementing.

Features

1 GraphGen Framework

2 Multi-Modal & Multi-Omics

  • 🧬 Define ImageNode, AudioNode, ProteinNode, etc.
  • 👁️‍🗨️ Vision–language fusion extraction: use open VLMs to generate "image–caption–entity" triples and write them into the graph: feat: add vqa pipeline #69
  • 🧪 Multi-omics extraction: process genomics/transcriptomics/proteomics with automatic node-property alignment

3 Data Quality & Curation

  • 📊 Multi-dimensional quality metrics with a unified scoring API
  • 💓 Graph-quality assessment similar to KGHeartBeat: Feat: kg evaluator #135
  • 🎯 One-click export of high-quality sub-graphs and high-quality data
  • ⚙️ Configurable pipeline: entity disambiguation, fact verification, redundancy removal, schema validation

4 Graph Construction

  • 🚀 Incremental & resumable construction: Feat: add trace info & task storage #168
  • ⚖️ Automatic data ratio optimization: dynamically adjust the mixing ratio of different data based on quality scores and training feedback to optimize model performance

5 Engineering

6 Community Detection & Data Synthesis

  • 🔎 Apply multiple community-detection algorithms; generate data from communities and provide typical samples plus visualizations
  • 🧠 Community summary → CoT data: use community summaries as few-shot examples to synthesize high-quality chain-of-thought data
  • 💬 Multi-turn dialogue synthesis: random-walk sampling → multi-turn Q&A while maintaining context consistency
  • 📈 Complexity grading for curriculum learning
  • 🕵️‍♂️ Support comparison with baselines

7 UX, Docs & Community

  • 📦 Streamlined pip install and usage
  • 📓 Jupyter tutorial suite
  • 📚 Comprehensive documentation
  • 🗃️ Data & user case library
  • 🤝 Contributor guide & roadmap: clear labels, branching strategy, PR template, code of conduct
  • 🌐 More user-friendly web interface

8 Others

  • 📝 More standardized prompt & post-processing management; post-processing should be bound to prompts
  • 🌐 Improve online connectivity
  • 🔗 Enhanced coreference resolution during chunking

Further feature ideas are welcome—feel free to suggest and join the plan!

Activity

  1. mdjhacker commented on Dec 24, 2025

    @mdjhacker
    Contributor

    Would it be possible to consider adding checkpoint support, so that pipelines can resume from an intermediate state after interruptions (e.g., node failures, manual stops, or preemption)?

    Although Ray already provides lineage-based retry and fault recovery, in practice, for long-running GraphGen workloads—often spanning multiple hours or even days, especially when combining KG construction with LLM generation—restarting the entire pipeline from scratch can still be quite expensive🤔.

  2. ChenZiHong-Gavin commented on Dec 24, 2025

    @ChenZiHong-Gavin
    CollaboratorAuthor

    Would it be possible to consider adding checkpoint support, so that pipelines can resume from an intermediate state after interruptions (e.g., node failures, manual stops, or preemption)?

    Although Ray already provides lineage-based retry and fault recovery, in practice, for long-running GraphGen workloads—often spanning multiple hours or even days, especially when combining KG construction with LLM generation—restarting the entire pipeline from scratch can still be quite expensive🤔.

    yeah, I'm working on it. This fearure will be implemented soon.

  3. ChenZiHong-Gavin commented on Jan 29, 2026

    @ChenZiHong-Gavin
    CollaboratorAuthor

    @mdjhacker Hi, GraphGen now supports resuming generation. See #168.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

documentationImprovements or additions to documentation

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions