ReLaG: Open Source Framework for Relation-Aware Data Splitting Released
ReLaG provides a scalable, modality-agnostic framework for generating independent train–test splits in datasets with latent sample relations, addressing overoptimistic generalization caused by standard random splits. It is open source and available for installation via pip.
When datasets contain related samples—such as in biochemical studies—standard random train–test splits can create overlapping sample groups, leading to non-independent evaluation and overoptimistic results. ReLaG is a newly released, open-source framework addressing this issue by generating independent splits based on inferred sample relations.

Key Features
- Models latent sample relations using a hierarchical latent-variable process.
- Infers groups of related samples via proximity graphs and community detection.
- Creates independent train–test splits better aligned with real-world deployment.
- Scalable to large datasets—enables splitting at previously impractical sizes.
- Includes a label-free method to match splitting resolution with production data.
- Provides a quick estimate of effective dataset size for diversity-aware scaling.
For developers working on applications where relatedness among samples could bias train/test splits—especially in chemical, biological, or similar datasets—ReLaG offers a practical, scalable solution. It is modality-agnostic and matches the effectiveness of existing relation-aware methods while being substantially more scalable.
