Modern research organizations generate large volumes of data through scientific instruments, computational workflows, and laboratory ecosystem e.g. laboratory notebooks. These data must be transferred reliably from the point of origin to the researchers, teams, and infrastructure services that need to analyse, process, share, and eventually publish them. In practice, this process often remains fragmented: data are copied manually, stored on instrument computers, transferred via external drives, or kept in local laboratory folders without consistent metadata, provenance, or access control.
This creates both operational and scientific risks. Research data are valuable institutional assets, not merely temporary files owned by individual users. They should therefore be stored in a managed organizational environment where they can be protected, described, shared, and connected to the context of their creation. Such an environment should capture technical metadata from instruments, link datasets to users, projects, facilities, and experiments, and preserve provenance information needed for traceability and later reuse. A central data management layer also enables more advanced workflows. Active datasets can be shared with collaborators, transferred to computational environments for analysis, and prepared for publication in domain-specific repositories once they have been curated and used, for example, as evidence supporting a scientific publication. At the same time, institutional storage must support lifecycle management: actively used data should remain accessible, published datasets may be removed from working storage, and unpublished datasets may require archiving or retention for a defined period.
The motivation for this system is therefore to replace manual and inconsistent data handling with an integrated information system for managing research datasets across their active lifecycle. The system connects scientific instruments, storage infrastructure, metadata management, access control, computational workflows, and later data publication into a coherent institutional process.
The primary users are core facility operators, laboratory managers, data stewards, and infrastructure administrators. The solution gives them operational control over active research data: acquisition from instruments, storage orchestration, access management, metadata capture, and preparation for later publication. Researchers and principal investigators benefit through more reliable access to measured data, clearer dataset context, and easier reuse or sharing.
The DAREG Ecosystem is a data management platform for research datasets produced by scientific instruments. It supports the active phase of the data lifecycle: acquisition from instrument environments, dataset registration, metadata capture, storage orchestration, access management, sharing, and preparation for later processing, publication, or preservation. Its purpose is to replace manual and fragmented data handling with a structured workflow connecting laboratory instruments, institutional storage, computational environments, and repository services.
Onedata – external component provided by e-INFRA CZ. Research data management and storage layer. Onedata provides the technical basis for storing, synchronizing, sharing, and accessing large and distributed research datasets produced by core facilities. In this role, it serves not only as a storage backend for local workflows, but also as a reusable layer for other institutional services, including integration scenarios (such as VUT Booking). More information: https://onedata.org
DAREG API – the backend layer responsible for the data model, metadata templates, access control, dataset and experiment records, and integration with external services such as Onedata, instrument reservation systems, and electronic laboratory notebooks. Repository: https://github.com/sb-ncbr/dareg-api. DAREG Web – the browser-based graphical interface for researchers, facility staff, and administrators. It allows users to register, browse, annotate, manage, and search datasets using facility-specific metadata structures. Recent development also introduced a reusable dynamic search component that supports structured, metadata-based filtering over evolving schemas, improving the findability of datasets in growing registries. Repository: https://github.com/sb-ncbr/dareg-ui.
DAREG Lab Client is the desktop component deployed close to laboratory instruments. It supports automated acquisition of data from instrument computers, synchronization of measured data into Onedata, and automatic registration of corresponding dataset and experiment records in DAREG. The client reduces manual copying, manual upload, and ad hoc folder handling. It also improves traceability by linking data transfer to dataset metadata and by supporting integrity checks during synchronization. Repository: https://github.com/sb-ncbr/dareg-lab.



