Toward a Trustworthy and Accessible Scientific Data Workflow Platform with StreamCI
Scientific research workflows increasingly involve not only structured streaming data but also raw artifacts and derived products, requiring platforms that provide trustworthy data protection and accessible interfaces beyond simple ingestion and storage. We present extensions to StreamCI, a cloud-based streaming data management platform, that evolve it into an end-to-end scientific data workflow platform. Key advancements include expanded data lifecycle support (blob ingestion, raw data refinement, and derived data generation), a migration from RabbitMQ to Apache Kafka for scalable dataflow infrastructure, automated backup mechanisms for data reliability, and a researcher-friendly web portal paired with a Python client API library (PyStreamCI). We demonstrate these capabilities through end-to-end workflows from active research use cases and report on operational experience following deployment on Purdue University’s Anvil cloud infrastructure. In production, these extensions enable researchers across domains—such as building energy sustainability, pavement condition monitoring, and precision audiology—to manage complete data workflows, from raw artifacts to analysis-ready products, without requiring backend expertise.