A Data Engineering Environment for Distributed Research Teams at Tessa Therapeutics
- A full pharma data engineering environment for distributed research teams
- Repeat engagement across three iterations on Google Cloud
- Data, imaging, and model-tooling workstreams delivered
- Customer
- Tessa Therapeutics
- Year
Tessa Therapeutics was a Singapore biotech developing cell therapies for cancer. Its research teams worked from different locations and handled large genomic and imaging datasets, so they needed one shared place for data, pipelines, and compute. Sakura Sky designed and delivered it on Google Cloud, and Tessa brought us back for two further iterations. The third added JupyterLab as a common workspace for Tessa’s distributed teams. Today this kind of work is part of our Data & AI practice.
The brief
The engagement covered three workstreams:
- Big Cloud & Data: data pipelines, databases, and web-based visualisation for large-scale genome analysis.
- Embedded Systems & Tools: deep learning libraries for image processing and segmentation, with automated cloud storage for data from embedded devices.
- AI Library Toolkit: a version-controlled code base with workflow automation and scheduling for interdependent analysis models.
Tessa also needed to combine its private research data with public reference datasets, and every connection to its cloud assets had to be encrypted and authenticated, through a VPN or a bastion host.
Iterations 1 and 2: the data engineering foundation
The first two iterations built the foundation on open-source components running on Google Cloud. Kubernetes hosted separate namespaces for ingestion, analytics, and orchestration. Apache Kafka carried near-real-time feeds, Apache Airflow scheduled and monitored the analysis workflows, and Apache Beam handled batch and stream processing in Python. Cloud Storage, BigQuery, and PostgreSQL held the data, and HashiCorp Vault managed secrets.
Iteration 3: JupyterLab for a Python team
Tessa’s research engineers worked in Python, and JupyterLab was a common choice for that work at the time. The third iteration gave them JupyterLab on Google Cloud, with the pipelines and storage from the first two iterations still in service. From inside their notebooks, the engineers could use worker instances, Cloud Storage, and the gcloud SDK. Access was limited to trusted users over the Tessa VPN. Sakura Sky also built a number of custom components that extended JupyterLab for Tessa’s work. Iteration 3 went into production.
What Tessa ended up with
By the end of the engagement, Tessa’s research engineers worked in one notebook environment on Google Cloud, backed by the pipelines and storage from the first two iterations, deep learning libraries for image segmentation, and a version-controlled code base for their analysis models.
If your research teams need a shared cloud analysis environment, contact us.