Durable pipeline orchestration · open core
Data doesn't just move here. It drives everything.
Cohestra's first product, DataFlow, is a visual, durable data pipeline platform powered by Go and Temporal. Build on a real canvas, ship with AI-assisted authoring, and trust every run to finish — even the ones that crash halfway.
Capabilities
Built for the whole journey.
Every run is durable by default and every stop is visible — build it once on the canvas, trust it to finish.
For building the pipeline
- ✓Visual canvas instead of a wall of YAML — drag, connect, ship.
- ✓AI-assisted authoring (beta) drafts the pipeline — you approve every step.
- ✓Open core — self-host or run managed.
For running it
- ✓Durable execution on Go + Temporal — a crash mid-run resumes, it doesn't restart.
- ✓Every step is inspectable — see exactly where a run is and what it did.
- ✓Connectors for databases, CDC, object stores, streams, and lakehouse tables.
Compose pipelines, not YAML files.
Drag sources, transforms, and destinations onto a real canvas. Every node is inspectable, every edge is a versioned connection.
- Stripe, Postgres CDC, and warehouse sinks as first-class nodes
- Draft → loaded versioning, visible right in the toolbar
- Minimap and zoom for pipelines with dozens of steps
Trace lineage across every active run.
Follow the runtime graph from source to sink — bronze, silver, and gold — with SLA and SLO tracked at every hop, not just at the end.
- Medallion layers tagged right on the lineage graph
- SLA (freshness) and SLO (success rate) tracked per pipeline
- Every row carries its execution, run, and trace ID
Know it's broken before a person tells you.
Runs, success rate, failures, average duration, SLO breaches, and quality rejects — tracked per workspace, not buried in a log tail.
- Execution activity feed with the exact node and error
- Data quality issues flagged next to the run that caused them
- Same durable execution data, just surfaced instead of hidden
Nothing reaches production without proof.
Pipelines move draft → integration → production, and each promotion is gated on evidence, not a person's word.
- A gate only opens once a run has actually passed at that stage
- Promotion history is versioned — you can see what shipped and when
- Rollback restores a prior version, not a guess
The architecture behind the bet.
Pre-launch, building in public. Here's what the core is made of.
Not every pipeline tool solves the same problem.
Most teams start on Airflow and outgrow it — DAGs that restart from scratch on a crash, YAML sprawl, no visual layer. DataFlow's bet: a real canvas instead of DAG files, durable execution instead of restart-on-crash, self-hostable instead of vendor lock-in.
We're pre-launch. Rather than publish a feature-matrix bake-off we haven't run, we'd rather you read the code and judge for yourself.
Pricing, once we've earned it.
We're pre-launch and haven't priced managed plans yet — want real usage data first, not a guess. Self-hosting is free today, no limits.
Managed cloud
- Durable execution, AI-assisted authoring, full canvas
- Talk to us if you want in early
AGPL-3.0 · run it on your own infrastructure, no run limits.
Infrastructure you can inspect, extend, and own.
DataFlow is developed in public with transparent architecture, a public roadmap, and a real contribution path.
Visit the repo ↗