Skip to main content
← Back to the trail log

Jul 2026 – present

FlowLab

In private beta

Role Solo. Design, backend, execution engine, and frontend.

    A browser-based, drag-and-drop workflow builder for tabular data analysis, with AI that proposes edits you can accept or reject.

    Stack: TypeScript · Next.js · React Flow · Python · FastAPI · DuckDB · PostgreSQL

    Started July 2026

    ready-made steps to load, clean, join, analyze, and chart, with no code
    29
    median time to preview a step, on a hosted database
    18.5 ms

    The problem

    Analysts who live in spreadsheets end up rebuilding the same load, clean, join, and summarize steps by hand, and then cannot explain where a number came from. FlowLab lets them compose those steps as a graph of nodes and run it on a server, so every result can be traced back through the graph.

    The graph validates as you build it. A node that refers to a missing column turns red before anything runs.

    The hardest stretch

    Keeping two languages honest. Schema prediction lives in TypeScript so the canvas can flag problems at design time. Execution lives in Python and DuckDB. If the two ever disagree about a node's output schema, the UI lies to the user.

    I wrote each concern once and made the disagreement detectable instead. At run time the worker asserts that the predicted schema equals the actual one for every node, and treats a mismatch as our bug, not the user's. In CI, a golden-fixture harness runs every node type through both sides and fails a pull request that has no fixture.

    The AI repair loop taught me where this breaks down. The model cannot converge on a node whose warning only clears after a sample run, so pivot burned all four repair attempts on one unchanged diagnostic. I removed pivot from the AI-generatable catalog rather than hide the failure.

    Decisions

    Chose Server-side execution with a warm session worker

    Rejected DuckDB-WASM previews in the browser

    A browser preview needs a second config-to-SQL compiler in TypeScript, caps dataset size at the browser's memory, and turns preview-versus-run drift into a user-facing bug. I set a tripwire instead: a warm preview at p50 under 500 ms.

    Chose Predict schemas once in TypeScript, execute once in Python

    Rejected A shared schema-transform DSL

    A declarative DSL handles select and rename, then joins, group-by, and pivot turn it into a bad programming language. The cost of my choice is real: a new node is a TypeScript definition, a Python executor, and golden tests.

    Chose A Postgres-backed job queue

    Rejected Kafka

    Kafka is a log, not a job queue. It has weak per-job retry and acknowledgement and no priorities or delays. It only earns its place if I build continuously running workflows later.

    Chose AI edits as structured patches, shown as a canvas diff to accept or reject

    Rejected Regenerating the whole workflow

    Regeneration destroys the user's layout and quietly changes unrelated nodes, which defeats the point of an auditable tool.