We design, build and operate the large-scale data pipelines your business runs on.
Principal-level data engineering, hands-on. We take ownership of outcomes, from architecture decisions to the pipeline running in production.
We start from the questions the business needs answered, then pick the layers that get there: lakehouse or warehouse, batch or streaming, buy or build. You end up with a written architecture, the options we rejected and why, and a cost you can defend in a budget meeting.
Pipelines that move billions of rows without waking anyone at 3am. Apache Spark, Apache Flink and Apache Kafka, in Python, Kotlin or Scala, with tests, safe reruns and predictable failure modes, so a bad night means a retry rather than a week of reconciling data.
Real-time on Apache Kafka, Apache Flink or Spark Structured Streaming, picked for the job rather than out of habit. We handle what decides whether it holds up: state and windowing, late and out-of-order events, replay after an outage, and alerting on lag before anyone downstream notices.
The foundation underneath: Kubernetes, Docker, and Terraform. Reproducible environments, infrastructure as code, and CI/CD for data workloads, treated with the same rigor as application software.
Freshness, quality, and trust: testing strategies, monitoring for pipelines and the runs that silently never happen, AI-assisted incident triage under human control, and giving stakeholders visibility into the data they depend on.
Ongoing senior capacity without a full-time hire: architecture guidance, code review, mentoring for your data team, and a steady hand on the parts of the platform nobody else wants to own.
Core stack: Apache Spark · Apache Kafka · Apache Flink · Apache Iceberg · dbt · Python / Kotlin / Scala · Kubernetes · Terraform
We build sharply-focused tools and operations services for data teams, born from problems we've hit in fifteen years of production pipelines.
Reliability tooling for dbt teams. Answers the question every data team gets asked, "is the data fresh?", and catches the failure your pipeline can't see. Currently in private validation with early design partners.
The fractional platform team for companies that run Apache Kafka but can't justify hiring one. AI-assisted triage with human accountability: incidents diagnosed with evidence, changes gated behind approvals, every action logged, and a senior engineer answerable for every intervention. If you run Apache Kafka and feel this pain, we'd like to talk.
Further tools for data reliability and cross-system consistency are in research. If your team feels a pain in this space, we'd like to hear about it.
Consultancy shaped by what actually makes data projects succeed, or quietly fail.
We start from the decisions your data needs to support, not from the tooling. Architecture follows the problem, never the other way around.
We choose tools by how well they fail, not by how new they are. Usually that means Apache Spark, Apache Kafka and Postgres: when something breaks at 2am, a decade of answers already exists.
Documentation, tests, infrastructure as code, and knowledge transfer are part of the deliverable. Success is your team owning the system confidently without us.
A pipeline that runs isn't enough: the business has to be able to rely on it, visibly. Reliability and stakeholder confidence are engineered in, not bolted on.
Codenexum is an independent data engineering consultancy and product studio, founded and led by Tiago Palma, a data engineer with over fifteen years designing and delivering large, distributed data pipelines across industries.
That experience spans the full lifecycle: greenfield platform builds, rescues of pipelines that grew faster than their architecture, streaming systems processing events at scale, and the infrastructure (Kubernetes, Terraform, CI/CD) that keeps it all reproducible.
The products we build come from the same place: recurring problems seen across many teams, solved once, properly, as focused tools.
Occasional writing on data platforms, streaming, and the parts of this work that only reveal themselves at 2am.
Whether it's a platform to build, a pipeline to rescue, or senior capacity your data team is missing: the first conversation is free and useful either way.
hello@codenexum.io · We reply within one business day.