I'm Diego, and I've been developing grounded AI systems over the last 11+ years
-
Are you frustrated by brittle AI prototypes that fail outside of development environments?
-
Is critical business knowledge scattered across large volumes of unstructured documents?
-
Are your technical teams facing unpredictable framework behavior or messy data extraction?
-
Do you need secure, event-driven AI assistants that integrate seamlessly into your operations?
-
Looking for a rigorous, PhD-level engineering standard to maximize worker productivity?

About me
I build end-to-end RAG systems and custom knowledge workflows for engineering and operations leaders who need production-ready automation but are frustrated by brittle prototypes. If your team is struggling to extract clean data from large volumes of scattered, unstructured documents, I help fix that. I build with LangChain and LangGraph when the orchestration justifies it, and with custom Python and FastAPI pipelines when it doesn't — always keeping the architecture explicit, so the LLM handles the unstructured reasoning while your backend retains total control over the execution loop.
My approach is backed by over 11 years of experience as an AI/ML Scientist across Europe and the US, along with a PhD in Bioinformatics. Having spent more than a decade managing massive data pipelines and training machine learning models, I bring a rigorous, scientific standard to modern AI production. I take projects systematically from an initial data audit and custom architecture design to containerized cloud deployment, ensuring your systems handle operations reliably from secure webhook ingestion to native vector retrieval.
🛠️ Tech Stack: Python, FastAPI, Async/Await, OpenAI API, Vector Databases (pgvector), LangChain, LangGraph, Langfuse, Pydantic (Structured Outputs), Celery, PostgreSQL, Azure, Hetzner, Docker, GitHub Actions (CI/CD), Data Pipelines, Evals, and Machine Learning.
Why work with me?
Here's what sets me apart, and how I can help your business:
-
Scientific Rigor
With a PhD in Bioinformatics and over 11 years as an AI/ML Scientist, I treat AI integration as an empirical engineering problem, not a guessing game. I bring over a decade of experience designing robust statistical frameworks and validating machine learning models to your production pipelines.
-
Explicit Architecture, Framework When Warranted
I develop AI agents with LangChain and LangGraph — LangGraph for the orchestration graph, LangChain components inside the nodes for retrieval and tooling — and I know the stack well enough to know when it earns its place. When a workflow doesn't need that machinery, I build custom pipelines directly in Python and FastAPI with strict validation schemas. Either way, the architecture stays explicit: the LLM is confined to the reasoning it's good at, while your backend keeps full control over the execution loop. Framework when it's warranted, framework-free when it reduces technical debt.
-
Built for Production Scale
My background includes handling massive datasets and setting up complex distributed architectures. I design production-ready automations that handle everything from secure webhooks to asynchronous job queues and cloud hosting.
-
Continuous Systematic Evaluation
I don't just deploy an AI wrapper and leave it running blindly. I integrate systematic evaluations to track, score, and benchmark model outputs over time. This keeps your workflows accurate and reliable as your documentation scales.
What they say about my work
-
David J Bishop
Research Leader at iHeS (VU - Australia)
"I had the pleasure of working with Diego on a large-scale genomic study investigating genetic predictors of physical performance in military recruits. Diego led the analytical side of the project, processing and modelling data from over 3 million genetic variants across more than 1200 participants. His work involved rigorous quality control, population structure analysis, and the development of polygenic scores.
Diego’s technical and analytical skills are exceptional. He independently built and validated complex pipelines and implemented robust statistical models. His ability to translate raw data into meaningful insights was critical to the success of the project. He also communicated findings clearly to a multidisciplinary team, bridging the gap between computational biology and applied military research. Beyond his technical expertise, Diego is proactive, thoughtful, and collaborative. He consistently sought feedback, proposed improvements, and showed a strong sense of ownership. His contributions were instrumental in shaping the direction of the study.
I highly recommend Diego for any data science role that demands analytical rigor, scientific curiosity, and a collaborative mindset. He would be a valuable asset to any team working with complex datasets and looking to extract actionable insights.
Beyond his data analysis skills, he is also a great guy and easy to work with."
-
Jonatan R Ruiz
Director at iMUDS (UGR - Spain)
"I have worked with Diego over the last nine years across multiple projects. He has consistently led the processing and analysis of massive datasets, encompassing tens of millions of genetic variants across thousands of participants.
While I could praise Diego for many reasons, what stands out most is the rigor and neatness of his work. It is striking that a person who self-taught advanced programming and machine learning during his PhD reached such high standards in data processing, statistical analysis, and model interpretability. In every project, Diego took a proactive role in technical strategy, from data quality control and processing to selecting the machine learning architectures best suited for our needs.
Regardless of the challenges we encountered, Diego always identified the right path forward to achieve our project goals. I highly recommend him for any lead role in Data Science, Machine Learning, or analytical leadership. He deeply investigates the best approach for the project at hand and delivers results at the highest industry standards.
Happy to continue! show must go on!"
Frequently asked questions
How quickly can you start working on my project?
I can typically begin new projects within 1-2 weeks of contract signing. For urgent matters, I keep some flexibility and can potentially start sooner — just let me know your timeline during our initial consultation.
Do you require a minimum project size or commitment?
While I can accommodate projects of any size, I find that engagements of at least 20 hours allow for meaningful impact. This gives us enough time to understand your data, implement solutions, and deliver results you can act on. We can start with a small pilot project to ensure we're a good fit.
What industries do you have experience in?
My background bridges biomedical and regulated-data domains with production AI engineering, so I am strongest where critical knowledge is buried in dense, unstructured documentation and where correctness actually matters — life sciences, biomedical, and technical operations. That same domain grounding makes me particularly effective on evaluation-heavy work, where knowing what a correct output looks like is half the problem. That said, the underlying methods — retrieval, structured extraction, systematic evaluation — apply across sectors.
How do you handle data security and confidentiality?
I take data security extremely seriously. I sign comprehensive NDAs before starting any project, use enterprise-grade encryption for all data transfers, and follow industry best practices for data handling. I can also work within your existing security infrastructure and policies.
How do you communicate progress and results?
I maintain clear communication through weekly progress updates and regular check-in meetings. You'll receive detailed documentation of all analyses, findings, and recommendations. For ongoing projects, I provide interactive dashboards and reports that allow you to track progress and results in real-time.
-
Let's have a virtual coffee together!
Want to see if we're a match? Let's have a chat and find out. Schedule a free 30-minute strategy session to discuss your AI challenges and explore how we can work together.