27 February 2025
Evaluating retrieval-augmented generation (RAG) is easier than ever, but
you need to keep a close eye on the LLMs that drive the evaluation metrics.
We discuss some common pitfalls and solutions.
Joe Neeman, Nour El Mawass, Maria Knorps, Solomon Ajani
24 September 2024
An empirical analysis of Python packages on PyPI and biomedical journals in 2023, with a focus on the quality of dependency declarations.
Zhihan Zhang, Dorran Howell, Maria Knorps
12 October 2023
Use monad-bayes and rhine in your interactive machine learning application
Manuel Bärenz
21 September 2023
FawltyDeps 0.13.0 introduces a brand new mapping strategy. In this post, we'll delve into the mechanics of how dependencies and imports are matched, as well as how you can leverage these new features to boost your Python dependency management workflow.
Johan Herland, Nour El Mawass, Maria Knorps, Zhihan Zhang
14 March 2023
FawltyDeps is a new tool to help you identify undeclared and unused dependencies in your Python code, making your projects leaner and more reproducible.
Johan Herland, Nour El Mawass, Maria Knorps, Vince Reuter
2 March 2023
Tweag releases the full source code of the Chainsail web service, for sampling multimodal distributions, first announced in August 2022. This blog post gives a tour of the Chainsail service architecture, links out to the relevant parts of the source code and proposes possible extensions to Chainsail for which the Tweag team would welcome contributions from the community.
Simeon Carstens
8 December 2022
A practical introduction to using Sparkle to write Haskell programs which interface with Delta Lake.
Zhihan Zhang
25 October 2022
Tweag intern Abdellatif summarizes his internship, in which he augmented Chainsail with a Bayesian Replica Exchange scheme to improve sampling of multimodal distributions.
Abdellatif Kadiri
18 October 2022
My experience improving monad-bayes, the probabilistic programming language package, as a Tweag fellow.
Reuben Cohn-Gordon
11 August 2022
Etienne Jean, Simeon Carstens
9 August 2022
Tweag announces Chainsail, a simple-to-use web service for better sampling of multimodal distributions with a scalable and auto-tuning Replica Exchange algorithm at its core.
Simeon Carstens, Dorran Howell, Etienne Jean, Saeed Hadikhanloo, Guillaume Desforges
26 May 2022
How to get reproducible development environments for probabilistic programming packages such as PyMC3, Theano or TensorFlow using Nix.
Etienne Jean, Mohamed Nidabdella, Simeon Carstens
17 November 2021
On design choices to build a resource-safe interface for Sparkle using linear types
Noah Goodman
30 September 2021
A discussion and benchmark of an alternative integrator for Hamiltonian Monte Carlo.
Arne Tillmann, Simeon Carstens
23 September 2021
Introducing a library for writing data pipelines which compose well and fail early
Dorran Howell, Guillaume Desforges, and Vince Reuter
28 October 2020
In the final post of Tweag's four-part series, we discuss Replica Exchange, a powerful MCMC algorithm designed to improve sampling from multimodal distributions. An illustrative example and, as always, an interactive Python notebook with easy-to-modify code lead to an intuitive understanding and invite experimentation.
Simeon Carstens
23 September 2020
Meet Lagoon, a new open source tool for centralizing and querying semi-structured datasets.
Dorran Howell
6 August 2020
Learn about Hamiltonian Monte Carlo, and how to implement it from scratch.
Simeon Carstens
26 February 2020
In this blog post series, we're going to lead you through Bayesian modeling in Haskell with the monad-bayes library. In the third part of the series, we setup a simple Bayesian neural network.
Siddharth Bhat, Simeon Carstens, Matthias Meschede
9 January 2020
In this second post of Tweag's four-part series, we discuss Gibbs sampling, an important MCMC-related algorithm which can be advantageous when sampling from multivariate distributions. Two different examples and, again, an interactive Python notebook illustrate use cases and the issue of heavily correlated samples.
Simeon Carstens
8 November 2019
Here's Part 2 in Tweag's Series about Bayesian modeling in Haskell with the monad-bayes library.
Siddharth Bhat, Matthias Meschede
30 October 2019
We're happy to announce the first release of Porcupine, an open source framework to express portable and customizable data pipelines.
Yves Parès
25 October 2019
In this first post of Tweag's four-part series on Markov chain Monte Carlo sampling algorithms, you will learn about why and when to use them and the theoretical underpinnings of this powerful class of sampling methods. We discuss the famous Metropolis-Hastings algorithm and give an intuition on the choice of its free parameters. Interactive Python notebooks invite you to play around with MCMC yourself and thus deepen your understanding of the Metropolis-Hastings algorithm.
Simeon Carstens
20 September 2019
In this blog post series, we're going to lead you through Bayesian modeling in Haskell with the monad-bayes library. In the first part of the series, we introduce two fundamental concepts of `monad-bayes`: `sampling` and `scoring`.
Siddharth Bhat, Simeon Carstens, Matthias Meschede
1 August 2019
We visualize large collections of Haskell and Python source codes as 2D maps using methods from Natural Language Processing (NLP) and dimensionality reduction and find a surprisingly rich structure for both languages. Clustering on the 2D maps allows us to identify common patterns in source code which give rise to these structures. Finally, we discuss this first analysis in the context of advanced machine learning-based tools performing automatic code refactoring and code completion.
Simeon Carstens, Matthias Meschede
17 July 2019
Every day we write repetitive code. A lot of it is boilerplate that you write only to satisfy your compiler/interpreter. But how do languages differ in their boilerplate content? We explore these questions using data sets of Python and Haskell code.
Simeon Carstens, Matthias Meschede
10 April 2019
Inspired by the Event Horizon Telescope images, we develop a quick exploratory study about future possibilities of this technology called the Sneakernet: Could massive data transfer give a new live to the homing pigeon industry? How about using transportation means that are optimized to carry incredible amounts of weight? Or transportation means that are designed to be fast as a bullet?
Matthias Meschede
28 February 2019
Millions of Jupyter notebooks are spread over the internet - machine learning, astrophysics, biology, economy, you name it. What a great age for reproducible science! Or that's what you think until you try to actually run these notebooks. Then you realize that having understandable high-level code alone is not enough to reproduce something on a computer. JupyterWith is a solution to this problem.
Juan Simões, Matthias Meschede
6 February 2019
The repositories of distributions such as Debian and Nixpkgs are among the largest collections of open source (and some unfree) software. They are complex systems that connect and organize many interdependent packages. In this blog post I'll try to shed some light on them from the perspective of Nixpkgs, mostly with visualizations of its complete dependency graph.
Matthias Meschede
23 January 2019
Matthias Meschede, Juan Simões
20 June 2016
Alp Mestanogullari, Mathieu Boespflug
25 February 2016
Alp Mestanogullari, Mathieu Boespflug