This could be the largest synthetic code dataset yet IBM Research 13 July 2026 at 12:00 Introducing CodeAlchemy, a synthetic data pipeline that has already produced nearly 1 trillion tokens of open-source code