For over two years, the largest open code dataset was The Stack v2… until today.
For over two years, the largest open code dataset was The Stack v2… until today. The Stack v3 is out: the largest open code dataset ever released: 114 TB, 770 languages, 224M repositories, ~5T tokens of deduplicated and filtered source code. Fully open, no res