2026-08-31
In 1985, a Cambridge startup called Thinking Machines Corporation shipped a black cube five feet on a side, studded with blinking red LEDs, that contained 65,536 processors. Each was a 1-bit ALU with 4 KB of local RAM, wired into a 12-dimensional hypercube network. It was the CM-1 Connection Machine, and it was Danny Hillis's PhD thesis at MIT turned into a commercial product.
Hillis had noticed something obvious that nobody was building for: the brain has ~10¹¹ neurons operating in parallel, not one Cray processor doing 250 million ops/second serially. His 1985 MIT dissertation and book The Connection Machine argued that data-parallel computing — same operation, thousands of data elements, simultaneously — was the future of everything from vision to physics simulation to what we now call machine learning.
The technical roster was surreal. Richard Feynman spent summers at TMC from 1983 until his death in 1988, personally analyzing the router-chip buffer requirements for the CM-1's hypercube network and deriving the correct queue depths from first principles. Marvin Minsky was a co-founder. Case designer Tamiko Thiel gave the machine its iconic look — the red LEDs weren't decoration; they showed per-processor activity, and Feynman insisted they stay because they made parallelism visible.
The CM-2 (1987) added Weitek floating-point coprocessors and hit 28 GFLOPS. The CM-5 (1991) pivoted to MIMD with SPARC nodes and vector units — a 1,024-node configuration cracked 65 GFLOPS on Linpack, briefly the fastest computer on Earth. The CM-5's fat-tree tower with red LED panels was so photogenic it played the InGen supercomputer room in Jurassic Park (1993). Los Alamos, NSA, and NCSA bought them for weather, cryptanalysis, and fluid dynamics.
Then it collapsed. On August 15, 1994, Thinking Machines filed Chapter 11. The kill chain:
Sun bought the software group in 1996 for ~$16 million. The hardware line simply ended.
Why the Connection Machine deserves a second look in 2026: every argument Hillis made in 1985 turned out to be correct. An NVIDIA H100 is 16,896 CUDA cores doing SIMT data-parallel execution — architecturally, it's a CM-2 on a single die. A Cerebras WSE-3 has 900,000 cores on one wafer — that's a CM-1 you can hold in your hands. CUDA kernels are C* with different syntax. JAX's pmap is CM Fortran's FORALL. Transformer training is exactly the "connectionist" workload Hillis built the machine to accelerate.
What we lost wasn't the hardware — silicon caught up. We lost the programming culture. TMC's languages made data parallelism first-class; today's CUDA is grafted onto C++ and leaks abstractions everywhere. A modern Connection Machine — wafer-scale silicon plus a *Lisp-style host language plus visible per-core activity indicators — would be the honest AI accelerator we keep pretending we have.
