64 lines
2.6 KiB
Markdown
64 lines
2.6 KiB
Markdown
# micro-gopt
|
||
|
||
A go hand-reimplementation of <https://karpathy.github.io/2026/02/12/microgpt/>.
|
||
|
||
Original python is included in the repo for reference against bitrot.
|
||
|
||
To use: `go run cmd/main.go input.txt`
|
||
|
||
Differences between the Go and the Python, as well as notes more generally:
|
||
|
||
- The GPT is implemented as a package and, separately, as a command-line wrapper that calls it, just to keep the algorithm separate from the invocation details.
|
||
- The Value class is more type-safe in go, using values everywhere as opposed to mingling floats and values in the localgrad tuple.
|
||
- The Value struct has actual tests confirming the backward propagation logic.
|
||
- When writing the Value struct and its methods, I accidentally swapped the order of the values in the `localGrads` slice in `Mul` and tore my hair out trying to figure out where the bug was. When I broke down and asked copilot to "compare these two implementations and tell me how they differ," it managed to find the error -- but also reported three non-existent differences and told me that `slices.Backward()` doesn't exist.
|
||
- Initial pass translating the linear algebra functions has me worried that all those value structs aren't going to be very fast...
|
||
- Had to implement weighted random choice. <https://cybernetist.com/2019/01/24/random-weighted-draws-in-go/> made that relatively straightforward; it's a neat algorithm.
|
||
|
||
First proper run:
|
||
|
||

|
||
|
||
Something's not right here, unless the hit new baby name is `kaaaaasehaaeaaal`.
|
||
|
||
...
|
||
|
||
After a few more rounds of debugging, I'm stumped. There must be some subtle pythonic behavior that my rewrite isn't capturing that's causing my results to all be nonsense like `eadaaaaannnaanba` and `oetlaaceta`, but I can't see it (and I don't know enough python to find it).
|
||
|
||
This was still a useful learning opportunity, although a frustrating one in the end.
|
||
|
||
## Update September 2026
|
||
|
||
I asked Claude the same question I asked Copilot six months ago and after thinking for a bit, it pointed out two places where the go program's reference semantics were different than python's. With those two things fixed:
|
||
|
||
```plaintext
|
||
❯ go run cmd/main.go input.txt
|
||
num docs: 32033
|
||
vocab size: 27
|
||
num params: 4192
|
||
step 10000 / 10000 | loss 2.6872
|
||
--- inference (new, hallucinated names) ---
|
||
sample 1: breya
|
||
sample 2: kariste
|
||
sample 3: kari
|
||
sample 4: elyna
|
||
sample 5: aliann
|
||
sample 6: alayn
|
||
sample 7: asari
|
||
sample 8: amara
|
||
sample 9: kadili
|
||
sample 10: avan
|
||
sample 11: aarie
|
||
sample 12: amari
|
||
sample 13: keli
|
||
sample 14: kericy
|
||
sample 15: areta
|
||
sample 16: kailyn
|
||
sample 17: kona
|
||
sample 18: daley
|
||
sample 19: avile
|
||
sample 20: alion
|
||
```
|
||
|
||
Success.
|