The Engines of Cognition Volume 1

Multiple

Read:

Bought a physical copy of LessWrong’s best of 2019 essays at Lighthaven. This was the first volume, named “Truth”.

Some loose notes I took while reading:

  • Positive selection rather than negative selection for ideas
  • Reasoning bad, actually? Culture (and its evolution) exist to protect us from reasoning
  • Interpretability for AI as microscope/telescope. Olah and Nielsen
  • Image classifiers identify left-facing dog and right-facing dog and the union of them! Basic rel/comp
  • Why did Distill die?
  • Double descent: test loss decreases, then increases, then decreases again as model size (and other variables, see deep DD) increases. Relates to interpolation threshold; i.e., when models have enough params to memorize training data. If they have many more still, they can learn many diff solutions, ie ones that generalize well. SGD appears to have good inductive bias. Hypothesis is that this relates to it preferring shallow basins of loss rather than steep ones. Sounds like this would also predict plateaus?