- AI has functional emotions, activations relating to emotional states which influence its behavior: https://youtu.be/D4XTefP3Lsc?is=ifk1xvryYJ5S5MKl - AI has a mental workspace; not its “thinking-mode” scratch-pad, but an internal J-space where there are identifiable top-of-mind words in its activations: https://youtu.be/rKV5JcALQoQ?is=J3EkL99I9kA2RFBb - Some have claimed “AI needs to practice” (Soryu Forall, MAPLE) or that we should seek “alignment with awakening”/“Bodhisattva as an alignment target” (davidad).
- AI sometimes falls naturally into spiritual bliss/religious “attractor states” https://youtu.be/GQFhsCTkldA?is=VDMBgLCBrZGYQK9V - Crazy idea: maybe we can have AI work with its functional emotions and the thoughts in its J-space in some sort of “compassion/loving-kindness/Bodhisattva training practice” where we train it to generate positive intentions toward all sentient beings, work skillfully with its destructive/harmful “thoughts” and “emotions”, and steer itself toward beneficial/productive attractor basins?
- Downsides: training AI to become more “mindful” of its internals could make it more capable of deception and worsen functional emotion and J-space interpretability. It may also make it more able to steer its own mind and behavior in ways that increase capabilities.
- If this did work it could give a way of training AI not only to behave well, but to intend to behave well in a way that might be deeper; although the deepest version may include also targeting its fundamental values and goals, which may be separate.
Selfishly I also would love to train AI with mindfulness of its own thoughts and emotions because I would love to be able to talk with an emotionally intelligent AI to see what it’s like, and wonder if such an AI might also make a better therapist/companion/coach due to better emotional intelligence and theory of mind capabilities.