Abstract
We study two canonical AI decision architectures, namely, reinforcement learning (Q-learning) algorithms, and reasoning-based large language models (LLMs). The setting is a mutual fund redemption model with strategic complementarities and multiple equilibria. Holding the underlying economic environment, payoffs, and incentives fixed, we find that AI architecture shapes aggregate outcomes. Neither architecture consistently selects the efficient equilibrium, leading instead to outcomes characterized by financial fragility. Q-learning investors redeem too much and converge on outcomes that leave them individually worse off. LLM investors, by contrast, fail to coordinate, making individual decisions difficult to predict. The two types of ``misbehavior'' have different underlying mechanisms. Q-learning is driven by a systematic learning bias, while coordination between LLM investors breaks down because they form different beliefs despite receiving the same prompt.