Empirical Evaluation of Concept Probing for Game-Playing Agents

Útdráttur

Concept probing is one prominent methodology for interpreting and analyzing (deep) neural network models. It has, for example, formed the backbone of several recent works to understand better the high-level knowledge learned and employed by game-playing agents, particularly in chess. However, some recent theoretical and empirical studies have questioned the methodology's reliability and highlighted some limitations. Here, in the game-playing domain of chess, we investigate the effectiveness of several different probing architectures and look into the reliability of methods for interpreting their results. We use a world-class chess-playing agent as our test domain, which allows us, via self-play, to quantify the importance of the concepts identified in the agent's neural network by the concept probes. Our results demonstrate that the widespread practice of using linear probes and interpreting their accuracy to indicate concept importance is somewhat unreliable and needs to be revised. We demonstrate several ways of doing that in our domain, particularly by using more complex probes and amnesic-like probing.

Lýsing

Publisher Copyright: © 2024 The Authors.

Efnisorð

Artificial Intelligence

Citation

Pálsson, A & Björnsson, Y 2024, Empirical Evaluation of Concept Probing for Game-Playing Agents. in U Endriss, F S Melo, K Bach, A Bugarin-Diz, J M Alonso-Moral, S Barro & F Heintz (eds), ECAI 2024 - 27th European Conference on Artificial Intelligence, Including 13th Conference on Prestigious Applications of Intelligent Systems, PAIS 2024, Proceedings. Frontiers in Artificial Intelligence and Applications, vol. 392, IOS Press BV, pp. 874-881, 27th European Conference on Artificial Intelligence, ECAI 2024, Santiago de Compostela, Spain, 19/10/24. https://doi.org/10.3233/FAIA240574
conference