In this paper, we describe and demonstrate TEAIS (Test and Evaluation of AI systems), a principled approach for assessing the performance of large language models. We walk through an educational example of TEAIS, where we leverage ideas from property testing of functions on the boolean hypercube to probabilistically verify a relevant security property of a PDF malware classifier. As part of this, we develop a novel property tester for monotonicity that works much like a mutation fuzzer. While the results in this report are preliminary, we hope to spark discussion around the challenges and opportunities associated with the testing and evaluation of complex AI systems.