Before we score anyone, here is the ruler.
Personal AI assistants now read our mail, hold our calendars, spend our money and remember our lives. Benchmarks measure what they can do. Almost nothing measures whether they should be allowed to.
This is the method we will use, published before a single assistant has been scored, so that anyone can argue with the measuring stick before it is used on them. It is version 0.1. It will change, and every change will be dated and kept.
We have not published a score for any product. Everything below is the test, not the result.