The monthly report that says whether your AI still works
You built it, you bought it, or a platform vendor deployed it. A support and run contract puts each named production system under continuous evaluation and sends you one report a month: whether output quality has held, what it is costing, and what broke. The panel below is that report, in the format you receive it.
What the report tells you before you have to ask
A mock of our own reporting format, filled with sample data, for one AI system over one month. Pass rate is the share of sampled cases that met the agreed definition of correct. Cost per case is inference spend divided by the work done. The dip is a supplier changing model version without saying so, which is the kind of event this contract exists to catch.
Interface mock. Figures are illustrative sample data, not a client system.
Somebody has to own whether it still works
That is what the contract is. No vague retainer: a severity model, response times, and a metric we report against whether it flatters us or not.
Continuous evaluation
Every release and a sampled share of live traffic scored against agreed acceptance thresholds, with the judge itself calibrated against human labels.
Drift, including judge drift
Input distribution drift, silent model version changes by your supplier, and movement in the evaluator’s own behaviour. The third is the one most people never watch.
Cost control
Token spend per unit of work, budget alerts, caching and batching, and a recommendation each quarter on whether a cheaper model would hold quality.
Incident response
Defined severities, a named responder, rollback procedure, and a written post-incident note that goes in your audit trail.
Quarterly improvement
One release a quarter aimed at the metric in your baseline, with the change in that metric reported rather than the work delivered.
Service levels
Scoped by the number of production systems under contract and the severity cover they need.
One production system. Monitoring, evaluation, monthly report, SEV-2 response.
Up to three systems. Adds SEV-1 response, cost optimisation and a quarterly improvement release.
Four or more systems, or a regulated environment needing named oversight. Audit and certification support is scoped separately with Pixelette Certified.
It does not have to be something we built
We will take on systems we did not build, once a baseline tells us what we are inheriting.
Questions worth answering
How is an AI support and run contract priced?
Against the number of production systems under contract and the severity cover they need, not against headcount or hours. Watch covers one production system with monitoring, evaluation, a monthly report and SEV-2 response. Operate covers up to three systems and adds SEV-1 response, cost optimisation and a quarterly improvement release. Estate covers four or more systems, or a regulated environment needing named oversight. Every contract is scoped and quoted after a conversation.
What are the response times?
SEV-1 within 1 hour, SEV-2 within 4 hours, SEV-3 by the next working day. Each incident gets a named responder, a rollback procedure and a written post-incident note for your audit trail.
Will Pixelette support an AI system it did not build?
Yes, once a baseline establishes what is being inherited. Running what somebody else wrote is the clearest proof that this is a capability rather than a warranty on our own work.
What is judge drift?
Movement in the behaviour of the model doing the grading, as distinct from drift in the input data or a silent model version change by a supplier. Almost nobody watches it, and it quietly invalidates your quality measurements when it happens.
Already have something in production?
The baseline works just as well on a system that exists as on one that does not. We measure what it is doing now and tell you what it would cost to keep it honest.