INDUSTRY INSIGHTS
Smart, Safe, and Boring AI
Neelanjan Sinha
VP – Product & Technology
GenAI can produce impressive output in seconds. The harder challenge is making that output reliable enough for real publishing workflows.
That was the idea behind our session at SSP this year. Drawing on our experience across operations, sales, and product, we wanted to spend less time on what AI can do, and more time on what has to happen around it before people can trust it.
That is what we meant by boring AI.
Not boring as in unimportant. Boring as in deliberate, repeatable, and often painstaking. It is the work that sits around the model: orchestration, guardrails, evals, reviewer workflows, measurement, and human judgement.
In practice, we have found ourselves coming back to two ideas:
- Automation: removing unnecessary effort from the system
- Intelligence: helping improve necessary human judgement, with better signals, context, and support.
The hard part is knowing which is which. Some work should disappear. Some work should stay with people, but become easier, better informed, and more consistent.
What Arty taught us
Sandeep used Arty as one example of a broader publishing problem: how do we take AI output from plausible to dependable enough for real content workflows?
Arty is an AI-powered alt-text generation product that provides image descriptions for more accessible research content. The use case is image accessibility, but the questions are familiar across publishing:
- Does the system understand the content, rather than just the surface?
- Which output actually needs human review?
- How do we make review more consistent without turning every reviewer into a copyeditor of AI text?
In earlier versions, 69% of Arty’s output had to be manually reviewed and changed. It would have been easy to treat that only as a model-quality problem. We came to see it as a system-design signal.
Three patterns stood out.
- Context gaps: some outputs described what was visible but missed what was meaningful. A bar chart could become “red, blue, and green bars” instead of a chart comparing quarterly revenue.
- Review bottlenecks: reviewers looked at everything. That protected quality, but it also made every output look equally risky.
- Calibration gaps: some edits improved accessibility or accuracy; others mostly changed style. “Red and white aircraft” becoming “A red and white aircraft” did not add value for a reader relying on alt text.
None of these felt unique to image accessibility. Similar questions appear wherever AI enters editorial, production, sales, or customer operations.
The model matters. So do the workflow and the human judgement around it.
When our teams encounter problems like these, we try to engage with them across three layers: technology, process, and people.
Technology: teach it what only an expert would catch
A generic model sees “a graph.” A useful one asks what the image is for: is it decoration, or is it carrying data the reader needs? We taught Arty to describe meaning, not just what is visible. “A graph” became “a vertical bar chart comparing quarterly revenue.”
Process: send review where it matters
Reviewing everything equally means reviewing nothing well. So a second AI now checks Arty’s output first and flags only what needs a person. Reviewers spend their time on the hard calls, not the easy ones.
People: agree on what to leave alone
Our best reviewers cared so much that they rewrote text that was already correct: an article here, a preposition there, a matter of taste. So we drew a line. Stylistic edits (wording, tone, preference) are not required. Substantive edits still matter: a wrong noun, a missing fact, a phrase that is not accessible.
Better intelligence is not only about improving the model. It is also about giving people enough context, examples, and feedback to make better decisions.
The result
Today, 73% of Arty’s image descriptions are accurate and useful as they are, and only 27% are flagged for human review.
For us, this is not simply a story about a better model, or even about image accessibility. It is a story about workflow.
Technology improved the output. Process focused attention where it mattered. People established the judgement that made both useful.
That is the more precise meaning of boring AI.
It is the deliberate work of making AI dependable after the demo: the wiring, measurement, review design, training, and iteration that allow people to trust the system.
What our customers saw
We did not have to make the case ourselves. Rob O’Donnell shared RUP’s experience with Arty, and Bill Kasdorf, sitting in the audience, offered a second opinion. We did not plan that, we promise!

“We’ve been impressed by the quality of descriptions from Arty. Our scientific editors reviewed everything before we made a decision, and they were impressed. Authors are reading the descriptions and, when they do edit them, they are not making significant changes.”
Rob O'Donnell | Senior Director of Publishing at Rockefeller University Press

“I’ve reviewed thousands of image descriptions from scores of publishers over the years. When I saw the recent results, I was blown away by how good they were. I’d say this is one of the best I’ve seen.”
Bill Kasdorf | Accessibility Consultant at Kasdorf & Associates
The same lens, across our products
Arty is one example. We find ourselves asking the same questions across every AI product we build:.
- What should automation take away?
- What should intelligence help a person do better?
- Where does a human still need to decide?
None of this is glamorous. That is the point.
The boring work is what lets the smart work be trusted.
We’d love to hear where you land on this, especially if you disagree. Write to us at communications@tnq.co.in.
Watch the full session here.
Want to learn more about Arty? Read the full case study here and a more technical whitepaper here.
This post is based on our Industry Breakout Session at SSP 2026 titled Smart, Safe, and Boring AI: From Prototype to Production’, given by Sandeep Dhawan, Byron Laws, and Neelanjan Sinha, along with Rob O’Donnell, who spoke about using Arty in RUP’s workflow.

