When I sit down behind a modern cinema camera, the first thing I do is start turning things off.
The camera is brilliant. It will meter the scene, track the face, smooth the motion, flatter the light, and decide a dozen things faster than I can name them. And I switch off as much of it as the thing will let me, until I’m down to one instrument: the histogram. I keep the histogram because it tells me what’s actually there and then gets out of my way. It shows me the light. It doesn’t take the picture. Everything else on that camera wants, a little, to take the picture for me.
Many working cinematographers I know do the same. Keep the histogram and shed the rest. And the cameras fight us for it — they make the plain setup weirdly hard to reach, and when they do hand back “manual control” it arrives as a thicket of menus you can’t quite trust.
I came up as an “analog” filmmaker, and my craft runs from the outside in. Where am I standing? How bright is the light, and how great is the difference between the brightest thing in the frame and the darkest? What will the film — or the sensor — forgive? What’s in the frame, and which lens serves the picture I already have in my head? The decision starts in the world, and in me, and the camera is the last link in the chain. Digital cameras run the other way, from the inside out. The decision starts in the box.
I’ve spent years trying to name what bothers me about that. It comes down to this: the tool’s intelligence was never the thing to measure. The question is where the decision lives.
About fifteen years ago, I helped build some “smarts” for those newfangled smart cameras.
By 2010 cameras had already gotten intelligent enough that auto-exposure and autofocus had quietly retired the two great sins of the snapshot era, the under/over-exposed frame and the soft shot. The machine had won the old arguments about light and sharpness. So a few of us in a Palo Alto lab asked a different one: we’ve got face recognition now — can we put even more expertise in the shooter’s hands without the machine taking the wheel? We called the thing NudgeCam.1
It watched what you were shooting and pushed back, gently, while you still had time to do something about it. Face too small, it drew a box and told you to step in. Face just off dead-center, it nudged you off to one side. Camera tilted, hands shaking, room too dark — it said so, quietly, and left the decision to you. The app suggested. You decided. We never let it grab the camera out of your hands.
In many ways, versions of it are in wide use today.
The framework we borrowed was Thaler and Sunstein’s libertarian paternalism: guide the hand, never grab it. But it was never only a framework. It was a motive. Even then the friend-or-foe question about smart machines was getting its legs, and NudgeCam was an early, stubborn answer — a genuinely smart tool earns its place by making you more the author, not less. Keep the decision with the human. Build the coach, not the boss.
And I mean coach in a particular way. Not the one who takes the ball and says “just do this,” but the one who helps you get your own body working so the ball goes over the plate.
I didn’t know it at the time, but that was the first draft of everything I think now. NudgeCam was the AI Elder fifteen years before I had a name for him.
My job on that team, more or less, was to be the resident photography teacher — to work out which rules would help the most and could actually be built. Under the hood we used the ones any teacher would sign off on: the rule of thirds, a level horizon, steady hands, decent light. In those early days of face-detection and just barely fast-enough processing, that was hard enough.
But a teacher called in to coach a “newbie” carries a longer, deeper list than the one written on the board. And through every image project we did at that lab, one item on my list kept ringing, one that seemed to reach out and shake hands with the computer: the Golden Ratio. In other words — a rule of beauty that happens to be a formula you can compute.
Put the subject not at a third of the way across the frame but a hair tighter, at 0.618: the proportion the old painters chased, the one you get when you cut a line so the small part is to the large part as the large part is to the whole. Luca Pacioli wrote a book on it in 1509 and had Leonardo draw the figures. Call it the divine proportion or call it a good place to put a face; either way it has been in the water of Western picture-making for five hundred years.2
I think it works. Not as number mysticism — I’ve seen enough of that, the Parthenons and seashells measured after the fact to fit the legend — but as a working painter’s habit that earned its keep over centuries of people trying to make a surface hold your eye.
Belief is a cheap thing to hold and an expensive thing to test, and in 2010 I couldn’t test it. The golden ratio nudge sat just past the edge of what the phone could do. Pulling the raw frames and finding the face fast enough to feel “live” ate the whole budget. Phi was an unaffordable luxury. So I filed it mentally under “someday.”
Six months ago I tried to raise the dead and rebuild NudgeCam on a modern Android tablet.
I failed for a dozen reasons. The toolchain had moved. The camera interfaces I knew were three generations gone. The tablet treated everything as an intrusion. So much the same, so much different. I spent a weekend as an apprentice again to a craft I thought had gotten simpler, and came away with nothing that ran. (I am, at best, a prosumer level coder. Even AI can’t seem to get me very far past that block.)
And while I was failing at my small thing, the image world was going somewhere else entirely. As one example among very many, this month Meta shipped Muse Image, a model that can take anyone’s public photo — your face, tagged straight off your own Instagram, with no warning to you — and drop it into a picture you never posed for. It took a Hollywood union and a talent agency leaning on the company to get the worst of it pulled, days after it launched. The piece that let the machine reach for other people’s faces is gone, for now.
That got me thinking along two lines.
The first starts with the old experiment and the rules it applied. With NudgeCam, the camera is the canvas — a canvas that comes with a few helpful, half-robotic brushes informed by rules — and the artist is the photographer, the human, standing in the world and choosing. With Muse, the canvas and the artist are the same thing. You don’t stand in the world and choose; you say a few words, and the machine is both the surface and the hand. A photograph was always a picture of something that was really there — you point the glass at the world and the light comes back off a real face. Muse doesn’t capture a face. It borrows one, off your own feed, to build a moment that never happened. We’ve gone from getting the real into the box to conjuring the whole thing from nothing; and once there’s no there there, even the plain question of who owns the picture turns strange. It doesn’t need you. It needs your face.
The second line I’d been chasing all along, without quite knowing it: who holds the decision in the moment. Muse keeps the image-making locked inside itself; NudgeCam’s little nudge does some calculating and hands you back the camera.
When you no longer have to take the picture — when you can speak one and it comes out gorgeous — everything you used to do with your hands collapses into a single act: choosing. The craft falls away, and the choosing is all that’s left, which drops the whole weight of the thing onto choosing well. That’s a heavier obligation. But the craft informs the choosing. A DaVinci-coded NudgeCam keeps both powers in your hands and lets each do what it’s good at — the machine measures, the human sees. It respects the natural power of both contributors instead of asking one to swallow the other.
But some questions remain: Does this nudge actually improve the picture? And, in the spirit of that libertarian paternalism, does it leave the artist in charge — or has a heuristic this sophisticated crossed the line from coaching into deciding?
Is the tool better than the maker?
You could settle the value-of-the-rule question with a plain contest. Since the pros would only muddy it with what they already know, recruit ordinary people, not photographers. Each shoots the same few scenes — a face against a wide landscape, a group of people at a party — while the camera cycles through the guides: thirds, phi, the golden spiral, and no nudge at all. Log where the camera actually lands, so you can tell whether the golden grid moved the hand or just decorated the glass. Hide which photo came from which rule, and let people rate them cold. There are even machines now that will score a photograph’s composition, trained on a quarter-million pictures that real people already rated, willing to take the first pass before you trouble a single human eye.
But a deeper question follows: is the golden ratio an invisible hand that overwhelms the user, so strong that it takes charge? Or does it do something else? Like the good coach, is a better-informed nudge bringing about a behavioral change that helps improve the user’s own skills? One man in our old study said the prompts made him feel he had to account for everything in the frame, so he slowed down and planned it. Framed that way, the nudge makes you stop and think — it makes you compose on purpose instead of grabbing the shot and moving on.
Which is the same thing the histogram does: it shows you where the light is and leaves the judgment to you. Done right, the phi nudge is a histogram for composition — it shows you where the golden line falls and lets you decide whether to honor it. It informs; it doesn’t decide. A nudge underpinned by the golden ratio becomes another of the coach’s tools, not the boss. As Joe Gideon says in Bob Fosse’s All That Jazz: “I can’t make you a great dancer. I don’t even know if I can make you a good dancer. But if you keep trying and don’t quit, I know I can make you a better dancer.”
As AI gets more powerful, as the margin between human intelligence and the machine’s closes — with the two vying for precedence — maybe ownership, usefulness, and value are all converging on a single question: who decides? What was in the background with NudgeCam is now at the center of a grand experiment, testing the character of digital empowerment.
The ratio I couldn’t reach in 2010 is closer now than it has ever been — closer than I’d have guessed even six months ago. The AI tools have come far enough, this time, that standing up a working demo in a browser took a few minutes, not a weekend. I fed it a few rough pictures and I liked what I saw. But it did what I suspected it would: it pried the question wider open instead of answering it. Which is better — to let the tool make the picture, or to make the picture with the tool? I didn’t make this demo with the tool, in the end. I let the tool make it — the very question, wearing its own answer.
Taking a long look at where we’ve been against where we are, I can make a fair guess at where we’re going. The cameras, the models, all of it will keep getting smarter — within its own toolkit. The only question that will still matter is the one NudgeCam was already asking, quietly, in 2010 — who’s holding the camera.
Seen this question land somewhere in your own work? Write me. I’d like to hear where it takes you.
1 Scott Carter, John Adcock, John Doherty, and Stacy Branham. 2010. NudgeCam: toward targeted, higher quality media capture. In Proceedings of the 18th ACM international conference on Multimedia (MM ’10). Association for Computing Machinery, New York, NY, USA, 615–618. https://doi.org/10.1145/1873951.1874034
2 I have a wonderful book on my shelf that takes one deep into this subject: The Painter’s Secret Geometry: A Study of Composition in Art by Charles Bouleau, 1980, Hacker Art Books, New York; first published 1963, Harcourt, Brace & World, Inc., New York.