Sin Chew DailyOctober 2026

Everything Reads Like AI Now

Yuan-Sen Ting / 丁源森View original →

We went out for pan mee the other week. The noodles came up in the soup looking immaculate, every strand identical. I asked the auntie whether anybody still tore them by hand.

My mother answered before she could. Of course not. Labour costs what it costs. They come out of a machine.

I looked at the bowl again. It was, for the record, very good.

Eight in ten

A PhD student and I recently picked up a slightly gossipy question. In astrophysics, our own field, how many papers have been through an AI?

More than we expected. Our estimate is that around a quarter of papers had been AI-edited in 2024, close to half in 2025, and roughly eight in ten this year. The share that says so anywhere on the page is under one percent.

We are not unusual. An analysis released in August swept the full text of more than a million open-access biomedical papers in PubMed Central and found the vocabulary fingerprints of a language model in about 52% of those published in 2024, rising to 89% by the end of 2025.

So even in academia, an industry not short of self-regard, almost everybody is using it and almost nobody is saying so. I would guess the ratio is worse everywhere else.

My student put it well. He has gone back to reading old novels. Open your eyes in the morning and every paper, every LinkedIn post, carries the same faint smell, and he wants something to rinse his eyes with.

What the smell actually is

To explain how we counted, I have to back up.

As earlier columns have covered, a language model is at heart a word-guessing machine. It reads what came before and picks what usually comes next, following patterns absorbed from an enormous pile of text. A person who only ever reads one kind of writing starts to write that way, and so does a machine.

Early ChatGPT wore it on its sleeve. It underscored things. It leveraged things. It delved into things. A study in Science Advances last year counted the words in fifteen million biomedical abstracts and found that after ChatGPT arrived, "delves" was turning up twenty-eight times as often as before.

The method is cruder than it sounds, and that is the point of it. It never tries to guess whether any single paper came from a machine. It watches the vocabulary of a whole population and notices when the distribution lurches. Our astronomy count works the same way.

Where "delve" came from

There is a well-travelled story about why ChatGPT picked up those particular words.

Training a language model involves a stage where human beings rank its answers. The work is tedious and requires no doctorate, so a great deal of it is contracted out to places where labour is inexpensive and English is strong. Parts of Africa are among them. In April 2024 a journalist at the Guardian floated a theory. "Delve" and words like it are considerably more common in formal Nigerian English than in British or American English, roughly the way we reach for lah and loh. The people doing the ranking preferred that register, and the model learned it from them.

This is a theory, and it has never been established.

What is documented is the row that broke out a few days before it appeared. An American venture capitalist posted that he had received a cold email containing the word "delve" and concluded it had come out of ChatGPT. Nigerians on the internet were not amused.

The interesting part is that they did not dispute the premise. They insisted on it. Delve is ordinary English, they said, the English we were taught and have written all our lives. What they refused was the step after that, the inference that a word on a page means a machine put it there.

The objection has aged better than the theory it was attached to. English speakers can now watch the same thing happen to a punctuation mark. Since early 2025 the em dash has been the most notorious tell of all, to the point where novelists and journalists who have used it since long before ChatGPT existed are being accused of running their work through a machine. Some have started stripping em dashes out of their own prose to avoid the suspicion. A mark that Emily Dickinson and Virginia Woolf leaned on now has to keep its head down.

The tells are disappearing anyway

People used to joke that these words were the giveaway, so anyone who cared learned to steer around them. Academia holds AI writing in open contempt, which is exactly why so many use it and so few admit it, and the workarounds travelled fast. Our own study found the same. Ask a model to tidy your sentences today and the conspicuous words mostly stay away.

But style was never only a word list. How words sit beside each other, how sentence lengths alternate, these are statistical quantities too, and statistics happens to be my day job. Everybody has a verbal habit. So does a machine. Its habits simply live a few decimal places down.

Here is the limitation that matters, though. This kind of method measures populations, not people. I can say with real confidence that roughly eight in ten of ten thousand papers have been through a model. I cannot point at any one of them and say it was you. A census will give you the average height of a city. It will not tell you how tall your neighbour is.

And the moment anyone insists on using it against individuals, the bill arrives. In 2023 a Stanford team ran seven widely used AI detectors over TOEFL essays written by non-native English speakers. More than half were flagged as machine-generated. The same detectors read essays by American schoolchildren and waved almost all of them through. The reason is not mysterious. Non-native writers reach for safer vocabulary and steadier sentence structures, and so does a machine. For those of us who learned English second, that is a hard bill to be handed.

So they built a watermark

If you cannot catch it afterwards, mark it at the source.

Each time a model produces a word it usually has several candidates that would serve equally well. A watermark uses a key, held only by the company generating the text, to settle those inconsequential choices. Nothing about the writing looks unusual, but across a whole document the pattern of word choices carries a hidden thread. Think of the security strip in a banknote, invisible until you hold it up to the light, except that this time somebody else owns the light.

None of this is hypothetical. Google DeepMind's SynthID-Text appeared in Nature in 2024 and already runs inside Gemini. In August this year, Anthropic said future versions of Claude will carry a watermark built on the same method, along with an interface for checking.

The push is coming from Brussels. The EU AI Act, whose transparency rules took effect the same month, requires AI-generated content to be marked in a machine-readable form. The rule governs Europe, but the model is the same model everywhere, so the rest of us get it as well.

Catch them, and then what

The hawks are delighted. About time, they say. Flush out everyone who has been quietly using this stuff and let hand-torn noodles have their day.

The defence is thinner than it looks. Papers on how to strip a watermark out were appearing well before watermarking went live. Reword it, run it through a translation, swap a few synonyms, and the thread fades. Others are working on recovering the keys outright. And it binds only the companies that agree to take part. Open-weight models, the ones anybody can download free and run on their own machine, sit entirely outside the scheme.

More damaging is a limitation Anthropic sets out in its own announcement. A watermark can tell you Claude was involved somewhere. It cannot separate "Claude wrote this" from "Claude heavily edited this." It fails from the other direction too. If you only ask a model to smooth your sentences, nearly all the words remain yours and there is very little for the watermark to hold onto. It can establish that a machine was in the room. It cannot establish how much of the work was the machine's.

Step back further and a harder question sits underneath. If that house style was assembled out of tens of millions of documents and the linguistic instincts of a lot of cheaply paid annotators, who exactly is entitled to say the style belongs to them? Which brings us back to delve. The words being convicted as machine-writing were, and still are, words that living people use. Whether you can catch it is one question. What you are entitled to do once you have caught it is another, and nobody has a good answer to the second.

Checking the mirror

For my own part, I have never objected to running a piece of writing through an AI.

The words you are reading were drafted a sentence at a time by me, and then I let a model read them over. It is what anyone does at the door on the way out, checking the mirror for hair sticking up or a collar folded under. Checking the mirror is not an attempt to become somebody else. It is a small courtesy to whoever you are about to meet.

What I do object to is the purism, the position that only hand-torn counts. Hand-torn noodles really are different, uneven in thickness, better under the tooth, and I would not pretend otherwise. But a shop that pours all its effort into tearing noodles and lets the broth, the ikan bilis and the sambal look after themselves is still serving a bad bowl.

That is roughly how I run my research now. I am insufferable about the parts that matter and clear-eyed about the parts that do not deserve my hours, because the hours saved are supposed to go somewhere harder.

None of which means anything goes. Earlier columns have made the case. A student who leans on these tools too early and too heavily loses their own taste first, and after a while cannot tell good work from bad. And a machine can turn out unobjectionable prose without limit, filling your screen and taking the time you might have spent on something worth reading.

The noodles or the broth

In this era the useful skill is not deciding whether a passage came out of a machine. It is the harder one. Deciding whether there is anything in it.

The first of those, machines will shortly make impossible. The second, no machine will ever do for you.

I finished the bowl, incidentally, and drank the soup. The noodles came out of a machine. The broth had been on the stove since before dawn.

You can live with a machine-made noodle. You cannot fake the broth.

About the author: Graduate of Chong Hwa Independent High School, Kuala Lumpur. Earned his PhD in Astrophysics from Harvard University in 2017, then held a NASA Hubble Fellowship at the Institute for Advanced Study (IAS) in Princeton. Former joint professor of Astronomy and Computer Science at the Australian National University, and currently Associate Professor of Astronomy at The Ohio State University, where his work focuses on using AI to accelerate discovery in astronomy and the fundamental sciences. Now on sabbatical in Malaysia, watching the local scene as an outsider, and welcomes letters from readers. Personal website: https://www.ysting.space