Reader submission from @rain落在鱼儿下. These are the contributor’s own views.

In this wave of model startups, almost every team trying to stretch a coin has heard of an “API relay.” Telegram groups post “10% of official price,” “pay 1000, get 1000,” and “high concurrency, zero risk control” all day.

There is no pie falling from the sky. In the gray chain under the model ecosystem, teams think they found a shortcut and step into a prepared pit.

1. A bright afternoon, and a quiet swap

In a maker space in Nanshan, Shenzhen, the afternoon was bright. Zhang Chi (a pseudonym) held an iced Americano and planned the next step. Winning the South China maker contest would mean 500,000 in cash support and three years of tax exemption in the park. Enough, he thought, for a five-person team to survive the prototype stage.

To make the demo faster and cheaper, he connected a relay that the circle was talking about. The product was an AI data-exploration workbench for enterprise analysts, built around multimodal understanding and adaptive reasoning over hard business logic. It was meant to shine at the contest.

“It felt great, until just before the contest we tried to generate cross-department dashboards and attribution reports automatically.” The system threw a fatal format-parse error with no warning. For three full days, two core backend engineers went through prompts, memory leaks, and data structures. The labor alone cost more than ten thousand yuan.

On the third night Zhang Chi skipped the relay console, attached a packet capture, and compared HTTP response headers byte by byte.

The office went quiet except for the fans. Token speed and stream jitter pointed at a different, cheaper open model. The relay’s router was simple and rude: once the prompt grew a little long, it redirected, without saying so, onto the cheap model. The compute they had saved out of meal money had paid tuition to an account reseller’s downgrade trap.

2. Accounting tricks on the bill

If the model swap is a hidden arrow, the billing multiplier is a visible abacus.

The senior architect everyone calls Zhou Ge shakes his head at the pattern. The resellers are no longer people who only sell accounts. They brought over the arithmetic of financial derivatives.

In a second-tier city not far from Shenzhen, the independent developer Lin Feng (a pseudonym) hit that wall. He runs an end-to-end AI marketing workbench. Monthly API spend was several hundred dollars, and it hurt. A relay advertised “pay 1000, get 1000.” He transferred 2,000 yuan on the spot.

Five days later the balance was gone. He recomputed 120 hours of call logs:

  • A hidden multiplier. The charge rate for higher models such as Claude 5 Fable had been quietly set to 2.8×.
  • Inflated tokens. The tokenizer was changed so a Chinese passage of 100 tokens was billed as 150 fragments.

The sticker said 50% off. A single call cost nearly 30% more than the official API. Zhou Ge’s line was that you think you found a bargain, and a few keystrokes drain the balance.

3. Broken memory, and 9:01 on a Monday

In production, the damage of that “low price” compounds.

Half a year later, at the Shenzhen startup final, another team demoed a multi-agent development workbench. The first three rounds of requirement splitting looked clean. On the sixth round, a deep question about the architecture, the agent answered, mechanically, that the uploaded document did not seem to contain that data.

The room reacted. The judges frowned. A night of packet captures found two causes:

  1. A forced context cut. To save gateway bandwidth, the relay used a sliding window. Past 16k tokens it chopped the earliest context, and the agent lost its memory.
  2. Kept cache savings. The relay turned on prompt caching and took the official cache discount, cited here as 50%–90%, while still billing the developer the full amount.

At almost the same hour, an enterprise AI sales-copilot automation engine in that second-tier city met a 429 disaster at 9:01 on Monday morning.

The relay had no regular keys. The pool was web-session credentials taken by reverse engineering. That morning OpenAI deployed a new generation of anti-bot controls, and the pool died at once. The sales system was down for two full hours. Dozens of salespeople stared at an error page. The commercial loss could not be priced.

4. When the data reaches a dark-web forum

The last act of the “save money” game is often a bleak joke.

As official controls tightened, several relays lost their number supply and started losing money. A maintainer left “the controls are too tight, goodbye” in a Telegram group, pulled the plug, and deleted the database. Unused balances vanished.

Two weeks later a dark-web forum listed a database archive of 50GB. Many open-source relay panels had been deployed on cheap VPS hosts with no real protection. After the operators left, the hosts were scanned and the databases taken. Internal code, commercial prompts, multi-agent flow configs, and business data that had never been deeply redacted sat in the open.

At a technical salon, Zhou Ge offered one equation:

Real cost of a model API = listed token fees + the cost of retries after a downgrade + time spent firefighting + the risk of a data leak.

There is still no free lunch. A price far below the market floor turns the visible bill into an operations nightmare and a security bomb that the team cannot size in advance. For anyone who wants a long company, staying away from gray-market relays is not about saving a few coins. It is about not handing over the cards yourself.