Yue Yang’s Test: Mid-Small Cloud Providers Lose 400 Million Monthly on MaaS

CategoryOpportunities

Editor's Note: An AI Serial Entrepreneur's Perspective (Content distilled by AI; viewpoints belong to the original author; reading this summary is sufficient)

You Yang, founder of Lucen Technology, reported monthly losses of 400 million yuan selling DeepSeek API on mid-tier cloud platforms after hands-on testing, due to throughput far below official figures. This reveals that mid-tier MaaS is a losing game, but conversely proves the high value of private deployment. For those chasing profit, the opportunity lies not in selling tokens, but in providing "high-throughput private deployment plus tuning" services. Next step: study inference optimization algorithms, then sell private deployment solutions to mid-tier AI companies.

  • Stay away from mid-tier MaaS services—the track is monopolized by giants with razor-thin margins
  • Enter private deployment, serving SMEs without in-house compute capability
  • Mastery of DeepSeek inference optimization code as a high-premium technical selling point
  • Validation: test private deployment willingness with three compute-lacking AI startups

1. What Kind of Opportunity Is This

Serving small and mid-sized enterprises that lack in-house compute, providing high-throughput model private deployment and inference optimization. The core proposition is using technical tuning to solve the pain point of low vendor-side throughput and severe losses, billed on a project basis or annual retainer.

2. Independent Judgment

Worth pursuing? The verdict is highly worth it, but with very high barriers. The key reason is the massive gap between DeepSeek's official benchmark data (14.8k tokens per second on 8 H800 nodes) and mid-tier vendors' measured data (~300 tokens/second), which directly caused mid-tier API resellers to lose 400 million yuan monthly. From the editor's standpoint, the market urgently needs intermediaries who can bridge this technical chasm, not simple cloud-resource resellers.

3. Cold-Start Path

Step one: immediately read DeepSeek's open-source inference optimization code, reproduce the official throughput metrics in a local simulation environment, and produce a quantifiable optimization plan. Cost scale: renting 8 H800 servers for 2–4 weeks of tuning tests, estimated at 100–150 thousand yuan. Timeline: complete technical validation and produce a case report within one month.

4. Biggest Risks and Pitfalls to Avoid

Critical pitfall 1: failed technical reproduction. If you can't reach near-official throughput in a private environment, you can't prove "high throughput" value, and customers will simply buy commodity cloud at market rates. Countermeasure: hire a senior inference optimization engineer focused on overcoming efficiency drops caused by sequence volatility.
Critical pitfall 2: dimension-reduction strikes from giants. OpenAI or major cloud providers could offer the same optimization service directly. Countermeasure: focus on vertical models or specific industries (e.g., video generation, code assistance) with customized deployment, staying clear of the red ocean around general-purpose LLMs.

5. Case Review (How Others Approached It)

  • Technical benchmarking: Analyzed DeepSeek's official data to identify an efficient baseline of 608 billion input tokens and 168 billion output tokens processed within 24 hours, locking onto the operational logic of 226.75 nodes (1,814 H800s) supporting 20–30 million daily active users.
  • Pain-point validation: You Yang's team at Lucen Technology found that after deploying DeepSeek, mid-tier platforms averaged only ~300 tokens per second per node—far below the official theoretical figure—directly calculating the 400 million yuan monthly loss conclusion and confirming that "high throughput" is a matter of life and death.
  • Strategic pivot: On March 1, Lucen Technology announced the termination of MaaS services including DeepSeek API, acknowledging that单纯 selling tokens is a high-risk business of "buying GPU time, selling call volume," and decided to shift toward focused private model deployment for companies of all sizes.
  • Commercial positioning: No longer reselling API access, but becoming the "AI version of Databricks"—offering an integrated private deployment solution from model deployment through inference optimization to operations, charging enterprises that have compute but lack technical tuning capability.
  • Tackling technical difficulties: (Inferred) Addressing online service efficiency degradation caused by input-output sequence volatility and diverse user request patterns, building a dynamic resource scheduling mechanism to maintain high utilization rates.

6. Dual-Track Executability

Cross-border: low feasibility. The overseas market is dominated by AWS/Azure/GCP and OpenAI, and high-performance cards like the H800 are subject to export controls, making hardware acquisition costly and launch timelines long.
Domestic: high feasibility. A large number of Chinese SMEs currently possess compute resources but lack inference optimization capability, DeepSeek's open-source ecosystem is mature, and private deployment readiness testing with three compute-lacking AI startups can begin immediately.

Original article · Late Talk: Read full episode →

Get the Creator Daily by email
Hand-picked opportunities, tools & insights for indie makers — free.
中文读者?订阅中文频道 →
iMessage 邮件 Contact us
中文