Qualcomm CEO's Disclosure and the 2028 Target
In an interview on Alex Heath’s Sources podcast, Qualcomm Chief Executive Officer Cristiano Amon disclosed that leading frontier AI research labs have formally approached the company to request that future mobile platforms support continuous on-device execution of models with at least 100 billion parameters by 2028.
The requested performance threshold represents an aggressive leap beyond current mobile silicon capabilities. Qualcomm’s contemporary flagship platforms, including the Snapdragon 8 Elite Gen 6 and Extreme Gen 6 manufactured on a 2-nanometer process node, are architected to support on-device Mixture-of-Experts (MoE) models up to approximately 30 billion parameters. Scaling to the 100B parameter tier demands a more than threefold increase in execution density on smartphone form factors in under three years.
Amon attributed this client demand to autonomous, always-on agentic workflows. Leading AI developers require local, continuous inference to sustain low-latency responses, preserve personal user context, and safeguard sensitive data without relying on continuous backhaul to central cloud infrastructure.
Architectural Strategies and Physical Hardware Bottlenecks
To satisfy the 100B parameter requirement within mobile constraints, Amon cited several architectural strategies. These include the implementation of tiered, low-power memory subsystems, the dynamic paging of weights from high-speed flash storage, and expanded memory allocation dedicated directly to the neural processing unit (NPU).
Crucially, Qualcomm did not issue a binding product roadmap or delivery guarantee for 2028. Amon framed the disclosure strictly as customer demand signals rather than a committed specification for upcoming Snapdragon generations. Furthermore, the chief executive did not identify specific AI labs by name, referring only generally to companies at the forefront of frontier AI development.
Physical and thermodynamic realities present formidable barriers to continuous edge inference. Mobile form factors operate within strict thermal dissipation budgets and rely on battery packs that cannot sustain heavy compute loads for extended periods without aggressive throttling or thermal saturation.
Practitioner Reaction: VRAM Feasibility and Economic Pressures
Systems engineers and edge hardware practitioners met the disclosure with immediate technical skepticism, breaking down the mathematical constraints of running 100B models. Technical calculations indicate that even when quantized down to 4-bit or 4.5-bit precision, a 100-billion-parameter model requires roughly 50 to 57 gigabytes (GB) simply to store base weights in memory.
Once activations, extended context windows via key-value (KV) caches, and core mobile operating system overhead are factored in, engineers calculate that an operational environment requires between 70 and 90 GB of high-bandwidth memory. That allocation dwarfs the physical capacity found in top-tier consumer smartphones today.
Practitioners noted that equipping smartphones with 70 to 90 GB of high-speed LPDDR RAM would introduce prohibitive manufacturing costs, likely limiting such configurations to hyper-premium tiers. Observers also pointed out that frontier AI labs are strongly incentivized to shift inference workloads onto end-user silicon to escape escalating data center electricity, compute, and token serving expenses.
Strategic Implications for Thailand's Enterprise Sector
For enterprise technology leaders in Thailand, the shift toward heavy on-device parameter execution carries strategic significance, particularly for banking, fintech, telecommunications, and healthcare sectors subject to Thailand's Personal Data Protection Act (PDPA). Executing 100B parameter intelligence entirely within client hardware eliminates network transmission exposure and addresses strict data sovereignty mandates.
Furthermore, Thai businesses facing volatile foreign currency expenditures on dollar-denominated cloud API tokens could dramatically lower recurring operating costs if user-facing agents run locally on customer devices without pinging central data centers.
Nevertheless, enterprise decision-makers should maintain a pragmatic outlook. Because the 2028 timeline reflects frontier lab demands rather than committed silicon deliverables, organizations should not base short-term IT procurement cycles on ubiquitous 100B edge hardware. Strategic planning should prioritize optimizing small language models (SLMs) on current architectures while tracking memory density advancements.
If achieved, enterprises could run state-of-the-art language models directly on mobile edge silicon, slashing recurring cloud inference bills and resolving data sovereignty constraints under local privacy regulations.