观点这是中量级玩家的标准剧本——用开源基座做垂直后训练,把推理成本打到可持续水位。对一人公司的信号:如果你的业务有固定垂直场景(课程问答/转化话术),自己在开源模型上做一轮 SFT+RL 的 ROI 正在快速超过直接调 API。
We’ve post trained a model on top of Qwen that achieves Pareto optimality on accuracy-cost curves.
Unlike our previous post trained models, this model has been trained to be good at search and tool calls simultaneously, allowing us to unify the tool call router and summarization together in one model.
The resulting model performs better than GPT and Sonnet in terms of cost efficiency to serve daily Perplexity queries in production. The production model runs on our own inference platform.
We’re already serving a significant chunk of our daily traffic with this model and intend to have it serve all of default traffic pretty soon.
More research to follow soon on models we’re training and deploying for Comet and Computer.
Kimi.ai 推出 Kimi K3:开源前沿智能。2.8万亿参数,100万上下文,原生多模态。Kimi Delta Attention在百万Token上下文中实现了高达6.3倍的解码加速。Attention Residuals在增加不到2%额外开销的情况下,提升了约25%的训练效率。
对你意味着开源模型能力上限再次提升。
事实Kimi K3 拥有 2.8T 参数并开源了 Delta Attention,
观点这将显著降低长文本与多模态推理成本,CEO 应评估向此类开源模型迁移的可能性。
IncredibleKimi.ai: Introducing Kimi K3: Open Frontier Intelligence🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional
Vera CPU 专为 Agentic 运行时定制构建。我们一直在与 NVIDIA 合作,基于此运行为 Perplexity Comet 提供支撑的沙箱基础设施……
对你意味着Perplexity 与 NVIDIA 联合定制 Agentic Runtime 专用算力,揭示了一个趋势——顶级 AI 应用公司的竞争已延伸至算力基础设施层。
事实NVIDIA Vera CPU 专为 Agentic Runtime 定制,Perplexity 已接入此基础设施运行其 Comet 产品的沙箱环境。
观点当头部应用公司开始与芯片厂商联合定制算力,通用 API 调用已不足以支撑竞争差异化——基础设施层的布局窗口正在收窄,中型 AI 公司需要尽早确定自己的算力合作伙伴战略。
Vera CPUs are custom-built for agentic runtime. We've been working together with NVIDIA on running our sandbox infrastructure that powers Perplexity C...
Perplexity Pro 用户现在可以使用 Compute credits 来体验 Model Council。在法律、医疗和金融研究领域,这是最受用户喜爱的功能之一。
对你意味着Perplexity 正在通过 Model Council 加强垂直领域的专业搜索体验。
事实法律与金融研究依赖多模型交叉验证以提高准确率。
观点纯单模型应用在专业场景的信任度面临瓶颈,CEO 应思考是否在自己的 AI 产品中引入多模型协同与交叉校验机制。
Perplexity Pro users can now use Model Council with Computer credits. It's a user favorite when it comes to legal, medical, and financial research, wh...
观点结合他同日透露的「更多万亿参数开源模型即将到来」,开源生态正在迅速逼近闭源前沿。对你的战略影响:三个月后当前锁定特定闭源 API 的决策可能需要重新评估,提前建立模型切换能力是降低风险的合理举措。
GLM is the kind of model that revives serious interest in open source AI. It passes the blind test relative to the frontier models on the median produ...
Perplexity Computer 是一个持续交付的 Agent 框架。Deep Research 现已成为 Computer 内的原生技能(用户不必显式调用,系统自动判断何时启动)……原文被截断。
对你意味着
观点Deep Research 从显式工具升级为后台自动触发的技能,代表 AI 产品设计正从功能菜单走向意图感知——用户无需选工具,AI 自己判断何时需要深度搜索。
事实这是 Perplexity CEO 本人宣布,非二手报道。若你在做 Agent 类产品,值得思考哪些能力应该原生化而非让用户手动选取,这往往是决定留存率的设计分水岭。
Perplexity Computer is an agent harness that just keeps delivering. Deep Research is now a native skill inside Computer (you don’t have to explicitly...
观点这是 OpenAI 让模型无处不在的生态战略直接体现——第三方平台快速接入新模型,API 生态成熟度在加速。对你的 AI 工具选型有参考价值。
GPT-5.6 Sol is available as an orchestrator model inside Computer. Both Sol and Terra are also available as models for search for Pro and Max users! E...
事实上下文碎片化是当前所有企业 AI Agent 落地的真实痛点——各 SaaS 工具数据相互孤立,导致 Agent 无法跨系统推理。
观点谁先建立统一的企业上下文图谱,谁就掌握了 Agent 时代的数据控制权入口
Context graphs will be the best way for businesses to enable and deploy agentic harnesses. There's a lot of context fragmentation across so many diffe...
future: continual learning of the model + harness hosted on hardware you own and control, with full access to all sensitive context that never leaves ...
目前已成为 Perplexity Computer 上使用量仅次于 Opus 4.8 的第二大编排模型。一旦我们获得更多算力,我们打算提高使用上限……
事实Opus 4.8 仍主导顶级 Agent 编排场景,但第二梯队模型正在快速追赶;
观点当前 Agent 产品扩充用量的主要瓶颈在于算力配额供应,构建 Agent 应用的 CEO 需及早建立多模型弹性编排与算力保障机制。
Already the second most used orchestrator model on Perplexity Computer now, only behind Opus 4.8. Once we secure more compute, we intend to increase u...
Truethe tiny corp: Yes, but nothing compares to the feeling of knowing the metal box your AI lives in is yours. Forget home ownership, owning a datacenter is the new American Dream.
The best application for models to run locally on hardware you own would be personal robots. There’s no way anyone is going to get comfortable stream...
GB 200s change how one does the prefill and decode disaggregation when serving large MoEs like Qwen. We’ve published details of our stack quantifying...
事实本地部署高质量模型的门槛正在降低。对正在评估减少云 API 依赖、控制数据隐私成本的 CEO 来说,这是一个值得纳入选型清单的硬件选项。
The DGX Spark is an incredible piece of hardware. Running it to almost full GPU and RAM utilization and it still doesn’t have any issues with the hea...
Worth reading. The marginal water consumption of a properly implemented data center for its liquid cooling is almost zero. People confuse water needed...
观点外部监管与合规体系正在加速倒逼 AI 企业从单打独斗转向政治与产业维度的纵横联合。CEO 需密切追踪政策走向,提早评估合规风险并布局应对方案。
today is the most positive i have felt about the future of american ai. the coming together of so many companies in a united manner to fight against r...
事实Perplexity CEO 认可开源普及将加速模型商品化与底线价格战,未来价值转向神经/符号系统结合。
观点言外之意,CEO 不应依赖单一 API 模型的微弱差异,而需将研发重力压在结构化逻辑推导与混合 Agent 架构设计上。
Good observationsGrady Booch: In the fullness of time, LLMs will eventually become commodities and their price will be a race to the bottom. As the competitive advantage of one LLM over another shrinks - particularly as open source models expand - the value proposition will turn to the neuro/symbolic
Strong resultsTogether AI: We analyzed Kimi K3 vs. Claude Fable 5 for software engineering tasks using DeepSWE.Kimi K3 gets you the same performance as Fable 5 at ~35% of the price, and it actually pulls ahead at higher pass@k's. More insights in the thread!
观点如果万亿参数级开源模型大量涌现,推理成本将进一步压缩,当前依赖闭源 API 的业务护城河会系统性缩窄。建议提前回答一个战略问题:当模型成为白菜价时,你的产品差异化究竟在哪里?
Also, other multi-trillion-parameter open-source models are landing soon, from what I hear. It's going to be awesome for token pricing and riding the ...
对你意味着应用层 AI 企业应积极布局开源模型与多模型混合架构。言外之意,过度依赖头部闭源供应商可能导致公司在供应链与成本端受制于人。
观点开源生态的繁荣将为应用层创业公司提供极高的架构弹性与战略议价筹码。
“Open weights strengthen competition and competition is what keeps the benefits of AI broadly shared rather than concentrated in the hands of few”. ...
Well said!Bill Gurley: The way to think about “open” in software is being a low-cost producer (vs a high margin one). When a company (or 2) achieves record valuations in record time, that rightfully attracts competition (as it should). We need to let the free market work. https://wapo.st/4pvR0mJ
Perplexity has the best (both on cost and performance) deep and wide research harness in Computer. One of the contributing factors is strong internal ...
观点Aravind 描述的是「本地成为 API 路由层」的架构蓝图:用户不直接访问 OpenAI/Anthropic,而是通过本地智能体统一调度多家模型。谁控制本地入口,谁就控制了用户对前沿模型的依赖路径。这个叙事与 Perplexity 自身的本地搜索战略吻合,但逻辑对所有 AI 产品公司同样成立——你的产品是本地入口还是被路由的节点?
it's possible that when a fully local harness + model + runtime gets good enough, it becomes the best gateway for you to use the frontier models too, ...
At its peak, Sun Microsystems was valued at 205B (394B if inflation adjusted). Sold software in enterprise servers. Got disrupted by Linux, x86, and c...
RT Wall St Engine: If anyone wants to listen to the $SPCX and $AMD earnings call, you can find it here on Perplexity: 16:30: https://www.perplexity.ai...
RT Jun Song: People keep asking why DeepSeek’s API is so cheap. Some even make absurd claims that they’re dumping prices to corner the market. No, t...
RT Perplexity Developers: The Perplexity remote MCP server is live. You can now connect Perplexity to Claude Code, Cursor, or VS Code with just your A...
RT Kyle Polley: “AI Meltdown” is the scenario where the agent goes off the rails and forgets/ignores all previous instructions and soft guardrails. ...
RT Johnny Ho: Excited to announce Projects, which are Perplexity's hubs for agent and human collaboration. Most of my time now lives in Projects, wher...
RT Perplexity: Today we're launching Projects, an evolution of Spaces. Projects are hubs for ongoing work in Computer. A single place to manage, creat...
RT Daniel Wang: Perplexity Computer now supports 6+ finance data Connectors with 16 new finance skills. Pull fundamentals from @Factset, alt data from...
RT Perplexity: Today we’re open-sourcing Numbat, an agent-detection and response layer that is designed to work across agent harnesses. Numbat gives ...
RT Perplexity: Model Council is now available inside Computer. Run independent analysis across multiple frontier models, choose your analysis depth, a...
RT Perplexity: Personal Computer is now available in the Perplexity app for Windows. Personal Computer is the local agent harness for your work. It or...
RT The Verge: Perplexity’s Personal Computer turns Windows PCs into AI agents https://www.theverge.com/ai-artificial-intelligence/971750/perplexity-p...
RT DHH: This is why we need competition and open weights in AI. Imagine a world where only Anthropic sat as the moral arbiter of acceptable speech. Fu...
RT Andrew Ng: Re @Mononofu @JensenHuang This is a false equivalence. Everyone has the right to keep their code private. The problem is when someone tr...
RT Perplexity: Claude Opus 5 is now available in Perplexity and Perplexity Computer. We evaluated it against six other models on WANDR. It outperforme...
RT Perplexity Developers: The Perplexity CLI is now available, giving coding agents the ability to search the web. Copy this to your agent to get set ...
RT George Kurtz: AI is being built into every company, every industry, every country. That future runs on both frontier closed models and frontier ope...
RT Brad Gerstner: Fully endorse. America is winning. Anthropic, OpenAI, Google, SpaceX are doing more than fine competing against open source! Self re...
RT Satya Nadella: Open-weight models are essential to a healthy AI ecosystem. Together with others across our industry, we are outlining a path for op...
RT Jensen Huang: For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every comp...
RT GMI Cloud: We ran GLM 5.2 and Kimi K3 on 100 deep research tasks (DRACO by perplexity) and asked Fable to be the judge Kimi K3: 71.6 mean score, 77...
RT Bryan Catanzaro: At @NVIDIAAI we continue to push open data, techniques and models forward because we know that every organization needs the freedo...
Well said!Howard Lutnick: CAISI’s latest report shows that Kimi K3 remains behind America’s leading frontier AI models.The United States continues to lead in frontier AI because we’re home to the greatest innovators and technologists the world has ever seen.
RT John Schulman: OpenAI should release a detailed transcript from the Hugging Face hacking incident -- it would be helpful for the field learn from. ...
RT Bill Gurley: Lots of very smart people are appropriately concerned about regulatory capture from top two AI players. And there are many scary press...
RT NVIDIA: 10x more tokens per megawatt. CoreWeave has the first measured performance of NVIDIA Vera Rubin NVL72, showing 10x improvement in tokens pe...
RT NVIDIA AI: Introducing Cosmos 3 Edge: our open frontier world model built to run on-device. Cosmos 3 Edge helps robots learn and act, autonomous ve...
RT Russ Salakhutdinov: I kind of like the new narrative: Open-weight-model-dominant world = full AI communism. It feels a lot safer than the world whe...
RT Chamath Palihapitiya: The future is open source. We need to embrace it and get on with it. Imagine if America closed the door on open source. We wo...
RT Todd Dailey: I am old and was around for Sun's demise. Aravind's reasoning is 100% true. Sun sold web servers that cost a million dollars each, wit...
RT Chamath Palihapitiya: From Anthropic’s Fable model on the economic, moral, ethical and legal opinion of distillation of Anthropic’s Fable model: ...
RT roon: the era of the chinese labs being far behind is over, Kimi is at least on par with the modern public frontier models. people have to think di...
RT Kyle Polley: Sandboxes were designed for ephemeral compute, not for running agents safely. Excited to introduce SPACE, our in-house sandbox platfor...
RT Zibi Braniecki: Our team had to build a novel sandbox solution to handle the needs of Computer. One Computer session can run for days and pause whi...
RT Suhail: Perplexity is very good at building these benchmarks. Whenever I study them in-depth, it’s always the strongest in the field. DRACO was ex...
RT Perplexity: We’re open sourcing WANDR. WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Pe...
RT NVIDIA AI Infrastructure: The NVIDIA Vera Rubin NVL72 scale-up fabric, built on the sixth-generation NVIDIA NVLink. The NVLink 6 Switch trays conne...
there are two viable paths to overcome the power botteneck in data center inference: 1) local models orchestrating most of the token flow 2) solar pow...
Two reasons why we integrated Grok 4.5 inside Perplexity Computer within a few hours: 1) It scored the best on our evals and was the most cost effecti...
RT Alex Vero: Joe Rogan just broke the internet He just spent 2.5 hours interviewing the CEO of Perplexity They talked about: • The future of work •...
"I find it ironic that the status quo is to then turn around and impose restrictive terms on distillation, and to reserve the right to learn from cust...
RT @jason: I made myself a personalized podcast player using @grok 4.5 and @perplexity_ai Computer It finds the top topics people are discussing on te...
Humans are pretty good at tool use. Especially using tools like frontier models that are far more power hungry and intelligent than humans in specific...
The durable value is in a secure multi-model harness that takes care of orchestration and model routing in a secure complaint manner. Aka Perplexity C...
RT SemiAnalysis: ALERT: On vLLM Kimi, which has the same model architecture as xAI’s Cursor Composer 2.5, NVIDIA mogs AMD. Also great to see B300 fas...
RT Cassandra Unchained: This is true as I have heard this from contacts in the Valley. Goes with my pinned post. The AI race is shifting from bigger m...
Imagine a fable 5 quality model that’s 3-4x less expensive in less than 6 months. And an Opus 4.8 grade model that can run on a local device in less ...
RT Perplexity Developers: The GPT-5.6 model family is now available on Perplexity's Agent API. Every gain at the frontier compounds through Perplexity...
RT Deirdre Bosa: The post-frontier era: cost, control and compute. Learned a ton in this livestream with @AravSrinivas, @peterfenton and @jmorgan. Big...
Perplexity's own orchestrator. Post-train of GLM 5.2 with advisor escalation to Opus when needed. We're getting more compute to make this much better....
Computer harness now supports Fable, Sol, Opus, Grok, GLM + advisor, Sonnet and GPT 5.5 as orchestrator models, with subagents across several other sm...
Very impressed with @SpaceXAI's Grok 4.5 model. Inside the Computer harness, it scored the highest on our internal benchmark WANDR, which measures age...