57:16Open Models Change The Economics of AI
AI Model Trends Shift towards open models, especially in enterprise, driven by coding agents and AI assistants. Chinese origin models are currently dominant in cloud-hosted open model usage. US and German businesses are significant users of open model tokens via cloud services. Key Drivers for Open Model Adoption Cost reduction is the primary pain point solved by open models. Businesses seek better control and customization of AI for unique use cases. Example: AT&T shifted 40% of token consumption to open models, predominantly for coding agents. Open Model Capabilities & Growth Explosive growth in token usage per developer, driven by coding agents and workflow automation tools like OpenClaw and Hermes. Context windows have expanded significantly, enabling more complex tasks. Olama Cloud has seen 150x growth since the start of the year. Fine-tuning vs. Out-of-the-Box Models Early 2024 saw interest in fine-tuning, which waned but is now returning. The accelerating release cadence of open models makes custom training harder. Improved tooling now supports fine-tuning efforts. AI Safety and Security AI safety is a growing concern, potentially slowing frontier model development. Security and safety are major blockers for enterprise adoption of open models. Open models can be used for security testing where closed models refuse. Hugging Face used open-weight models to detect hacks. Olama's Role and Technical Challenges Olama acts as an "operating system" for AI, integrating hardware, inference engines, and application runtimes. Ensuring fast and accurate model performance across diverse hardware is a key challenge. A "playbook" exists for day-zero model launches, focusing on harness support, use cases, and hardware optimization. Challenges include model architecture changes, tool calling mechanics, and rigorous benchmarking. Ecosystem Dynamics: Bundling vs. Unbundling Trend towards specialized "best of breed" companies for knowledge, coordination, and execution layers, rather than bundled solutions. Open source fosters competition and innovation in these unbundled areas. Future of AI Workflows "Hidden layers" between models and applications offer opportunities for new companies. The next scarcity is above tokens: orchestrating agents and complex workflows. "Flash models" (e.g., Deepseek Flash) offer ultra-low cost per token and task, enabling "unlimited tokens" for common use cases. Chaining cheaper models together can solve complex problems through orchestration. Most AI tasks may not require "god models"; "good enough" models coupled with orchestration are sufficient. Geopolitics and Model Origin Customer concern varies: some prioritize model origin, others focus on where/how it's run. Data origin and model "voice" (communication style) are important factors. For critical tasks (e.g., power plant analytics), model origin is paramount for trustworthiness. IT and security teams in large enterprises are experienced with managing dependencies, similar to open-source software security. Olama's Origin and Growth Founded by ex-Docker developers, focused on developer experience. Pivoted to open-source LLMs after realizing the difficulty of running them locally. Rapid growth on GitHub (100k stars quickly) indicated strong product-market fit. Transitioned from hobbyist users to 85% of Fortune 500 adoption in approximately 18 months. Monetization Strategy Initially focused on a privacy-focused AI product. Waited for product-market fit with open models for complex use cases (e.g., coding agents). Monetization began with Olama Cloud, leveraging the abundance of open models. Emphasis on customer connection and understanding evolving needs. Local vs. Cloud Models A hybrid approach is expected, with easier tasks running locally for lower latency and cost. Modern hardware (Apple Silicon, Nvidia DGX Spark) can run large parameter models effectively. Coding agents often perform best with large cloud models due to task complexity. Document processing and simpler workflows run well locally. This hybrid model further reduces costs by utilizing existing hardware. GPU Market and Infrastructure High demand and volatility in the GPU market. Access to high-end GPUs (e.g., B200, B300) is challenging for startups. Inference providers are crucial for accessing and scaling these resources.















































