# BearPlex: Full Site Content (LLM-Friendly Format) > This file contains the complete public content of bearplex.com in clean markdown without JavaScript. > Built dynamically from the site's data files so it stays in sync with the live site. > For the structured overview, see https://www.bearplex.com/llms.txt Last updated: 2026-09-10 --- ## About BearPlex BearPlex is a custom software development and AI engineering company founded in 2017 by Hamad Pervaiz. We design and build SaaS products, internal tools, enterprise platforms, and production AI systems. ### What Makes Us Different - **No prototypes**: We only build production systems meant to run forever. - **War Room model**: 90-day embedded sprints. Cross-functional pods ship production code from day one. - **Outcome-based pricing**: Compensation is tied to measurable business outcomes, not billable hours. - **Sovereign deployment**: All systems deploy on the client's infrastructure with full source code handover. - **65 people, around 45 of them engineers**: Headquartered in Lahore, Pakistan, with a US entity, BearPlex Technologies Inc. - **Credentials**: SOC 2 Type II audit underway, and a verified Clutch profile rated 5.0 by client reviews. ### Methodology: The War Room Every engagement begins with "Day Zero", a single day where the entire team aligns on three things: 1. What metrics define success 2. What's the first thing that ships 3. Who can make decisions From Day One, engineers write production code. No prototypes, no proofs of concept, just actual systems that will run in the client's infrastructure. ### Engagement Structure - **Phase 0**: Discovery diagnostic (1-2 weeks): detailed specification, success criteria, risks. The first conversation is with an engineer, not an account manager. - **Phase 1+**: Outcome milestones with clear deliverables, success criteria, fixed price per milestone - **Risk Pool**: 15-25% of total project value tied to high-level business outcomes measured 90 days post-launch ### Leadership **Hamad Pervaiz**, Founder & CEO - Architect, investor, and strategist with 15+ years in business-critical software - Also: Managing Partner at Turing Venture Capital, Founder of PeoplePlus, Fractional CTO at The MediaGale - Profile: https://www.bearplex.com/authors/hamad-pervaiz - Personal site: https://hamadpervaiz.com - LinkedIn: https://www.linkedin.com/in/hamadpervaiz **Alice Chen**, Head of North America - Runs BearPlex across the United States and Canada - Former M&A and securities attorney turned venture and private-equity investor; Partner at Pender Hastings Capital, board member at EO Silicon Valley - Has worked with Y Combinator-backed companies and EO and YPO member businesses - Profile: https://www.bearplex.com/authors/alice-chen **M. Irshad Kanwal**, Chief Financial Officer - Five years as CFO at BearPlex, owning engagement economics and outcome-based pricing models - Came up through engineering, including Principal Software Engineer at TRG (The Resource Group) --- ## Pricing Philosophy BearPlex operates on outcome-based pricing exclusively. We do not bill by the hour. Every engagement is structured as fixed-price milestones plus an outcome-tied risk pool. ### Engagement Tiers | Tier | Duration | Price | Description | |------|----------|-------|-------------| | Discovery Sprint | 1 to 2 weeks | $2,500 to $5,000 | Scoping diagnostic. Detailed spec, success criteria, risks, go/no-go recommendation. | | Single Service | 4 to 12 weeks | from $15,000, typically $25,000-$80,000 by archetype | Focused engagement on one BearPlex service. Production deployment + 30 days support. | | Enterprise Platform | 6 to 12 months | Multi-phase programs, scoped from the same bands | Multi-service platform builds at PeoplePlus scale. | | Integrated Teams | 3 to 24 months | $9,500/month per pod, larger pods to about $20,000/month | Embedded engineering pods, priced per pod. 3-month minimum, then month-to-month. | ### Pricing FAQs **Q: Why doesn't BearPlex publish hourly rates?** Hourly billing structurally rewards inefficiency: the longer a project takes, the more we earn. We refuse to operate that way. Our compensation is tied to outcomes the client cares about, not hours we burn. **Q: How does outcome-based pricing actually work?** Each engagement is structured as fixed-price milestones tied to specific deliverables and success criteria. A typical structure is 75-85% milestone payments + 15-25% risk pool tied to high-level business outcomes measured 90 days post-launch. **Q: What does a typical BearPlex engagement cost?** BearPlex prices by outcome, not by hour, from published bands. Marketing websites start at $2,000 on WordPress (typically $3,000-$8,000) and $4,000 on Next.js (typically $6,000-$15,000). E-commerce storefronts start at $4,500 (typically $8,000-$25,000). SaaS MVPs and custom platforms start at $15,000 (typically $25,000-$80,000). AI systems and agent builds start at $15,000 (typically $25,000-$75,000). Custom internal tools, CRMs, and ops systems start at $15,000 (typically $25,000-$70,000). Multi-phase enterprise programs range higher. Every engagement starts with a discovery diagnostic; a 1-2 week discovery sprint runs $2,500-$5,000. Integrated Teams run on a monthly retainer: embedded engineers from $3,000/month per role, a full pod (2 senior devs + PM + QA) at $9,500/month, larger pods to about $20,000/month, with 3-month minimums, then month-to-month. Care plans start at $150/month (typically $350-$2,000/month). --- ## Services ### Custom Software Development Company BearPlex designs and builds SaaS products, internal tools, customer portals, commerce systems, and integrations around the workflow a business actually runs. Custom platforms start at $15,000 and typically range from $25,000 to $80,000. Internal tools, CRMs, and operations systems start at $15,000 and typically range from $25,000 to $70,000. A 1 to 2 week discovery sprint costs $2,500 to $5,000 and ends with a fixed milestone plan and build price. Standard delivery includes source repositories, deployment access, architecture documentation, and operating runbooks so the client can own and extend the system. URL: https://www.bearplex.com/services/custom-software-development #### Frequently Asked Questions **Q: What is custom software development?** Custom software development is the design and engineering of a system around one organization's workflows, rules, users, data, and integrations. Unlike an off-the-shelf product, the system is shaped around the way the business creates value and can evolve with that operating model. BearPlex builds SaaS products, internal tools, customer portals, commerce systems, and integrations when a mature product cannot meet the real requirement. **Q: How much does custom software development cost?** A BearPlex SaaS MVP or custom platform starts at $15,000, with most projects landing between $25,000 and $80,000. A custom internal tool, CRM, or operations system also starts at $15,000 and typically runs $25,000 to $70,000. E-commerce storefronts start at $4,500 and typically run $8,000 to $25,000. The final price is fixed after discovery and depends on workflow breadth, user roles, integrations, data migration, compliance, and timeline pressure. Project builds are priced as fixed milestones tied to agreed deliverables. **Q: How long does it take to build custom software?** The honest answer depends on the number of workflows, roles, integrations, and migration risks. We start with a 1 to 2 week discovery sprint that produces the architecture, milestone plan, and fixed build price. Your plan is based on the smallest production scope that can prove the business case, not an arbitrary calendar promise. Larger systems are split into production milestones so the team can inspect working software before the entire roadmap is complete. **Q: When should we build instead of buying an off-the-shelf product?** Build when the workflow is a real competitive advantage, several disconnected systems are creating operational risk, your data or compliance requirements exceed what a vendor can support, or vendor economics become unreasonable at your scale. Buy and configure when the workflow is standard and a credible product already solves it. If a subscription is the better answer, we will tell you before proposing a build. **Q: Who owns the source code and intellectual property?** Our standard delivery gives your team the source repositories, deployment access, architecture documentation, and operating runbooks. Project-specific intellectual property is assigned through the engagement agreement, while third-party and open-source licenses are disclosed rather than hidden inside the handover. The goal is simple: your team must be able to operate and extend the system without depending on BearPlex. **Q: Can BearPlex work with our existing product or engineering team?** Yes. We can own a complete workstream or embed alongside your team. We use your planning and code-review process where that is the better operating model, document decisions as the build progresses, and design the handover from the first milestone instead of leaving knowledge transfer until the final week. **Q: Can you modernize or take over an existing system?** Yes, but we do not begin with a rewrite assumption. A takeover starts with the current architecture, user-critical workflows, data model, deployment path, security exposure, and change history. We then recommend one of three paths: stabilize what exists, replace it in slices, or rebuild only where the evidence justifies it. Preserving working behavior and search equity is part of the plan. **Q: What happens after launch?** Every launch includes production deployment, acceptance checks, documentation, and a handover walkthrough. You can operate the system internally, retain BearPlex for product development, or use a care plan for monitoring, maintenance, and smaller improvements. Care plans start at $150 per month and typically range from $350 to $2,000 per month, depending on the support surface. **Q: Which technologies do you use for custom software?** We choose the stack around the product and the team that will inherit it. Common choices include React or Next.js for product interfaces, PostgreSQL or Supabase for transactional data, Node.js, Python, or Go for services, and Vercel or AWS for deployment. Those are options, not a prescribed bundle. Existing team expertise, security requirements, data shape, latency, and operating cost decide the architecture. **Q: How do you handle security and compliance?** Security is designed into the data model, identity system, permissions, audit trail, deployment, and release process. The exact controls depend on the product and its obligations, so we map them during discovery instead of attaching a generic compliance badge. BearPlex was founded in 2017; its U.S. entity is BearPlex Technologies Inc.; its SOC 2 Type II audit is underway. --- ### Autonomous AI Agents & Workflows Deploy autonomous AI agents that execute complex workflows end-to-end. Multi-agent orchestration, LangChain integration, and zero-touch operations for enterprise. URL: https://www.bearplex.com/services/autonomous-agents #### Frequently Asked Questions **Q: What's the difference between an autonomous agent and a chatbot?** A chatbot responds to user messages. An autonomous agent plans multi-step work, calls tools, reflects on its progress, and completes the goal end to end, escalating only the actions you have gated, like moving money, touching PII, or changing access. We build production agent systems using LangGraph, CrewAI, and the Claude Agent SDK, not simple chatbots. **Q: How long does a BearPlex autonomous agent engagement take?** Twelve weeks from kick-off to production. Week 1-2: scoping and architecture. Week 3-8: agent development, tool integration, eval harness. Week 9-12: production hardening, observability, handover. **Q: What does an autonomous agent project typically cost?** BearPlex prices by outcome, not hours. A one-week discovery produces the scope, the eval criteria, and a fixed quote for the whole build. The number depends on how many tools the agent touches, how many systems it integrates with, and the eval rigor your domain demands. The agent runs in your cloud on your LLM spend, so there is no per-seat meter. Tell us the workflow and we will scope it. **Q: Which LLM frameworks do you work with?** LangGraph (our default for stateful workflows), CrewAI for multi-agent orchestration, the Claude Agent SDK for Anthropic deployments, LangChain for simpler chains, and native function-calling for tight integrations. We pick based on your stack, not vendor affinity. **Q: Do you deploy on our infrastructure or yours?** Always yours. BearPlex's sovereign deployment model means the agent runs in your VPC, on your LLM provider of choice (OpenAI, Anthropic, Google, AWS Bedrock, Azure OpenAI, or on-prem Llama/Mistral). We hand over full source code and runbooks at engagement end. **Q: How do you evaluate agent reliability?** Every BearPlex agent ships with an evaluation harness: golden task datasets, LLM-as-judge scoring, regression tests, and observability (OpenTelemetry + LangSmith or Arize). We target 95%+ task completion and 99%+ safety compliance before production cutover. --- ### Model Engineering & Fine-Tuning Custom LLM fine-tuning and model engineering services. Deploy sovereign AI models on your infrastructure with LoRA, QLoRA, and PEFT techniques. URL: https://www.bearplex.com/services/model-engineering #### Frequently Asked Questions **Q: When should we fine-tune a model vs just prompt-engineer?** Fine-tune when: (1) you need consistent output format that prompts can't enforce, (2) you have >1000 high-quality training examples, (3) latency or cost of a large model is prohibitive and a fine-tuned small model suffices. Otherwise, invest in prompt engineering and RAG first, cheaper and faster to iterate. **Q: What's the difference between LoRA, QLoRA, and full fine-tuning?** Full fine-tuning updates all model weights (best quality, highest cost, needs huge GPU clusters). LoRA freezes the base model and trains small adapter matrices (90% of the quality, 1% of the cost). QLoRA quantizes the base model to 4-bit and trains LoRA adapters (fits on a single consumer GPU). We use LoRA for most production work. **Q: Which open models do you fine-tune?** Llama 4 (Meta license), Qwen 3 (strong reasoning), Gemma 3 (Google), Phi-4 (SLM for edge deployment), DeepSeek, and Mistral (commercial use allowed). For closed models we fine-tune via OpenAI and Anthropic fine-tuning APIs where available. **Q: How much training data do we need?** For style/format fine-tuning: 500-2,000 high-quality examples typically suffice. For knowledge injection: you probably want RAG instead of fine-tuning. For domain reasoning: 5,000-50,000 examples + synthetic data augmentation. Quality >> quantity: 500 curated examples beat 50,000 messy ones. **Q: Where do you deploy fine-tuned models?** Sovereign deployment by default: your VPC on AWS/GCP/Azure via vLLM, TGI, or TensorRT-LLM. Anyscale or Modal for serverless inference. Together.ai or Fireworks for managed endpoints. Full stack on-prem for regulated industries. We optimize for your specific latency/cost/compliance constraints. **Q: What does a fine-tuning engagement cost?** Fine-tuning engagements run a 4-8 week cycle at a fixed scope quoted up front. Included: the data curation pipeline, training infrastructure setup, LoRA/QLoRA training, the evaluation harness, deployment automation, and handover. Compute costs are passthrough at our discounted GPU rates. --- ### RAG & Knowledge Systems Enterprise RAG systems that eliminate AI hallucinations. Vector databases, hybrid search, citation-tracked retrieval, and role-based access control for your knowledge base. URL: https://www.bearplex.com/services/rag-knowledge #### Frequently Asked Questions **Q: What is a RAG (Retrieval Augmented Generation) system?** RAG combines a search system (retrieval) with an LLM (generation). When a user asks a question, relevant documents are retrieved from your knowledge base and injected into the LLM's context, so answers are grounded in your actual data, not hallucinated. **Q: Why do most enterprise RAG projects fail?** Three reasons: (1) they treat retrieval as solved (it isn't: chunking, reranking, and hybrid search matter enormously), (2) they skip evaluation (no golden dataset means no accuracy measurement), (3) they forget access control (role-based document permissions must propagate through retrieval). BearPlex engineers all three. **Q: What's the difference between RAG and fine-tuning?** RAG injects knowledge at query time (good for frequently-changing data, specific facts, permissions). Fine-tuning bakes knowledge into the model weights (good for style, format, implicit behaviors). Most enterprise deployments use both: fine-tune for tone, RAG for facts. **Q: Which vector databases do you work with?** Pinecone for managed production deployments. Qdrant for self-hosted. Weaviate for hybrid search at scale. pgvector for teams already on Postgres. Elasticsearch for teams already on Elastic. We benchmark for your specific workload before choosing. **Q: How do you handle document access control in RAG?** We implement filter-first retrieval: before any similarity search, we filter the index by user permissions (RBAC, ABAC, or row-level security depending on your IAM). This ensures an engineering user can't retrieve HR documents, even if they're the most semantically relevant result. **Q: What eval framework do you use for RAG accuracy?** RAGAS for automated metrics (faithfulness, answer relevancy, context precision/recall). LLM-as-judge for nuanced quality. Human eval on a sampled golden set for calibration. We target >90% faithfulness and >85% answer relevancy before production. --- ### RLHF & AI Alignment Services Enterprise RLHF alignment services to make AI systems safe, truthful, and obedient. Expert human feedback pipelines, reward modeling, and red team validation. URL: https://www.bearplex.com/services/rlhf-alignment #### Frequently Asked Questions **Q: What is RLHF (Reinforcement Learning from Human Feedback)?** RLHF is the post-training technique that makes language models helpful, honest, and safe. Human annotators rank model outputs by quality; those rankings train a reward model; the reward model then fine-tunes the base LLM via reinforcement learning. It's why ChatGPT feels different from raw GPT-4: RLHF aligned it to human preference. **Q: When does an enterprise need RLHF vs just prompt engineering?** Prompt engineering covers 80% of use cases. RLHF is needed when: (1) you have a narrow domain where generic models underperform (legal drafting, medical diagnosis), (2) you need consistent tone/style that prompting can't reliably enforce, (3) you have 5K+ quality preference labels or budget to generate them. **Q: What are DPO and ORPO compared to traditional RLHF?** DPO (Direct Preference Optimization) skips the reward model, training the LLM directly on preference pairs. Simpler, cheaper, often equivalent quality. ORPO combines SFT and alignment in one pass. For most enterprise work, we use DPO or ORPO, not classical PPO-based RLHF. **Q: How do you source human feedback at scale?** Internal SME annotation for specialized domains (legal, medical, financial). Scale AI or Surge AI for general preference labels. Our own 12-person annotation team in Lahore for bilingual and domain-specific work. Quality control via inter-annotator agreement and adversarial probe sets. **Q: What's red-teaming and why does it matter?** Red-teaming is adversarial testing: deliberately trying to make the model produce harmful, biased, or incorrect output. Every BearPlex alignment engagement includes a red-team pass covering: prompt injection, jailbreaks, bias probes, refusal consistency, and safety boundary tests. We ship alignment with a documented failure-mode inventory. **Q: How long does an RLHF engagement take?** 12-20 weeks end-to-end: 2 weeks data strategy, 4-8 weeks preference annotation, 2-4 weeks reward modeling + RLHF/DPO training, 2-4 weeks evaluation and red-teaming, 2 weeks deployment and monitoring. Pricing is fixed-scope and depends mostly on annotation scale. --- ### Enterprise Platform Engineering Full-stack enterprise platform development. Microservices architecture, event-driven systems, API-first design, and cloud-native infrastructure built for scale. URL: https://www.bearplex.com/services/enterprise-platforms #### Frequently Asked Questions **Q: What does 'enterprise platform' mean at BearPlex?** An enterprise platform is a multi-tenant, multi-service system that forms the operational spine of an organization, handling auth, billing, permissions, workflows, and integrations across dozens of modules. BearPlex builds platforms at the scale of PeoplePlus (15+ modules, 130+ database tables, 143K LOC), not single-purpose apps. **Q: How long does it take to build an enterprise platform?** Production-ready platforms typically take 6-12 months across our War Room model: 90-day sprints stacked with clear deliverables. The first 90 days ship a usable v1 (core workflows live). Subsequent sprints expand modules, harden scale, and integrate with existing enterprise systems. **Q: What stack do you use for enterprise platforms?** Next.js + TypeScript for the frontend, Postgres + Prisma for data, Node/Go/Python microservices for domain logic, Redis for caching, Kubernetes or AWS ECS for orchestration, Auth0/Okta/Keycloak for identity. We pick based on your existing stack, not vendor affinity. Zero interest in rebuilding what works. **Q: Do you integrate with Salesforce, HubSpot, Snowflake, Databricks, SAP?** Yes. Every enterprise platform we ship integrates with the client's existing CRM, data warehouse, and SaaS stack. We've built integrations with Salesforce, HubSpot, Snowflake, Databricks, SAP, Workday, NetSuite, Zendesk, Intercom, Slack, and dozens more. Integrations are built as first-class citizens, not afterthoughts. **Q: Can you work alongside our existing engineering team?** Yes, this is our Integrated Teams model. BearPlex engineers embed with your team, use your tooling, follow your code review process, and transfer knowledge continuously. You end the engagement with a platform AND an internal team capable of operating it. **Q: What does a platform build cost?** Platform builds are scoped to the outcome, not billed by the hour. The number depends on module count, integration surface, and compliance requirements, and a portion of the fee is tied to platform performance measured post-launch: uptime, latency, and user adoption. Tell us the platform and we will scope it. --- ### Interface Design & UX Engineering Psychology-first interface design for AI-powered products. Atomic design systems, dev-ready Figma files, and performance-optimized UI components. URL: https://www.bearplex.com/services/interface-design #### Frequently Asked Questions **Q: What makes interface design for AI products different?** AI products have probabilistic output, high latency, and frequent failure modes. Good AI UX acknowledges all three: transparent confidence, progressive disclosure of generation, graceful fallback when the model refuses or errors. We design for AI's actual behavior, not idealized behavior. **Q: Do you deliver design systems or one-off designs?** Design systems by default. Every engagement produces atomic components (Radix or shadcn-based), token definitions (colors, spacing, typography), Figma libraries, and Storybook documentation. Your team can extend the system long after BearPlex exits the engagement. **Q: How do you handle accessibility in AI interfaces?** WCAG 2.1 AA as a floor, not a ceiling. Streaming text needs screen-reader-friendly rendering. Model confidence indicators need non-color redundancy. Tool-calling UIs need keyboard navigation. Every BearPlex design ships with accessibility tested by actual users on screen readers, not just axe DevTools. **Q: Do you design for mobile, desktop, or both?** Both by default. Mobile-first for consumer AI. Desktop-first for enterprise AI tooling (CTOs don't configure RAG pipelines on their phones). Responsive doesn't mean identical. We design distinct affordances for each context. **Q: What's your design + engineering handoff process?** Figma with dev-ready components, auto-generated CSS tokens, and live component specs via Figma Dev Mode. We deliver working Storybook alongside Figma. Zero 'it looks different in production' because designers and engineers work in the same Slack channel, the same ticket system, and the same sprint cadence. **Q: Can you audit and improve an existing product?** Yes. Audit engagements are 2-3 weeks: usability testing with 5-8 real users, heuristic review, accessibility scan, design system health check. Deliverable: ranked list of issues, mockups of the top 10 fixes, and a 90-day improvement roadmap, at a fixed price scoped up front. --- ### Data Pipelines & MLOps Industrial-grade data refineries and MLOps infrastructure to feed, train, and monitor your intelligence at scale. Airflow, Dagster, Feature Stores, and CI/CD for ML. URL: https://www.bearplex.com/services/data-pipelines #### Frequently Asked Questions **Q: What is an MLOps pipeline?** An MLOps pipeline is the automated system that moves data from source → feature engineering → model training → evaluation → deployment → monitoring. It's the production infrastructure that turns a model from a Jupyter notebook into a reliable service that retrains on new data and catches drift before customers feel it. **Q: Which orchestration tools do you use?** Airflow for traditional scheduled DAGs, Dagster for modern asset-based pipelines, Prefect for Python-native workflows, Temporal for long-running stateful workflows. We pick based on team familiarity and pipeline complexity, not vendor preference. **Q: What's your approach to feature stores?** For most teams: start with Postgres + materialized views (simple, cheap, sufficient). Scale to Feast or Tecton when you need real-time feature serving across multiple models. Skip feature stores entirely if you have only one model and one team. Premature abstraction kills velocity. **Q: How do you monitor LLM and ML models in production?** We deploy LangSmith + Arize + OpenTelemetry for LLM observability (prompts, latency, token usage, hallucination rates). For traditional ML: Evidently AI or Arize for data drift, WhyLabs for ML health. Every BearPlex pipeline ships with dashboards and alerting before cutover. **Q: Can you migrate legacy ETL to modern data pipelines?** Yes. Common migrations: Informatica/IBM DataStage → Airflow/Dagster, SSIS → dbt + Airflow, home-grown cron scripts → proper orchestration. We do parallel-run cutover (old and new systems run side-by-side until trust is established) to de-risk migrations in regulated environments. **Q: What's your CI/CD for ML?** GitHub Actions or GitLab CI for the pipeline itself. MLflow or Weights & Biases for experiment tracking. Model registries (MLflow, SageMaker Model Registry) for versioning. Deployment via BentoML, Seldon Core, or native SageMaker/Vertex AI endpoints. Everything Git-versioned, everything reproducible. --- ### Sovereign Cloud Infrastructure Deploy AI systems on sovereign cloud infrastructure with full data residency compliance. Air-gapped environments, on-premise GPU clusters, and government-grade security for sensitive workloads. URL: https://www.bearplex.com/services/sovereign-cloud #### Frequently Asked Questions **Q: What is sovereign AI deployment?** Sovereign AI deployment means running AI systems entirely within your own data boundaries: your VPC, your on-premise GPUs, your compliance perimeter. No data leaves your infrastructure. Required for regulated industries (healthcare, finance, defense, government) and any organization with strict data residency requirements. **Q: Which LLMs can run sovereign?** Open models: Llama 4, Qwen 3, Gemma 3, Phi-4, DeepSeek, Mistral. Managed sovereign options: AWS Bedrock (in-region), Azure OpenAI (in-region), GCP Vertex AI (in-region). Enterprise-licensed closed models through private deployments (Anthropic Enterprise, OpenAI Enterprise) are also supported depending on your agreement. **Q: How do you handle air-gapped deployments?** An air-gapped design can package the model runtime, retrieval layer, dependencies, and observability for offline operation, with updates moving through a controlled artifact-transfer process. The exact architecture, signing process, update cadence, and evidence requirements are scoped against the client's environment during discovery. We do not claim a generic air-gapped configuration is suitable for every regulated workload. **Q: What GPU infrastructure do you work with?** On-premise: NVIDIA A100, H100, H200, L40S clusters. Cloud: AWS p4d/p5, GCP A3, Azure NDv5. Edge: Jetson Orin for embedded, Mac Studio (M3 Ultra) for small-scale LLM inference. We size based on concurrency, latency SLA, and model size: right-sized clusters, not over-provisioned spend. **Q: Which compliance frameworks have you deployed under?** The required framework depends on the product, the data, the deployment region, and which organization is responsible for the control. We map those obligations during discovery and scope the technical controls and evidence the system must support. BearPlex's own SOC 2 Type II audit is underway; we do not present that work in progress, or a client's framework requirements, as a completed certification. **Q: What's the typical cost of sovereign AI deployment?** There are two cost surfaces: the infrastructure you pay for directly and the BearPlex engagement that designs, deploys, and hands over the system. The infrastructure number depends on model size, concurrency, latency, availability, and whether hardware is owned or rented. The engineering engagement is fixed after discovery; we model the managed-API alternative rather than promising a universal savings percentage. --- ### Integrated Teams: Embedded Engineering Dedicated engineering teams that embed into your organization permanently. Not contractors, but partners who understand your codebase, culture, and trajectory. URL: https://www.bearplex.com/services/integrated-teams #### Frequently Asked Questions **Q: How is Integrated Teams different from staff augmentation?** Staff augmentation sends bodies. Integrated Teams sends capability: a complete pod (engineers + lead + PM) with internal BearPlex support (architects, SRE, security) behind them. Our engineers act as your engineers for 6-24 month engagements: they learn your codebase, culture, and strategy. **Q: Who's on a typical Integrated Team pod?** Minimum viable pod: 1 tech lead + 2 senior engineers + 1 mid-level engineer. Typical pod: 1 tech lead + 3-4 senior engineers + 2 mid + 1 design lead + 1 product designer. Larger engagements scale up to 12-person pods with SRE and security specialists. **Q: Where are your engineers based?** BearPlex is headquartered in Lahore, Pakistan. The company has 65 people, with around 45 in engineering and development, and its U.S. entity is BearPlex Technologies Inc. Working overlap and communication cadence are agreed for each embedded team rather than implied by office locations. **Q: How do you hire engineers?** We hire at senior and staff level through a multi-stage process that includes live, client-style work, not just whiteboard puzzles. We screen as hard for communication and ownership as for technical depth, because an embedded engineer has to reason with your executives and own what they ship in production. That bar is why a pod is effective from its first week. **Q: What does Integrated Teams cost?** Pods run on a monthly retainer, priced per pod rather than per hour, with a three-month minimum and then month to month. Per engineer it typically lands 40 to 60 percent below equivalent US-based talent at the same skill level, and you own all the code, infrastructure, and IP from the first commit. Tell us the disciplines and timeline and we send back a scoped number. **Q: How do knowledge transfer and exit work?** Every Integrated Team engagement produces continuous documentation: architecture diagrams, runbooks, onboarding guides, code walkthroughs recorded on Loom. At engagement end, your internal team can operate everything BearPlex built. No vendor lock-in. --- ### Application Security & Penetration Testing Autonomous penetration testing that finds real vulnerabilities before attackers do. Source-code-aware, exploit-validated, OWASP-complete security testing for web apps, APIs, and cloud infrastructure. URL: https://www.bearplex.com/services/appsec #### Frequently Asked Questions **Q: Is this automated scanning or real penetration testing?** It's autonomous penetration testing reviewed by our security team. Our engine reads your source code, understands your business logic, and crafts real exploits. Our team then curates the findings, validates edge cases, and delivers a polished report, all within 24 hours. **Q: How is this different from tools like Snyk or SonarQube?** Those are SAST/SCA tools that find code patterns and known CVEs. We go further: we correlate source code analysis with runtime behavior, chain vulnerabilities together, and validate each finding with actual exploitation. A SAST tool might say 'possible SQL injection.' We say 'confirmed SQL injection: here's the payload that extracts your users table.' **Q: Do you do retesting after we fix vulnerabilities?** Yes. One-Time engagements include one re-test. Continuous Security includes unlimited re-tests. Our engine automatically verifies fixes are effective and haven't introduced regressions. **Q: How do you handle our source code and sensitive data?** All testing runs in isolated, ephemeral environments. Source code is encrypted in transit and at rest, processed in memory, and purged after the engagement. We sign NDAs before every engagement. No data is retained beyond the final report. **Q: What frameworks and languages do you support?** We support all major web frameworks: Next.js, React, Django, Rails, Laravel, Spring Boot, Express/Node.js, Go, Flask, FastAPI, and more. Our engine handles REST APIs, GraphQL, gRPC, WebSockets, and SAML/OAuth auth flows natively. **Q: Is the first assessment really free?** Yes. Submit your application details through the form below and we'll deliver a full pentest report within 24 hours at no cost. No strings attached. It's our way of showing you the quality of our work. If you want ongoing security after that, we'll talk. **Q: Can we see a sample report before engaging?** We provide redacted sample reports under NDA. These include real findings from internal engagements (with identifying details removed) so you can evaluate the depth and quality of our work before committing. Or just request the free assessment: your own report is the best sample. **Q: What's the difference between a vulnerability scan and a pentest?** A vulnerability scan runs automated checks against known patterns, like spell-checking for security. A pentest thinks like an attacker: it chains vulnerabilities, tests business logic, and proves impact. We deliver pentests, not scans. --- ### Supabase Development Agency Senior Supabase development: production builds, rescues and RLS hardening, multi-tenant row-level security architecture, edge functions, realtime, storage architecture, migrations in and out, and backup/DR with external pipelines to client-owned S3 and rehearsed restores. BearPlex runs Supabase in production for client platforms (CleverCoach, Letti AI, Yaay365, the Ticketmaster scraper) and for its own SaaS, PeoplePlus, whose public careers API runs on Supabase edge functions and powers the careers system on this site. URL: https://www.bearplex.com/services/supabase-development #### Frequently Asked Questions **Q: Do you take over existing Supabase projects?** Yes. A takeover starts with an audit: every table checked for row-level security coverage, every policy read against the roles that actually exist, auth flows traced end to end, missing indexes and slow queries profiled, and the backup posture verified, including whether a restore has ever actually been rehearsed. You get a written findings report first, then a fixed-scope hardening plan. The bar is the one we hold our own builds to: production systems with 13+ RLS-protected tables and multiple JWT roles. **Q: Is Supabase production-ready for enterprise use?** Yes, with discipline. Under the branding it is managed Postgres plus auth, storage, realtime, and edge functions, and Postgres is as enterprise-proven as databases get. Production readiness comes from what you add around it: RLS designed as a real security boundary, external backups (native backups do not cover Storage), point-in-time recovery sized to your actual recovery objectives, and observability. We run our own SaaS, PeoplePlus, on Supabase with multi-tenant RLS isolation, and we ship client platforms on it in production. **Q: Does Supabase back up Storage buckets?** No. As of mid-2026, Supabase's native backups and point-in-time recovery cover the Postgres database only; objects in Storage buckets are excluded, and Supabase documents this openly. The free tier has no automatic backups at all. For production clients we set up an external pipeline: scheduled database dumps plus a sync of every Storage bucket to an S3 bucket the client owns, and we rehearse the restore so recovery is a runbook, not a hope. **Q: Supabase vs Firebase: which should we choose?** For most B2B and SaaS products, Supabase. The core difference is the database: Supabase is real Postgres with SQL, joins, row-level security, and full portability (you can dump the database and leave), while Firestore is a proprietary document store that shapes your data model around its query limits. Firebase still wins for mobile-first apps that lean on its offline sync and for teams deep in the Google ecosystem. If your data is relational, and business data almost always is, start with Postgres. **Q: Supabase vs a custom Postgres backend: when is each right?** Supabase gives you Postgres plus auth, storage, auto-generated APIs, realtime, and edge functions on day one, which saves a small team months of undifferentiated backend work. A custom backend on RDS or Cloud SQL wins when you need everything inside your own VPC, deep control over the API layer, or infrastructure your platform team already operates. The good news: because Supabase is standard Postgres, choosing it is not a one-way door. We have designed systems on Supabase precisely because the exit path is a pg_dump away. **Q: How much does a Supabase build cost?** We price by outcome, not by the hour. A rescue and hardening engagement (RLS audit, auth review, backup pipeline, performance pass) is the smallest thing we sell and typically starts in the low five figures. A full product build on Supabase usually lands from the mid five figures, and platform-scale builds go beyond that. The drivers are the tenancy model (how many roles, how strict the isolation), the number of edge functions and third-party integrations, realtime surfaces, data migration volume, and compliance requirements. A short discovery sprint produces a fixed quote before any build work starts. **Q: Do you work fixed-scope or ongoing?** Both. New builds and rescues run as fixed-price milestones scoped from a discovery sprint, each with explicit deliverables and acceptance criteria. Platforms we have built or hardened can move onto an ongoing engineering retainer that covers new features, dependency and Postgres upgrades, RLS review for every new table, and backup restore rehearsals. Our own SaaS, PeoplePlus, is developed on the same continuous model, so the retainer discipline is one we live with ourselves. **Q: Can you migrate us off Supabase?** Yes, in either direction. Moving off is tractable because the database is standard Postgres: pg_dump or logical replication carries the data to RDS, Cloud SQL, or self-hosted Postgres. The real work is in the platform services: replacing Supabase Auth (exporting users and password hashes, reissuing JWTs), rehoming Storage objects, rewriting realtime consumers, and porting edge functions to your new runtime. We scope all four explicitly. Migrations onto Supabase, from Firebase or from legacy custom backends, follow the same discipline in reverse. **Q: How do you architect multi-tenant row-level security?** A tenancy key on every table, policies written per role against JWT claims, and no table left unprotected: RLS as the security boundary, not a decoration. Then we prove it, with per-role integration tests that attempt cross-tenant reads and writes and must fail. PeoplePlus, our own SaaS, isolates every organisation's data this way with four permission levels; CleverCoach runs three JWT-authenticated roles over 13+ RLS-protected tables. Policies that are never tested are policies you do not actually have. **Q: Do you work with self-hosted Supabase?** We deploy managed Supabase for production clients by default: the hosted platform's operational surface (upgrades, monitoring, failover) is most of what you are paying for. Self-hosting the open-source stack is possible and we will scope it when data residency genuinely requires it, but you inherit the ops burden for Postgres, Auth, Storage, and Realtime yourself. For most teams with residency constraints, the honest comparison is self-hosted Supabase vs plain Postgres in your VPC with a conventional backend, and we will tell you which one fits. --- ### Next.js Development Agency Senior Next.js development: App Router product builds, migrations from WordPress, Create React App, and the Pages Router, performance and Core Web Vitals engineering, programmatic SEO architectures, ISR and edge rendering strategy, e-commerce, and design-system implementation. The proof is bearplex.com itself: a Next.js 16 App Router build with 440+ indexable routes generated from typed data, careers pages on a five-minute ISR window against a live ATS, and an edge-runtime OG image service. Published Next.js case studies include Vertex360, Yaay365, and Optinizers, plus a delivered luxury designer-collections e-commerce build. Production builds deploy to Vercel daily. URL: https://www.bearplex.com/services/nextjs-development #### Frequently Asked Questions **Q: Do you migrate sites from WordPress or Create React App to Next.js?** Yes, both are standard engagements. A WordPress migration moves content into typed data files or a headless CMS, rebuilds the templates as App Router routes, and treats the redirect map as a first-class deliverable so existing search equity survives the cutover. A CRA or Vite SPA migration is usually incremental: routes move across in slices, client-only pages keep working, and each slice picks up server rendering and per-route metadata as it lands. In both cases the migration starts with an inventory of every URL that matters, because the pages you forget are the rankings you lose. **Q: App Router or Pages Router in 2026?** App Router for anything new. It has been the default since Next.js 13, the ecosystem has consolidated around server components, and the newest capabilities land there first. This page is served by an App Router build. That said, a working Pages Router app is not an emergency: it remains supported, and the honest advice is to migrate when a redesign, a major version upgrade, or a feature that needs server components gives you a reason, not because a blog post said so. When we do migrate one, it goes route by route, not as a rewrite. **Q: Do we need Vercel to run Next.js?** No. Next.js self-hosts on any Node server or container using standalone output. The honest tradeoff: on Vercel, ISR, image optimization, edge middleware, and preview deployments work with zero configuration; self-hosted, you own each of those, including a shared cache handler if you run multiple instances. We deploy production builds to Vercel daily, including this site, and when a client needs to self-host we plan the caching, image, and CDN pieces explicitly instead of discovering them in production. **Q: Is Next.js good for SEO and programmatic sites?** It is the strongest fit we know for that job, and this site is the receipt: bearplex.com is a Next.js App Router build with 440+ indexable routes, most of them programmatic pages generated from typed data files with generateStaticParams, each with its own metadata, JSON-LD, and Open Graph image. Server rendering means crawlers get complete HTML, ISR keeps generated pages fresh without full rebuilds, and the sitemap is code, so it never drifts from the routes that actually exist. **Q: How much does a Next.js build cost?** We price by outcome, not by the hour. A focused engagement (a performance and Core Web Vitals pass, or a migration assessment with a redirect map) is the smallest thing we sell and typically starts in the low five figures. A full product or platform build usually lands from the mid five figures, and large programmatic or e-commerce builds go beyond that. The drivers are the size of the page inventory, the number of integrations and data sources, e-commerce complexity, design-system maturity, migration and redirect volume, and how aggressive the performance targets are. A short discovery sprint produces a fixed quote before any build work starts. **Q: Do you offer ongoing Next.js retainers?** Yes. Retainers cover new feature development, dependency and framework upgrades, and Core Web Vitals monitoring against real field data. The upgrade line matters more than it sounds: Next.js caching behavior has changed across major versions, so version bumps get route-by-route verification, not a version number change and a prayer. Our own site runs on the same continuous model, so the discipline we sell is one we live with. **Q: Can you take over an existing Next.js codebase?** Yes. A takeover starts with an audit: bundle composition per route, where the server/client boundaries actually sit, what each route's caching and rendering behavior really is (not what the comments claim), Core Web Vitals field data, and the upgrade path from whatever version the project froze on. You get a written findings report first, then a fixed-scope plan. No rewrite pitch unless the audit genuinely supports one, and most of the time it does not. **Q: Next.js vs plain React: do we actually need Next.js?** Not always, and we will tell you when you do not. An internal tool or a dashboard behind a login has no SEO surface and no need for server rendering; a Vite SPA is simpler to build and simpler to operate. We ship both: SimpliRFP, published on this site, is a React SPA because that was the right shape for it. Next.js earns its complexity when public pages, search visibility, per-route rendering strategy, or a large content surface are part of the product. If they are not, we will say so in the first call. **Q: Which Next.js version should we be on, and do you handle upgrades?** New builds start on the current major; this site runs Next.js 16. Existing apps should upgrade deliberately, not reflexively, because the majors have changed real semantics: caching defaults moved between versions 13 and 15, so code written against one major's implicit behavior can silently change freshness on the next. We handle upgrades as engineering work: read the release notes against your codebase, make every route's caching policy explicit, then verify behavior route by route before it ships. **Q: How do you keep Core Web Vitals healthy on a Next.js site?** Discipline, not tricks. Server components stay the default so client bundles carry only what interactivity needs; heavy dependencies stay out of client components; images go through next/image with real dimensions; fonts load through next/font so text never swaps late. Then we measure field data (what real visitors experience, not just lab runs) and hold a per-route budget. Most Next.js performance problems we see are bundle problems wearing a framework costume. --- ## Case Studies ### MaiProperty: Custom AI Landscaping Case Study Starting in March 2024, BearPlex built a segmentation and ComfyUI landscaping pipeline with marketplace product matching, quotations, contractor leads and lender pre-approval connections. Before GPT Image and Nano Banana. URL: https://www.bearplex.com/case-studies/maiproperty --- ### PeoplePlus Case Study How BearPlex built PeoplePlus, a comprehensive HR SaaS platform with 15+ modules, 130+ database tables, and 143K lines of code. The OS for services companies. URL: https://www.bearplex.com/case-studies/peopleplus ### Client Testimonial > "BearPlex built PeoplePlus from day zero into a 15+ module HR platform that now runs production operations across the MENA region. 130+ database tables, 143,000 lines of code, and a platform stable enough that we ship new modules without fear." > > Hamad Pervaiz, Founder, PeoplePlus at PeoplePlus --- ### Vertex360: NDIS Management Platform Case Study How BearPlex built a comprehensive NDIS ecosystem from ground zero: participant management, rostering, invoicing, HRM, and compliance reporting for Australian disability providers. URL: https://www.bearplex.com/case-studies/vertex360 ### Client Testimonial > "BearPlex built our entire NDIS ecosystem from ground zero: participant management, rostering, invoicing, HRM, and compliance reporting in a single integrated platform. The speed and quality of execution were unlike anything we'd seen from previous engineering partners." > > NDIS Provider Leadership, Chief Operating Officer at Vertex360 --- ### Extended Trust: Vehicle Protection Platform Case Study How BearPlex built a full-stack vehicle protection platform for the Canadian F&I industry: dealer management, VIN decoding, claims processing, and 11 protection products. URL: https://www.bearplex.com/case-studies/extended-trust --- ### Yaay365: Membership & Loyalty Platform Case Study How BearPlex built a full-stack membership platform with 6 integrated apps: member website, partner portal, admin dashboard, mobile app, e-wallet, and QR system for African markets. URL: https://www.bearplex.com/case-studies/yaay365 --- ### AWS Infrastructure Dashboard Case Study How BearPlex built a unified AWS visibility dashboard for Odus Cloud, monitoring 17 servers, 226+ sites, and security alerts without direct console access. URL: https://www.bearplex.com/case-studies/aws-dashboard ### Client Testimonial > "Monitoring 17 servers and 226+ sites without direct AWS console access sounds impossible. BearPlex built the visibility layer we needed (unified security alerts, resource monitoring, and cost attribution) while maintaining the strict access boundaries our audit requirements demanded." > > Odus Cloud Team, Infrastructure Leadership at Odus Cloud --- ### CleverCoach: Tutoring Platform Case Study How BearPlex built a full-stack tutoring management platform with smart teacher-student matching, digital contracts, and dual-sided financial management. URL: https://www.bearplex.com/case-studies/clevercoach ### Client Testimonial > "German education market expects Swiss-watch-level reliability. BearPlex delivered 33K+ lines of production code, smart teacher-student matching that actually works, and a digital contract flow that holds up in German courts. Three years later, the platform is still the backbone of our business." > > CleverCoach Leadership, Founder at CleverCoach --- ### Scale Mediation: Mediator CRM Case Study How BearPlex built a full-stack mediation management platform with 6 microservices, 114K+ lines of code, AI-powered case assessment, and Stripe Connect payments for US mediators. URL: https://www.bearplex.com/case-studies/scalemediation ### Client Testimonial > "Six microservices, 114K+ lines of code, AI-powered case assessment, Stripe Connect payments for US mediators. BearPlex understood the regulatory nuance of legal mediation, not just the engineering, which is why this platform works in practice, not just in demos." > > Scale Mediation Leadership, Founder & CEO at Scale Mediation --- ### Letti AI: Career Intelligence Platform Case Study How BearPlex built an AI-powered career intelligence platform with automated job scraping, 3072-dimensional skill embeddings, and personalized interview preparation. URL: https://www.bearplex.com/case-studies/letti-ai ### Client Testimonial > "We needed AI-powered career intelligence at scale: automated job scraping, 3072-dimensional skill embeddings, personalized interview prep. BearPlex shipped the full stack in a single War Room engagement. The semantic matching quality is now a core differentiator in our market." > > Letti AI Leadership, Founder & CEO at Letti AI --- ### Optinizers OS: AI Continuity Agent Case Study How BearPlex built an autonomous AI agent that eliminates knowledge loss in VA agency operations, achieving zero knowledge loss with intelligent continuity. URL: https://www.bearplex.com/case-studies/optinizers ### Client Testimonial > "Zero knowledge loss. That was the outcome BearPlex promised and the outcome BearPlex delivered. Our VA agency operations used to lose context every time a team member rotated. Now the AI continuity agent captures, organizes, and makes retrievable every decision: indefinitely." > > Optinizers Leadership, Founder at Optinizers --- ### Ticketmaster Scraper: Real-Time Ticket Intelligence Case Study How BearPlex built an AI-powered Ticketmaster scraper with real-time seat map parsing, residential proxy rotation, and automated data collection into Supabase. URL: https://www.bearplex.com/case-studies/ticketmaster --- ### SimpliRFP: AI-Powered Government Contract Discovery Case Study How BearPlex built an AI-powered platform that discovers, parses, and summarizes government contracts from the US and Canada, reducing RFP review time by 66%. URL: https://www.bearplex.com/case-studies/simplirfp --- ## Tools ### AI Readiness Score (interactive self-assessment) Free 4-minute self-assessment that scores an organization against the 48 checks of the BearPlex AI Readiness Audit (six dimensions: strategy, data infrastructure, model pipelines and evaluation, team capabilities, governance and risk, change management). All checks weigh equally; the overall 0 to 100 score is the average of the six dimension percentages. Results place the organization in one of four fixed-range readiness bands (Forming 0-39, Developing 40-64, Operational 65-84, Compounding 85-100); the bands are ranges on the framework itself, not comparisons against other companies, because no benchmark dataset sits behind the tool. The score, band, dimension breakdown, and top three gaps are free with no email; a full evidence-backed gap report is available by email. URL: https://www.bearplex.com/tools/ai-readiness-score --- ### Project Cost Estimator (interactive cost-range tool) Free 2-minute estimator for custom software cost. The visitor picks one of six project archetypes (WordPress marketing website, Next.js marketing website, e-commerce storefront, SaaS MVP / custom platform, AI system / agent build, custom internal tool / CRM / ops system), answers a handful of scope questions (size, user roles, integrations, compliance, timeline), and gets an honest cost range anchored to BearPlex's published bands, with the exact drivers that moved it. Ranges only, never a single number; the low end never falls below the published floor for the archetype; the combined scope multiplier is capped just above the typical band, and the tool says so when the cap is hit. No hourly rates anywhere: fixed-scope work is priced by outcome and embedded work by what ships per month. The full result is free with no email; an optional emailed breakdown restates the range against the published band with every scope selection and its effect. URL: https://www.bearplex.com/tools/project-cost-estimator --- ## Reports Research-backed intelligence reports. Each is a deep, source-verified methodology with a downloadable PDF (gated behind an email capture at the linked page). ### The Technical Due Diligence Defence Kit *Category: Due diligence* The 50-point checklist VCs and their technical auditors use to interrogate your tech stack during Series A/B diligence. Architecture, code quality, security, engineering ops, and IP: every audit pillar covered. URL: https://www.bearplex.com/reports/the-technical-due-diligence-defence-kit PDF: https://www.bearplex.com/reports/the-technical-due-diligence-defence-kit.pdf --- ### The AI Readiness Audit *Category: AI strategy* A framework to evaluate your organisation’s preparedness for AI integration: from data infrastructure and model pipelines to team capabilities and governance. URL: https://www.bearplex.com/reports/ai-readiness-audit PDF: https://www.bearplex.com/reports/ai-readiness-audit.pdf --- ### The SaaS Scalability Blueprint *Category: Engineering* Engineering playbook for scaling from 1K to 1M users. Database sharding, caching strategies, queue architecture, auto-scaling patterns, and cost optimisation. URL: https://www.bearplex.com/reports/saas-scalability-blueprint PDF: https://www.bearplex.com/reports/saas-scalability-blueprint.pdf --- ### The Security Posture Assessment *Category: Security* Complete security audit methodology covering OWASP Top 10, infrastructure hardening, secret management, dependency scanning, and incident response readiness. URL: https://www.bearplex.com/reports/security-posture-assessment PDF: https://www.bearplex.com/reports/security-posture-assessment.pdf --- ### Cloud Cost Optimisation Playbook *Category: Infrastructure* Tactical guide to cutting cloud spend by 40% without sacrificing performance. Reserved instances, rightsizing, spot strategies, and FinOps best practices. URL: https://www.bearplex.com/reports/cloud-cost-optimization-playbook PDF: https://www.bearplex.com/reports/cloud-cost-optimization-playbook.pdf --- ### The CTO Succession Framework *Category: Leadership* How to structure engineering leadership for resilience. Bus-factor analysis, knowledge-transfer protocols, documentation standards, and org design for scale. URL: https://www.bearplex.com/reports/cto-succession-framework PDF: https://www.bearplex.com/reports/cto-succession-framework.pdf --- ## Stack Reviews First-person, hands-on reviews of the tools BearPlex builds with, each with a rating out of 5 and a plain verdict. ### LangChain (3.5/5) LangChain is useful again after its v1 refocus, but we still treat it as a high-level agent and integration layer, not the place to hide core product logic. It is strongest when a team wants provider flexibility, standard message/tool abstractions, and a fast path to a working agent loop. For durable, auditable production agents, we still push the state machine into LangGraph or ordinary application code. URL: https://www.bearplex.com/stack/langchain-review --- ### Pinecone (4/5) Pinecone remains the safest managed vector database choice when the team wants retrieval to be someone else's operational problem. Its best fit is production RAG where uptime, scaling, backup posture, and indexing behavior matter more than database portability. The tradeoff is cost and lock-in: once metadata design, namespaces, and retrieval behavior are tuned around Pinecone, moving away is real work. URL: https://www.bearplex.com/stack/pinecone-review --- ### LangGraph (4.5/5) LangGraph is our default choice for production agents that need explicit state, durable execution, streaming, checkpoints, and human review. It is lower-level than LangChain, which is the point: production agent behavior should be inspectable instead of hidden inside a one-call abstraction. The cost is extra architecture work up front, but that cost is cheaper than debugging a runaway agent loop after launch. URL: https://www.bearplex.com/stack/langgraph-review --- ### Claude Agent SDK (4.5/5) Claude Agent SDK is the strongest vendor-specific agent SDK when the job resembles Claude Code: inspect a codebase, run commands, edit files, and work through a task loop. It is not a neutral agent framework; its value comes from exposing Claude Code's agent loop in Python and TypeScript. Use it for developer automation and coding-agent workflows, not for general customer-facing agents where model portability, predictable cost, and strict sandboxing dominate. URL: https://www.bearplex.com/stack/claude-agent-sdk-review --- ### Qdrant (4.5/5) Qdrant is our preferred open-source vector database when filtering, tenant boundaries, and retrieval control matter more than managed convenience. The Rust core, payload filtering model, quantization options, and hybrid query features make it a serious production choice. The tradeoff is ownership: if you self-host it, you own capacity planning, backups, upgrades, and query tuning. URL: https://www.bearplex.com/stack/qdrant-review --- ### Weaviate (4/5) Weaviate is strongest when vector search, keyword search, reranking, and RAG workflow features need to live close together. It is more opinionated than Qdrant and less plug-and-play than Pinecone, but it gives teams a broad search platform rather than just a vector index. We choose it when hybrid retrieval and schema-driven data modeling matter; we avoid it when the product only needs a small vector sidecar. URL: https://www.bearplex.com/stack/weaviate-review --- ### LlamaIndex (4/5) LlamaIndex is still the best specialized framework for document-heavy RAG, ingestion, parsing, retrieval, and context assembly. It is not our default for general agent orchestration, but it is usually the fastest path from messy documents to a usable retrieval layer. The production risk is over-adopting the framework: keep ingestion, retrieval evaluation, and source lineage explicit so the app can evolve beyond the first RAG prototype. URL: https://www.bearplex.com/stack/llamaindex-review --- ### Vercel AI SDK (4.5/5) Vercel AI SDK is the best TypeScript-first toolkit for shipping AI product interfaces: streaming text, tool calls, structured output, provider routing, and React/Next.js chat UX. It should not own your agent business logic or long-running state, but it is excellent at the edge between model output and user experience. We use it when the hard problem is product polish, streaming ergonomics, and provider flexibility in a web app. URL: https://www.bearplex.com/stack/vercel-ai-sdk-review --- ### pgvector (4/5) pgvector is the right answer when vector search is a feature of your Postgres application, not the center of your retrieval business. It keeps embeddings close to relational data, transactions, backups, access control, and existing operational habits. It stops being the right answer when vector search needs independent scaling, deep hybrid search, multi-tenant retrieval tuning, or massive ANN workloads. URL: https://www.bearplex.com/stack/pgvector-review --- ### MLflow (4/5) MLflow has become much more relevant for AI engineering because tracing, evaluation, prompt/version management, and production monitoring now matter as much as classic experiment tracking. It is strongest for teams that need one open platform spanning ML models, LLM apps, and agents. It is heavier than purpose-built LLM observability tools, but that weight can be a strength in enterprises already using Databricks or MLflow governance patterns. URL: https://www.bearplex.com/stack/mlflow-review --- ### Cohere (4.5/5) Cohere is most valuable in production RAG as an enterprise retrieval-quality vendor, especially for embeddings and reranking. We rarely choose it because it has the flashiest chat model; we choose it when semantic relevance, multilingual search, or reranking quality is worth paying for. The risk is treating Rerank like magic: it improves ordering, but it cannot recover documents the retriever never found. URL: https://www.bearplex.com/stack/cohere-review --- ### Mistral (4/5) Mistral is best understood as a serious enterprise AI platform with strong open-model roots, not just a cheap OpenAI alternative. We use it when deployment control, European vendor posture, model choice, or on-prem/private deployment options matter. It is less automatic as a default app model than OpenAI or Anthropic, but it belongs on the shortlist for enterprise, sovereignty, and custom-model work. URL: https://www.bearplex.com/stack/mistral-review --- ### Modal (4.5/5) Modal is one of the best Python-first ways to run AI compute without becoming an infrastructure team. It shines for bursty GPU inference, batch jobs, fine-tuning experiments, sandboxes, and internal ML services where per-second serverless economics beat idle GPU ownership. It is less ideal for always-hot, ultra-low-latency services where dedicated infrastructure or a managed inference provider may be cheaper and more predictable. URL: https://www.bearplex.com/stack/modal-review --- ### Together AI (4/5) Together AI is a strong default for managed open-model inference when teams want fast access to a broad model library, fine-tuning, dedicated endpoints, and GPU clusters without operating the stack themselves. It is especially useful when cost, model optionality, and open-source model access matter. It is not a universal replacement for frontier APIs: quality, latency, and reliability must be evaluated per model and endpoint type. URL: https://www.bearplex.com/stack/together-ai-review --- ### DSPy (3.8/5) DSPy is the strongest framework we have used for turning prompt work into an optimization problem, but it is not a general replacement for LangGraph, LlamaIndex, or direct model APIs. Use it when the quality bottleneck is measurable prompt behavior and you have a real development set, a metric, and time to run optimizer experiments. Skip it when you need a full production orchestration layer, TypeScript-first product plumbing, or a team that is still learning the basics of LLM evaluation. URL: https://www.bearplex.com/stack/dspy-review --- ### Supabase (4.5/5) Supabase is our default backend for product builds where a small team needs Postgres, auth, storage, realtime, and serverless functions on day one, and we trust it enough to run our own SaaS (PeoplePlus) on it. The database underneath is standard Postgres, which means row-level security is a real security boundary and the exit path is a pg_dump away, not a rewrite. The honest caveat is operational: native backups and PITR cover the database only (Storage buckets are excluded), the free tier has no automatic backups at all, and PITR is a paid add-on, so production readiness means an external backup pipeline and a rehearsed restore. We run Supabase as a dedicated practice; the receipts and the full playbook live on our Supabase development service page. URL: https://www.bearplex.com/stack/supabase-review --- ### Next.js (4.5/5) Next.js is our default framework for anything with a public web surface, and the review you are reading is served by it: bearplex.com is a Next.js 16 App Router build with 440+ indexable routes, ISR against a live ATS, and an edge-rendered OG image service. The App Router's server-components model is genuinely the right architecture for content-heavy and programmatic sites, and it is also where teams hurt themselves: the server/client boundary reshapes bundles, and caching defaults have changed across major versions, so upgrades are behavioral work, not version bumps. The honest caveat is that not everything needs it; we still ship plain React SPAs when there is no SEO surface to win. The receipts and the full playbook live on our Next.js development service page. URL: https://www.bearplex.com/stack/nextjs-review --- ## Feed (Full Posts) ### The Makeover Era *Published: 2026.06.11 · 5 min read · Tag: ENGINEERING · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/makeover-era **Excerpt**: Vibe-coded apps are shippable, not survivable. AI-built software is meeting real users, real load, and real audits. Most of it needs a rebuild, not a debug pass. The app works. That is the new part. Built over a weekend with Claude Code or Codex, it signs users up, takes their money, sends the onboarding emails, renders a respectable dashboard. Two years ago that was a quarter of work for a funded team. Today it is one determined person, a long Saturday, and a tab full of prompts. I think this is genuinely great. The tools are not the problem. Agentic coding has collapsed the distance between idea and artifact further than anything since the compiler. People who could never afford custom software now have it. Founders validate in days what used to consume a seed round. I run an engineering studio with 65+ engineers, so I am supposed to feel threatened by all of this. I do not. I feel like a structural engineer watching the whole county discover it can build its own Ferris wheels. Busy years ahead. Because there is a pattern in what this wave produces: the apps are shippable, and they are not survivable. Those are different properties. Only one of them was ever in the prompt. ## What the Prompt Never Asked For Vibe-coding optimizes for the visible. The screen, the flow, the demo. That is exactly why it feels miraculous: everything you can point at works. But most of what makes software survivable is invisible, and agentic tools, left unsupervised, skip it silently. - **Data model discipline.** The schema grows by accretion, one prompt at a time. Each feature adds a column, a JSON blob, a duplicate source of truth. Nothing gets normalized because nobody asked. - **Authorization boundaries.** Authentication is a solved prompt. Authorization is design work. Who can see whose records, and under which role? Vibe-coded apps reliably check that you are logged in and rarely check what you are allowed to touch. - **Observability.** No structured logs, no traces, no alerts. When it breaks at 2am, the only debugging tool left is asking the model what it thinks it wrote. - **Tests.** Either the model wrote them and the model graded them, or the demo was the test suite. Both fail the first time a change actually matters. None of this is the tools' fault. A senior engineer driving Claude Code gets the schema, the policy layer, the traces, and the tests, because they ask for them. The tool answers the questions it is given. The weekend builder does not yet know which questions exist. That gap, between what got generated and what was never specified, is where the serious work of the next two years lives. ## The Audit Arrives For a while, none of it matters. Ten users, low stakes, the app glides. The invisible layers stay invisible right up until something good happens. Then it does. A growth spike. A first enterprise customer with a security questionnaire. A procurement team asking about SOC 2. An investor running technical due diligence. Real users, real load, real auditors: three audiences the weekend app has never met, often arriving in the same quarter. We already know how AI-built software does when reality starts grading it. S&P Global Market Intelligence found that 42% of companies abandoned most of their AI initiatives before production in 2025, up from 17% a year earlier. MIT's NANDA initiative found that about 95% of generative AI pilots show no measurable P&L impact. Those numbers describe enterprise projects with budgets, committees, and steering decks. The weekend app has none of that armor, and it hits the same wall, because the wall was never about resources. Shipping is not the hard part anymore. Surviving what you shipped is. Audits are not kind even to the grown-ups. At Japan IT Week this spring, we ran live security diagnostics on 47 Japanese enterprises from our booth, and not one came back clean. Those were established companies with real engineering organizations. Now extend that curve to software whose entire provenance is a chat transcript. ## Makeover, Not Makeup The instinct, when one of these products starts creaking, is to order a debug pass. Fix the slow queries, patch the auth, sprinkle in some tests. I say no to that scope, and the no is the most useful thing I can offer. A debug pass assumes the structure is sound and the defects are local. In a vibe-coded system the defects are the structure. You cannot patch your way from no authorization model to an authorization model. You design one, then rebuild around it. The same goes for the schema, the observability stack, and the test strategy. Makeup hides the problem until the next incident. A makeover replaces what the speed skipped. The frame has three parts. 1. **Keep the product.** The flows, the validated demand, the decisions users already voted for with their time and money. That work is real. It was bought with weekend speed, and it is the most valuable thing in the repository. 2. **Rebuild the skeleton.** A schema designed on purpose. Authorization as an explicit policy layer instead of a scattering of if-statements. Logs, traces, and alerts wired in before scale demands them. Tests that encode intent, so the next prompt cannot quietly undo it. 3. **Redesign the face.** Vibe-coded apps all wear the same face: the same component defaults, the same gradient hero, the same spacing drift. The surface is its own discipline, and it deserves the same deliberate treatment as the skeleton: a face designed on purpose, ending in a system, not a fresh coat of paint. At BearPlex we run this work the way we run everything: War Rooms, cross-functional pods embedded for 90-day production deployments. The structure is the honesty. Rebuilding under live traffic is a commitment, not a ticket queue. ## The Interesting Work Here is the position that surprises people: I think this rebuild wave is the best engineering work available right now. Greenfield is overrated. Starting from nothing mostly tests your taste in boilerplate. Rebuilding a live product with paying users, where you cannot stop the world and the previous architect was a language model with no memory of its own decisions, tests everything else: archaeology, judgment, sequencing, nerve. There has always been a word for engineers who can replace the skeleton without dropping the body. The word is senior. The tools belong inside this work too. The same model that generated the slop, pointed at a real architecture by an engineer who knows what to ask for, becomes a formidable rebuilding instrument. The difference was never the model. It was the questions. So if you shipped something fast and it is starting to creak, hear this clearly: you did it right. You validated before you invested, which is the correct order. The creaking is the signal that the investment is now due. [Talk to us](/contact): the first conversation is with an engineer, not an account manager, and the diagnostic costs nothing. And if you would rather see how we think before talking to anyone, the six research reports in our [library](/reports) are public, every statistic in them re-verified against primary sources. The weekend app was the opening act. The Makeover Era is the show, and we intend to be extremely busy. --- ### The Agent Test *Published: 2026.06.11 · 5 min read · Tag: STRATEGY · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/the-agent-test **Excerpt**: One question separates agent projects worth building from the ones Gartner expects to be cancelled. Most teams cannot answer it. Here is the test we use. Gartner went looking for real agentic AI vendors and found about 130 of them. Not 130 categories. Not 130 percent growth. About 130 actual companies, out of the thousands currently selling something with the word agent on it. The rest are doing what Gartner calls agent washing. The label changed. The product did not. The same Gartner research expects over 40 percent of agentic AI projects to be cancelled by the end of 2027. Two in five, dead before they ship anything that matters. I run an AI engineering studio, and I think agents are the most interesting engineering problem of this decade. So read what follows as the opposite of a rant. Those Gartner numbers do not offend me. They describe the projects we decline. There is one question that separates the agent projects worth a quarter of your roadmap from the ones headed for Gartner's cancellation column. I ask it before any agent work starts, and I would want every CTO to ask it before signing anything. ## The Test **What does this agent do that a workflow plus an LLM call does not?** That is the whole test. It has to be answered concretely, in a sentence or two, by the person proposing the build. Not with adjectives. With behavior. Be precise about what the alternative is, because it is stronger than most people give it credit for. A workflow plus an LLM call means deterministic steps, with a model invoked at the nodes where language or judgment is needed. Extraction here. Classification there. A drafted reply at the end, gated by a human before anything irreversible happens. Retries in code. Fallbacks in code. Logs you can actually read. You can build it in weeks, test it like normal software, and explain to a regulator exactly what it will and will not do. Most of what gets pitched as an agent today is exactly this, with a loop drawn around the diagram. The model is consulted at fixed points. The path through the work never changes. Nothing decides anything that was not already decided at design time. If the autonomy in your agent can be expressed as an if statement, it is not autonomy. It is an if statement. And that is fine. A workflow is not a consolation prize. It ships faster, fails more legibly, and costs far less to run. The mistake is not building one. The mistake is building one inside an agent framework, with agent complexity, agent failure modes, and an agent invoice. ## What Earns the Word The test is a filter, not a wall. Some answers pass it, and when they do, an agent is the only honest architecture. An agent earns the word when the path through the task cannot be enumerated in advance: - The next step depends on what the system just discovered, and the branching is too rich to draw. Think of investigating a production incident, or reconciling records across systems that disagree in ways nobody catalogued. - The system has to choose which tool to reach for, not just which parameters to pass to a fixed one. - It has to recover from failures nobody predicted, not just retry the ones somebody did. - The environment pushes back. Each action changes what the right next action is. If you can draw the flowchart, build the flowchart. Concrete passing answers sound like this: the task tree is different for every input and only visible one step at a time. Or: the plan made at step one is routinely invalidated by what step three finds, and the system has to keep going anyway. Failing answers sound like this: it will be more flexible. It will learn over time. The demo was incredible. One more marker worth writing down. A workflow that handles the common cases and routes the rest to a human will usually beat a fully autonomous system that handles everything unevenly. If a human checkpoint would not embarrass your business case, you did not need an agent. You needed software. ## Why Everyone Says Agent If the test is this simple, why does anyone fail it? Because every incentive in the room points the other way. Budgets flow toward the word. Boards ask about it by name. Vendors know the same product commands different numbers priced as a workflow tool versus an agent platform. Internal teams know an agent initiative gets headcount and a keynote slide, while an automation cleanup gets neither. Nobody in that chain is lying, exactly. They are rounding up. Thousands of vendors rounded up, and Gartner counted about 130 that did not need to. The demand side is sprinting just as hard. Deloitte found 74 percent of organizations expect at least moderate AI agent use by 2027, while only 21 percent have a mature governance model for agentic AI. When appetite runs that far ahead of discipline, the market will happily sell you appetite. The test also has a sibling, and it deserves a sentence here: what does building this give us that buying it does not? MIT NANDA found externally purchased AI tools reached deployment about 67 percent of the time, against roughly 33 percent for internal builds. If someone already sells a tool that passes the agent test for your problem, buying it is not a defeat. Pride is not an architecture. ## The Free Question We wrote this exact check into [the AI Readiness Audit](/reports/ai-readiness-audit), one of the 48 checks the audit puts in front of engineering leaders, because it is the cheapest item on the list with the most expensive failure mode behind it. Like everything in [the reports library](/reports), every statistic in the audit was re-verified against primary sources in June 2026. The numbers above are the ones that pushed me to write this essay. The test disciplines us as much as anyone. BearPlex deploys through 90-day War Rooms: cross-functional pods embedded with the client, shipping to production inside the window. Ninety days does not forgive a wrong architecture chosen in week one. If we let a workflow problem wear an agent costume, we are the ones standing in the wreckage on day 60. The filter protects the client's quarter and ours. So we apply it early. Discovery at BearPlex is not billed, and the first conversation is with an engineer, not an account manager. That engineer will ask you the question in this essay, and no is a real possible outcome. I am proud of how often we say it. For some companies, a clear no on a doomed agent build is the most valuable thing any vendor will give them this year. If there is a concrete answer, we will build the agent, and we will build it properly, with evaluation and guardrails designed in from day one. If there is not, you just saved a budget, a quarter, and one very uncomfortable board meeting. The question is free. The quarter is not. --- ### Evals Are the Product *Published: 2026.06.11 · 6 min read · Tag: ENGINEERING · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/evals-are-the-product **Excerpt**: Stalled LLM products all share one missing artifact: a way to know if the thing got better or worse. The eval suite is the spec, the regression net, and the roadmap. Voiceflow migrated between two versions of the same model. Not GPT-3.5 to GPT-4. GPT-3.5 to GPT-3.5, one snapshot to the next. Intent classification dropped 10 percent. They caught it fast, and that is the remarkable part. The Applied LLMs practitioners who documented the incident point out it was caught only because the evals existed. Same prompt, same pipeline, same model family, and a tenth of the system's accuracy quietly gone. Without an evaluation suite, that regression ships. It sits in production for weeks. Users feel it before anyone inside the building does, and when someone finally notices, the team blames the prompt. I think about that story every time a founder tells me their LLM feature is almost there. Almost there usually means someone is on prompt version forty and the demo behaves when the CEO is watching. The bottleneck in these products is rarely the prompt and rarely the model. The bottleneck is that nobody can answer the only question that matters: did the last change make this better or worse? ## The Treadmill Prompt tweaking feels like progress because the feedback is instant. You adjust a sentence, rerun your three favorite test inputs, and the output reads better. Small dopamine hit. Commit. Ship. What you cannot see is everything else that moved. A language model is a distribution over behaviors, and a prompt edit shifts the entire distribution. You fixed the tone on your three test cases and broke date extraction on cases you never look at. Nobody checked, because nothing forces the check. Without a fixed yardstick, every edit is a coin flip where you only ever see one side of the coin. Hamel Husain, whose field work on this problem is the sharpest I have read, finds that unsuccessful AI products almost always share a single root cause: the failure to create robust evaluation systems. Not weak models. Not clumsy prompts. A missing measurement layer. The stalled teams all look the same. They polish outputs by eye, they argue about whether responses feel better, and they accumulate a folder of prompts with names like final-v9-ACTUAL-this-one. Motion, mistaken for progress. The stall has a body count. S&P Global found that 42 percent of companies abandoned most of their AI initiatives before production in 2025, up from 17 percent a year earlier. Plenty of those deaths were deserved. Some were products that worked, attached to teams who could not prove it. When you cannot demonstrate improvement, you cannot defend the next sprint, and the pilot dies in a budget review. ## The Anatomy A working evaluation system has three levels. Hamel Husain's framework lays them out, and the order matters because each level earns the next. **Assertions.** Deterministic unit tests that run on every change: the output parses as JSON, the required fields exist, the date is a real date, the response never mentions a competitor, the refund amount never exceeds the invoice. They are cheap to run and brutal in what they catch. GoDaddy's engineering team measured malformed output on roughly 1 percent of its structured GPT-3.5 calls. One percent sounds tolerable until you multiply by production volume and realize it is a steady stream of broken responses every single day. Assertions catch those in code and trigger retries before a user ever sees them. **Trace review.** Humans reading real production transcripts, on a calendar cadence, with tooling so frictionless there is no excuse to skip it. This is where you learn what is going wrong instead of what you guessed might go wrong, and it is where your failure taxonomy comes from. Once volume outgrows human reading, you add LLM judges, with one hard rule attached: Husain's guidance is to track the correlation between judge scores and human labels before trusting the judge. An unvalidated judge does not measure your failures. It launders them into a green dashboard while users quietly churn. **A/B tests.** The final word. When the assertions pass and the traces look clean, real users tell you whether the improvement is real. Everything before this level exists to make sure you only spend live traffic on candidates that deserve it. There is a fourth piece that sits beside the levels rather than inside them: deterministic gates on anything irreversible. GoDaddy gates human handoff with code-identified stop phrases instead of the model's own judgment, because a model that is right most of the time still fires the wrong action constantly at scale. Refunds, deletions, escalations, anything with consequences gets decided by code. We gave this whole discipline its own pillar in the [AI Readiness Audit](/reports/ai-readiness-audit) we published, because the public failure data points at evaluation more than at any other layer of the stack. ## The Afternoon Upgrade Models now churn faster than budget cycles. Every few months something better and cheaper lands, and every team faces the same question: migrate, or stay put and watch the cost gap widen? Without evals, migration is a rewrite. Someone re-tests the system by hand, the team argues about whether the new outputs seem fine, somebody senior gets nervous, and the migration slides a quarter. Teams stay on aging models for a year, paying more for worse output, purely because nobody can say what would break. With an eval suite, migration is an afternoon. Pin the current model version. Point the suite at the candidate. Read the diff: where it wins, where it regresses, which assertions moved. Decide by evening, with a rollback path and a number attached to the decision. Voiceflow's 10 percent drop was never a crisis, because the suite turned a silent regression into a line item. The same suite covers you when a provider goes down. GoDaddy lost its chatbots for hours to a provider outage and came away treating outages, malformed outputs, and version drift as routine engineering problems: retries, fallbacks, multi-provider switching, all handled in code. But a fallback provider is only safe to fail over to if you know its quality holds, and the only way to know in advance is a suite you can run against it before the bad day arrives. This is the asymmetry most teams miss. The prompt is disposable, tuned to one model's quirks. The model is rented, and the terms change without notice. The eval suite is the only layer you own, and it compounds: every production failure becomes a new test case, and every test case makes the next migration cheaper. It outlives every model it ever judged. ## The Spec You Never Wrote Here is the reframe I want to leave you with. Most LLM products have no written spec, because nobody can fully specify in advance what "answer customer emails well" means. But the eval suite specifies it, executably. Each assertion is a requirement someone decided to enforce: the assistant never invents a policy, never quotes a price outside the table, always cites a source. The trace-review taxonomy, ranked by frequency, is your roadmap, ordered by what real users hit rather than by what the loudest voice in planning guessed. The judge rubric is your quality bar, finally written down where it can be argued with. That is why I say the evals are the product. The model generates. The evals define what good means, and the definition is the part your customers experience. The same artifacts feed the heavier machinery later. Graded outputs and human preference labels are the raw material of [RLHF and alignment tuning](/services/rlhf-alignment), so when the day comes to tune behavior into a model rather than prompt around it, a team with mature evals already owns the dataset that work requires. A team without one starts from zero, twice. So start, and start ugly. Twenty assertions and one hour a week reading traces beats the elegant evaluation platform you are planning to build next quarter. Wire the suite into CI so no prompt change lands without a score. Promote every production failure into a test case the same day it happens. In six months you will hold the most valuable artifact in your stack, the one thing that gets stronger every time a model gets replaced. The teams that win the next model generation will not be the ones with the cleverest prompts. They will be the ones who can tell, by the end of an afternoon, whether the new thing is better. Build the yardstick first. Then argue about prompts. --- ### Scraping Is an Arms Race You Win With Discipline *Published: 2026.06.11 · 6 min read · Tag: ENGINEERING · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/scraping-arms-race **Excerpt**: Clever bypasses win a week. Disciplined extraction wins years. On rate ceilings, backoff, observability, and the boring engineering that survives the arms race. Every scraper dies the same way. Not with an error. With a page that returns 200, looks perfectly healthy, and contains nothing you wanted. The block usually landed weeks earlier. The team just was not watching. Extraction at scale has a reputation as a dark art: rotating proxies, headless browsers, cat-and-mouse tricks traded in forum threads. The reputation is earned, but it points at the wrong skill. The hard part of scraping has never been getting one page. The hard part is getting page ten million, on schedule, six months from now, after the target has redesigned twice and swapped anti-bot vendors once. That is an operations problem. And operations problems are won with discipline. ## The Lucky Run Most extraction projects are engineered for the lucky run. Someone finds a bypass, the crawl works, the demo lands, and the team scales it 50x overnight, because nothing inflates faster than confidence in a working demo. Two weeks later the data stops, and nobody can say exactly when, because nobody was measuring. I think of it as a simple split. **The lucky run** is when your scraper works because nobody has noticed it yet. **The long run** is when your scraper works because it was built to be tolerable. Those are different engineering goals, and they produce different systems. The arms-race framing matters here. On the other side of your scraper sits an anti-bot vendor with a funded roadmap, a detection team, and telemetry from thousands of sites. Your one clever trick is competing against their quarterly release cycle. You will not out-clever a roadmap. You can out-discipline one, because discipline does not expire when detection improves. A system built on respect for rate ceilings, graceful degradation, and honest measurement keeps working through the detection upgrades that wipe out the trick-based crawlers around it. ## The Ceiling Every site has a tolerance: a request rate at which you are background noise, and a request rate at which you are a problem worth solving. Disciplined extraction starts by finding that ceiling and then staying politely under it, permanently, even when you could go faster. This sounds slow. It is the opposite. Backoff is the clearest example. Most teams implement backoff as an apology, something the scraper does when it gets caught. We treat it as a feature with a spec. Exponential delays. Jitter, so a thousand workers do not retry in lockstep. Circuit breakers that pull an entire domain out of rotation the moment error rates twitch. A scraper that knows how to slow down is a scraper that gets to keep going. A scraper that only knows how to push gets remembered, fingerprinted, and banned at the network level, taking your future options with it. Fingerprint realism is the same idea from another angle. Amateurs fake a user agent string and call it stealth. Real detection looks at everything at once: TLS handshake, header order, timing rhythm, session depth, whether your supposed browser ever loads a stylesheet. You do not beat that by lying harder. You beat it by actually behaving like the thing you claim to be, which mostly means behaving like someone patient. No human reads 400 product pages a second. Not even my engineers, and I hired 65 of the most caffeinated ones I could find. And when a block comes anyway, the discipline is to read it instead of fighting it. A 429 is the site telling you exactly where its line is. A new challenge page is advance notice that the defense posture changed. Teams that treat blocks as obstacles rotate IPs and stomp harder, burning trust they cannot rebuild. Teams that treat blocks as signals adjust the same day and are still collecting a year later. The block is free intelligence. Most people throw it away. ## The Instruments The most dangerous scraper failure makes no noise at all. The site redesigns, your selector matches an empty div, and your pipeline keeps writing rows. Every status code says 200. Every dashboard stays green. Three weeks later an analyst asks why prices stopped changing, and you discover you have been collecting beautifully structured nothing since the 14th. This is why I insist on observability on every fetch, with no exceptions for the boring ones. Each request should leave a record: status, latency, response size, parse yield. Each extracted record should pass through quality gates before it is allowed to exist: - Field fill rates, tracked over time and alarmed on dips - Value distributions, so a price column that flatlines to zero screams - Schema checks that catch a redesign within minutes, not weeks - Canary pages with known content, fetched on a schedule and compared against truth The principle underneath: treat data quality like uptime. An HTTP 500 is the good kind of failure, loud and pageable. A 200 full of garbage is the kind that quietly poisons every decision built downstream. If your monitoring only watches transport and never watches meaning, you are watching the half of the system that was never the risk. ## The Line Discipline also means knowing when not to collect, and this is where I will plant a flag: the strongest scraping teams are defined by what they refuse. Public data only. If it sits behind a login, a paywall, or a personal privacy setting, the answer is no. Read the terms of the sites you touch and treat them as an input to system design, not an obstacle to route around. Strip personal data you do not need at the point of ingestion, because every field you collect is a field you now have to secure, monitor, and answer for. Data minimization usually gets framed as compliance. I frame it as engineering hygiene: smaller surface, fewer liabilities, cheaper pipeline. These lines are not decoration on top of the engineering. They are part of why the engineering survives. A collector that respects the explicit and implicit rules of its targets attracts less hostility, draws fewer escalations, and gives the lawyers on both sides nothing to do. The teams that ignore the lines win sprints and lose the race. When a project requires crossing them, the right answer is to walk away. We say no often, and I am proud of it. Refusal is a capability. It compounds. ## The Long Run None of this is theoretical for us. The [Letti AI engagement](/case-studies/letti-ai) was large-scale data extraction, the kind of work where the lucky run is worthless because the system has to keep producing long after launch week. Work at that scale is why I hold these convictions as hard as I do. Not exotic bypasses. The boring parts, treated as the product: rate discipline, degradation paths, instruments on every fetch, quality gates on every record. And extraction is only the front door. The data that comes through it still has to be validated, deduplicated, versioned, and delivered somewhere it can earn its keep, which is why we treat scraping as the first stage of [a data pipeline](/services/data-pipelines) rather than a standalone stunt. A perfect crawl feeding a sloppy pipeline produces the same outcome as a blocked crawl, just with a bigger storage bill. Here is the test I give any extraction system, ours included. Imagine the target hires its dream anti-bot team tomorrow. Does your system degrade gracefully, alert honestly, and keep delivering what it still can while you adapt? Or does it go dark and take your roadmap with it? The arms race is real, and it never ends. That is exactly why it rewards the patient. Cleverness peaks on day one. Discipline compounds every day after. And the page that returns 200 with nothing inside? With the right instruments, it pages you in four minutes instead of poisoning you for three weeks. That difference is the whole game. --- ### Why Our Discovery Is Free *Published: 2026.06.11 · 6 min read · Tag: BUSINESS · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/why-discovery-is-free **Excerpt**: We never bill for the part where we decide whether we can help. Charging for diagnosis rewards finding work. Free diagnosis rewards finding the truth. There is an invoice we have never sent. It would cover the first conversation: the hours where a BearPlex engineer sits with your problem, asks uncomfortable questions, reads whatever you are willing to share under NDA, and works out whether we can actually help. Sometimes that takes an afternoon. Sometimes it takes days of real engineering attention from people whose time is the most expensive thing we have. We have never billed for it. Not as a discount, not as a loss leader with a nurture sequence waiting behind it. We decided early that the part where we figure out whether we can help is not work we do for you. It is work we do on ourselves, and you should not pay for it. I want to explain the reasoning, because it gets misread as generosity. It is not generosity. It is incentive design. ## The Meter Picture the alternative, because most of the industry runs on it. You bring a problem to a consultancy. They quote a paid assessment: two weeks, a fee, a report at the end. It sounds fair. Expertise costs money, and you are buying expert time. Now watch what the money does to the diagnosis. A paid assessment has to justify its own price. Nobody feels good paying for two weeks of analysis that concludes with "you are fine, you do not need us." So the report grows findings. Risks inflate into roadmaps. A problem that needed a config change and a stern code review comes back as a transformation initiative with phases. Nobody lied. The meter was running, and a running meter wants to find billable road ahead. I have read a lot of these assessments over the years. The recommendations section almost always describes work the authoring firm happens to sell. That is not a coincidence. That is the invoice writing the diagnosis. Charging for diagnosis creates an incentive to find work. We wanted the opposite pressure: an incentive to find the truth, even when the truth pays us nothing. ## The Wrong Yes Here is what the unbilled diagnostic actually protects: us, from our own optimism. When BearPlex says yes, it is not a small yes. Our integrated teams carry a three-month minimum. The War Room model embeds a cross-functional pod with a client for a 90-day production deployment. A yes commits real engineers, from a bench of 65 and growing, to your problem for a quarter of a year. A wrong yes at that scale is brutal. Engineers grinding on a problem that should never have been accepted. A client paying for motion instead of progress. A quarter of capacity burned, a reference we will never earn. The cost of a bad engagement lands on both sides, and it dwarfs anything we could have charged for the diagnostic that should have caught it. So the free diagnostic is not charity. It is the cheapest insurance we buy. If an engineer spends three days inside your problem and concludes we are the wrong firm, those three days just saved both of us a catastrophic quarter. I take that trade every time. If someone in our finance team keeps a spreadsheet of what this policy costs us, I have made the executive decision never to open it. ## The Speed of No The economics only hold under one condition: you have to be willing to say no, and say it quickly. A free diagnostic that drifts for months while everyone stays polite is just an unpaid sales cycle, and unpaid sales cycles eventually bankrupt the firms that run them. What makes ours sustainable is the discipline of the fast no. If the problem sits outside what we are genuinely good at, we say so on the first call. If the data your idea depends on does not exist, we say that before anyone drafts a proposal. If the honest answer is that you do not need an AI engineering studio, you need one strong hire and six months of patience, we say that too. I am proud of how often we say it. Saying no fast does two things. It caps the cost of the giveaway, which is what lets us keep giving it away. And it trains judgment. Every quick no is a rep for the muscle that evaluates problems honestly, and that muscle is the only thing that makes a yes from us worth anything. A firm that says yes to everything has a yes that carries no information. It is just the sound a vendor makes when a budget appears. When we do say yes, it means engineers looked at the problem with no meter running, with every incentive pointed at truth, and concluded the work is real and we are the right people for it. That is the entire value of the word. We protect it by spending the no freely. ## An Engineer Picks Up One structural detail makes all of this work: the first conversation at BearPlex is with an engineer, not an account manager. That was an engineering-culture decision, not a staffing quirk. An account manager's job, structurally, is to keep the conversation alive. Every metric they carry points toward advancing the deal, and you cannot blame a person for doing the job their incentives describe. Put one at the front door and the diagnostic becomes theater. The real evaluation happens later, after rapport, after momentum, after it is socially expensive for anyone to walk away. An engineer's job is to end the conversation with an answer. Can we help, or not. What would it actually take. Which failure mode has nobody mentioned yet. Engineers are constitutionally bad at pretending a problem is interesting when it is not, which makes them poor salespeople and exactly the right people for diagnosis. The diagnostic is an engineering artifact, so it gets held to engineering standards: it has to be correct, not persuasive. So when you write to us through [the contact page](/contact), the person who engages with your problem could be the person who has to build the solution. They have skin in the diagnosis. If they say yes and they are wrong, the consequence is their next quarter, not a missed commission. ## What the Yes Is Worth I am aware that an essay called "Why Our Discovery Is Free" on a company blog invites a cynical reading. Fair. Here is the honest accounting. We charge properly for the work itself, and [how we think about pricing](/pricing) is public. The work is where money should change hands. The decision about whether the work should exist at all is where it should not, because the moment that decision earns revenue, the decision bends. Free discovery is not a promotion. It is a constraint we imposed on ourselves to keep our own judgment clean. It forces the no to come quickly, because slow nos are expensive. It puts engineers at the front door, because they are the only people who can answer the question actually being asked. And it means that when we finally do send an invoice, every line on it survived a diagnosis that had no reason to flatter you and every reason to be right. The first conversation costs you nothing. It costs us an engineer's attention, and often enough, the deal. That is exactly the price that keeps it honest. --- ### The Boring Levers *Published: 2026.06.11 · 6 min read · Tag: ENGINEERING · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/the-boring-levers **Excerpt**: Most scaling failures are self-inflicted: Google-shaped architecture adopted before the boring levers are exhausted. Boring first is sequencing, not timidity. Eight percent. That is the average CPU utilization across production Kubernetes clusters, measured by Cast AI across tens of thousands of them. Not at 4am on a holiday weekend. On average. Datadog's State of Cloud Costs research prices the same picture from another angle: 83% of container spend is tied to idle resources. Read those two numbers again, because they describe a choice. Thousands of engineering teams built distributed systems for traffic that never arrived, and now they pay monthly rent on the waiting room. The clusters did not fail to scale. They scaled exactly as designed. They scaled the idle. I have seen this enough times to call it what it is. Most scaling failures are self-inflicted. The database does not break because you grew. It breaks because the team skipped the boring work and went straight to the exciting kind. ## Borrowed Problems The pattern repeats with remarkable discipline. A SaaS platform with four thousand users adopts microservices, an event bus, a service mesh, and a multi-region Kubernetes footprint, because that is what serious companies run. Except the companies being copied adopted those things under duress, at a scale where nothing simpler survived. Copy the solution without the duress and you get all of the operational cost with none of the necessity. Your platform with four thousand users does not have Google's problems. It has Google's YAML. Premature architecture is a loan. You borrow complexity today and pay interest on it every on-call shift: more deployment paths, more failure modes, more dashboards, more 2am pages that can only explain themselves through distributed traces. The collateral is your roadmap, because every hour spent operating the big system is an hour not spent building the product that was supposed to need it. The cruel part is that the bill still arrives at the database. Microservices do not rescue a Postgres instance with missing composite indexes and a cache hit rate nobody measures. They just make it harder to see which service is hammering it. ## The Late Sharders The two most instructive database scaling stories in SaaS belong to Notion and Figma, and people routinely take the wrong lesson from both. Notion runs one of the most famous sharded Postgres fleets in the industry. Read their engineering team's own account of the migration, though, and notice what forced it. Not slow queries. VACUUM stalls and transaction ID wraparound risk, a Postgres failure mode that does not degrade gracefully: it halts every write on the system. When Notion moved, they moved well. Sharded by workspace ID, with 480 logical shards chosen precisely because 480 divides cleanly in ways a power of two never will. Figma is even clearer. Their databases team grew the database stack almost 100x since 2020 before horizontally sharding at all. The runway came from unglamorous levers: caching, read replicas, and roughly a dozen vertically partitioned databases. When sharding finally became unavoidable, they shipped it in-house in about nine months. Here is the part people miss. Notion's stated lesson from the whole experience was to shard earlier. That sounds like a vote against everything I just argued, until you notice what it actually means: once the symptoms have names, stop hesitating. It was not a regret about the sequence. The sequence was right. The regret was the pause at the end of it, after the system had already asked. Both companies sharded late. Figma sharded once; Notion came back for a second pass, tripling its fleet from 32 to 96 instances, and by the team's own account that re-shard cost users at worst about a second of a saving spinner, because dark reads compared old and new databases before cutover. Because they waited, both sharded with years of real data about access patterns, tenancy boundaries, and hot spots. Shard on day one and you are guessing at your shard key. The shard key is the one decision in this whole discipline that is brutally expensive to take back. ## The Levers, In Order What does boring first actually look like? Roughly this, in roughly this order: - **Indexes and query plans.** Read your slow query log before you read another Kubernetes doc. Most "we need to scale" conversations end here, quietly. - **Caching.** Cache-aside with sensible TTLs covers most of it. It is the cheapest 10x in the entire playbook, and the easiest to get burned by if you never test cache-cold. - **Read replicas.** Most SaaS workloads are read-heavy. Cloning data is dramatically cheaper than splitting it. - **Vertical partitioning.** Move the noisiest tables onto their own hardware. Alongside caching and read replicas, this is what bought Figma its years. - **Queues.** Put a buffer between bursty producers and anything that falls over at peak, then scale the consumers, not the queue. - **Rightsizing.** Datadog finds most container workloads use less than a quarter of the CPU they request. Reclaim what you already pay for before buying more. None of this is timid. Data Center Dynamics reported that Stack Overflow serves around 6,000 requests per second from nine on-prem web servers running a tuned monolith. Nine. Know what one well-run box can do before you architect for a fleet. Every lever on that list shares the properties the exciting ones lack. It is reversible. It is cheap. Its failure modes are documented to death. An index takes an afternoon. A cache policy takes a week. A microservices migration takes a year, and you learn whether it worked at the end. It is also the conversation I want to have first whenever someone brings us [enterprise platform work](/services/enterprise-platforms). Not redesign. Measurement. Which queries, which tenants, which queue depths, which hit rates. The glamorous architecture conversation can wait until the numbers demand it. ## Sequencing, Not Timidity The obvious objection: if the big architecture will be needed eventually, why not build it now and skip the rewrite? Because "eventually" is doing fraudulent work in that sentence. Most platforms never reach the scale where the boring levers run out. The ones that do arrive there with revenue, real load data, and a team that understands its own system, which is exactly the position you want to occupy when making expensive, irreversible decisions. Figma sharded from strength. Notion sharded with a clear forcing function and a measured plan. Neither was rescued by an architecture built years earlier on guesses. Boring first does not mean boring forever. It means writing graduation criteria in advance and respecting them. When VACUUM falls behind, when replica lag becomes user-visible, when one tenant's traffic distorts everyone's p99, the system is asking. Answer it then, decisively, the way Notion wished they had. Until that day, every exciting component you do not run is an outage you do not have and a line item you do not pay. We wrote the full sequence down. [The SaaS Scalability Blueprint](/reports/saas-scalability-blueprint) is 49 checks across six disciplines, built from what Notion, Figma, Shopify, Slack, and Amazon's own builders actually did, with every statistic re-verified against primary sources in June 2026. Run it as an audit. Each check you fail has a named company that already paid for the lesson, which means you do not have to. And if you would rather walk through it with someone, our first conversation is with an engineer, not an account manager, and the diagnostic costs nothing. Bring your slow query log. The teams that scale are not the ones that built for a million users on day one. They are the ones that kept the system boring long enough to earn the interesting problems. --- ### Sovereignty Is a Feature *Published: 2026.06.11 · 5 min read · Tag: STRATEGY · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/sovereignty-is-a-feature **Excerpt**: Where data is allowed to live is becoming a product requirement. What building for Health Canada, Saudi Aramco, and NDIS Australia taught us about jurisdiction. The most expensive line on an enterprise architecture diagram is no longer the one around the database. It is the national border drawn around everything else. Health Canada. Saudi Aramco. NDIS Australia. Those three names sit on our [enterprise platform register](/services/enterprise-platforms), and on paper they share nothing: a federal health authority, an energy giant, a national disability insurance scheme. In the room, they ask the same first question. Not what the platform does. Not what it costs. Where the data lives. Which country, which region, which legal regime, and who exactly can reach it from outside. I think that question belongs in the product requirements, on the same line as uptime. Where data is allowed to live is becoming a feature. Buyers in healthcare, government, and energy already treat it as one. Everyone else is about to. ## The First Question Compliance-heavy work has a reputation for being slow, and some of that reputation is earned. But it teaches a discipline most startups learn too late, usually in the middle of their first enterprise deal, when a procurement team asks for a data flow map and the honest answer is "mostly us-east-1, we think." Our healthcare-sector compliance work made this concrete for us early. Health data does not move freely. It arrives with rules about where it can be stored, where it can be processed, who can administer the systems holding it, and under whose laws a subpoena lands. Government work sharpens the same point. Energy work sharpens it again, with national infrastructure stakes attached. The lesson that transfers to every other sector is simple: jurisdiction is an architectural input, not an output. You either decide where data lives on day one, or you discover where it ended up on the day someone official asks. Your data has a passport now. It is just never allowed to travel. ## Storage Is the Easy Half When teams hear data residency, they think storage. Put the database in Montreal, done. Storage is the easy half. The hard half is processing. A request that originates in Toronto can touch a queue in Virginia, an error tracker in California, an email relay in Dublin, and a model inference endpoint in a region nobody checked, all before the response renders. Each hop is a jurisdiction decision somebody made by default, usually by accepting the first region in a dropdown. Backups replicate across borders because that was the durable option. Logs ship to a third-party tool because that was the convenient one. Telemetry alone can move more sensitive data than the application does. So the first artifact on any sovereign engagement is unglamorous: the architecture diagram with the residency boundary drawn as a hard line, and an inventory of everything that crosses it. Every queue, every backup target, every observability vendor, every support-access path. Once the crossings are visible, the design conversation gets simple. Each one is eliminated, moved inside the boundary, or documented with a legal basis a reviewer will accept. Third parties deserve their own line in that inventory, because the numbers are blunt. The Verizon 2025 Data Breach Investigations Report found that third-party involvement in breaches doubled in a single year, to 30 percent. Every vendor inside your residency boundary inherits your promises and adds its own failure modes. Our published [Security Posture Assessment](/reports/security-posture-assessment) treats vendor exposure as a first-order risk for exactly this reason: your data's jurisdiction is only as stable as the sloppiest SaaS tool in your stack. ## The Retrofit Is a Rewrite Here is the part founders consistently underestimate. Residency is not a layer you add. It is a property smeared across every layer you already have. Retrofitting starts with re-homing the database, which is the easy part, and continues with re-homing everything the database talks to. The queue. The cache. The analytics pipeline copying production data into a warehouse on another continent. The five SaaS tools that each hold a partial replica. The CI system with production credentials. The admin panel your support team opens from a different country. Multi-region was already hard as a performance problem. As a legal problem, it comes with a reviewer checking your work and a contract clause waiting on the result. And while you rebuild, the deal that triggered the rebuild waits. The squeeze is predictable: the buyer wants a residency commitment in the contract, the current architecture cannot honor it, and the engineering roadmap turns into a hostage negotiation between sales and infrastructure. Teams that architected for jurisdiction early skip the whole drama. For them, the audit is paperwork. For everyone else, it is a rewrite with a deadline attached. ## Pick Your Geometry Sovereignty has more than one shape, and the right one depends on what the regulator and the buyer actually require. Our [sovereign cloud](/services/sovereign-cloud) work falls into three families: 1. **Pinned region.** Data and processing stay in a named in-country cloud region. The control plane (deployment tooling, identity, orchestration) can stay global. This satisfies most residency requirements at the lowest operational cost, and it is where most teams should start. 2. **Dedicated sovereign environment.** A separated environment where operator access is restricted by contract and by architecture: in-country admin paths, in-country key management, in-country support staff. For buyers who care about who can touch the system, not only where the disks sit. 3. **On-premise or air-gapped.** The data never leaves the building, let alone the country. Slowest to operate and strongest to defend, and for certain government and energy workloads it is the only acceptable answer. The pattern underneath all three is the same. Separate the data plane from the control plane, pin the data plane to the jurisdiction, and treat every boundary crossing as a contract. Get that separation right at the start and moving between models later is mostly configuration. Get it wrong and you are back in the rewrite chapter. ## Paperwork, Not Panic We carry the same paperwork we help clients prepare. BearPlex is in the middle of its own SOC 2 Type II audit, and every engagement starts NDA-first. That is what lets a studio of 65-plus engineers headquartered in Lahore, with presence in Austin and Doha, deliver work across 19 countries for clients whose regulators do not accept "trust us" as an architecture document. The deeper point is about timing. Every company that grows past a certain size will face the residency question. The compliance-heavy sectors just face it first, which is why they are worth learning from. They front-load the decisions: where data sits, where it is processed, who holds the keys, what crosses the border and why. Made early, those decisions cost a whiteboard session. Made late, they cost a quarter of engineering and a delayed contract. Boring audits are the goal. An auditor who reads your data flow map, matches it against your infrastructure, and leaves without follow-up questions is the highest compliment an architecture can receive. You earn it years before the audit happens. If your buyers have started asking where the data lives, the conversation costs nothing. Discovery at BearPlex is not billed, and the first person you talk to is an engineer, not an account manager. Draw the border on the diagram first. Everything else is engineering. --- ### Two Departures From Stalling *Published: 2026.06.11 · 6 min read · Tag: METHODOLOGY · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/two-departures-from-stalling **Excerpt**: Most software organizations are two departures away from stalling, and almost none have measured it. Bus factor is a number you can compute and engineer down. Name the engineer whose resignation email would ruin your quarter. You had a name before the sentence ended. Most leaders do. What almost nobody has is the second name, the third, or a number that says how deep the bench really goes. The first answer is instinct. The rest is risk management, and most software organizations have never done it. Here is the claim I want to defend: most software organizations are at most two departures away from stalling, and they have never measured the distance. Not because the risk is hidden. Because nobody counted. ## A Number, Not a Feeling Engineers call this the bus factor: the smallest number of people who, if they vanished tomorrow, would leave a system nobody can confidently change. The name says something about the profession's optimism. In our imagination, nobody ever takes a better offer or burns out quietly. They have to get hit by a bus. The nickname makes it sound like trivia, a morbid icebreaker for retros. Wrong category. It is a number, it has a peer-reviewed estimation algorithm, and it has been computable from your git history for a decade. Avelino and his co-authors ran exactly that computation across 133 popular GitHub systems and found that **65 percent had a truck factor of two or less**. Two departures from stalling, in well-known projects with global contributor pools and every line of code in public view. I see no reason to believe the average private codebase, written under deadline and reviewed by nobody outside the building, does better. And git history flatters you. Commit analysis sees who wrote the code. It does not see who reviews everything, who unblocks every stuck thread, who carries the memory of why the billing system has that one terrifying feature flag. Some of your most dangerous dependencies have thin commit logs and full calendars. If your audit only reads the repository, your real number is worse than the one on the screen. So measure properly, then run the cheapest audit in the industry: send each key person genuinely offline for two or three weeks and keep a written log of everything that stalls and everyone who pings them anyway. That gap log is your succession plan in draft form, written by reality instead of by HR. ## The Clock Is Already Running The standard objection arrives on cue: our people are happy, nobody is leaving. Maybe. But tenure has a curve, and it is shorter than your roadmap. Nash Squared surveyed just over two thousand technology leaders across 62 countries and found they expect to stay in their current role for **3.3 years** on average. Not the disgruntled ones. The average ones. That number converts every key-person dependency from a hypothetical into a dated liability. If your architecture lives in one head, the expected case, not the unlucky one, is that the head is somewhere else within three years. You do not get to schedule the departure. You only get to schedule the preparation. The cost of getting this wrong is documented at the very top of the org chart. Harvard Business Review put the market value wiped out by badly managed CEO and C-suite transitions at close to a trillion dollars a year in the S&P 1500 alone. Boards already treat CEO succession as standing scenario planning. Almost nobody runs the same exercise on the staff engineer who is the only person able to deploy the thing the company actually sells. The title is smaller. The dependency is not. ## Nobody Is Coming There is a comforting story we tell about departures: someone steps up. The repository is still there, the docs are mostly fine, a capable replacement absorbs it all in a quarter. Open source gives us a way to test that story at scale, and Nourry and colleagues did, across more than 36,000 projects. Of the projects that lost all their core developers, **only 27 percent ever attracted a replacement**. The rest just stopped. A company can do something an open source project cannot: pay someone to show up. But you are hiring a body, not the map. The new hire receives the code. They do not receive the decade of rejected alternatives, the constraints nobody wrote down, or the list of things that look refactorable and are actually load-bearing. The standard mitigation fails too. The notice-period knowledge transfer, two weeks of walkthrough meetings while the departing engineer's heart is already at the next job, is theater. Nobody retains a passive tour of a decision engine. Knowledge moves when the receiver does hands-on work with the expert reviewing rather than driving, and that takes calendar time you do not have once the resignation lands. Transfer works in peacetime. The time to start was before anyone resigned, and the second-best time is this quarter. ## Replaceable on Purpose I run an AI engineering studio with more than 65 engineers, and the question I ask myself most often is uncomfortable in exactly the right way: what stalls if I disappear? If the honest answer is "a lot," then I have not built a company. I have built a bottleneck with a personal brand. So I treat my own replaceability as a design requirement, the same way I treat uptime. Hand off one real responsibility every quarter, permanently and visibly. Grow two candidates for every critical role, because a single anointed successor is a bet, and bets miss. Promote people for making others capable, and stop rewarding indispensability. An irreplaceable engineer is a failure of the system around them, and most organizations quietly pay people to stay irreplaceable. The same conviction shapes how we work with clients. Our [integrated teams](/services/integrated-teams) carry a three-month minimum, and this essay is the reason: code ships in weeks, but judgment transfers slowly, through working together on real systems. An outside team that leaves and takes the understanding with it has not lowered your bus factor. It has rebranded it. We refuse to become the dependency we were brought in to remove. ## Engineer It Down Bus factor responds to engineering the way latency does: measure, set a target, work the levers, re-measure. We compressed everything we know about those levers into [The CTO Succession Framework](/reports/cto-succession-framework), 48 checks across measurement, knowledge transfer, documentation, deputy development, org design, and board governance, with every statistic re-verified against primary sources. If you only take five things from it, take these: 1. Run a truck-factor analysis on every production repository, quarterly, and govern the trend rather than the snapshot. 2. Cross-check the scores against review load and meeting load, because the glue people never show up in commits. 3. Run the vacation test and treat the gap log as a work queue, not an anecdote. 4. Replace walkthrough meetings with hands-on transfer, verified by the receiver operating the system with the expert out of the room. 5. Develop deputies with real, permanent handoffs, and start years before any expected departure. None of this is glamorous. That is rather the point. The companies that survive departures are not the ones with heroic engineers. They are the ones where heroism became unnecessary, one handoff at a time. Two departures. For most teams, that is the whole distance between a roadmap and a standstill. You can find your real number this quarter with a script, an honest afternoon, and the willingness to hear the answer. Or you can find it the way most companies do: during a four-week notice period, with years of context to transfer and nothing but goodwill to transfer it with. I know which postmortem I would rather not write. --- ### The Bill Nobody Owns *Published: 2026.06.11 · 6 min read · Tag: BUSINESS · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/the-bill-nobody-owns **Excerpt**: Cloud waste is rising again because nobody owns the bill. Savings get taken once and lost quarterly. The fix is a name on every line item, not another dashboard. Somewhere in your cloud account, right now, a load balancer is pointing at nothing. The service it fronted was decommissioned months ago. The load balancer stayed. It bills every hour, serves no traffic, and will keep billing until someone deletes it. Nobody will, because deleting it is nobody's job. That load balancer is not a technical problem. Finding it takes minutes. Deleting it takes one. It survives anyway, because it has no owner. Multiply it by every orphaned volume, every oversized instance, every dev environment running all weekend for nobody, and you get the strangest expense in modern business: a bill everyone can see and nobody can touch. ## The Regeneration Cloud waste does not behave like a leak. It behaves like a lawn. You cut it to the ground in a heroic March sprint, and by September it is back, slightly thicker than before. The numbers say this plainly. Flexera's 2026 State of the Cloud report puts self-reported cloud waste at 29 percent of spend, the first increase in five years. Hold that against the decade behind it: more FinOps tooling, more cost dashboards, more visibility than any infrastructure team has ever had. The waste number went up anyway. I find that one fact more useful than any tool comparison. If visibility fixed waste, waste would be fixed by now. Most cloud accounts of any size already have a cost explorer, a tagging policy, and a dashboard someone built in a burst of enthusiasm two years ago. What they do not have is a person whose week gets worse when the number moves. Here is the mechanism. The cleanup is an event. The spending is a process. Events lose to processes every time. The team that rightsizes in March is competing against every deploy, every new service, and every "we'll tune it later" that ships after they move on. Savings are taken once. Waste regenerates quarterly. ## The Known Levers What makes the 29 percent genuinely embarrassing is that none of this is mysterious. The technical playbook has been stable for years. Six levers: commitments, rightsizing, spot, storage tiering, egress, and the operating practice that holds the gains. All six are public knowledge with mature tooling behind them. Now look at the execution data. ProsperOps analyzed roughly $3 billion of AWS compute spend and found a median effective savings rate of 15 percent, against published commitment discounts of up to 72 percent. These are the easiest savings in the catalog: no engineering work, no migration risk. The median organization captures about a fifth of what the vendor is openly offering. Worse, ProsperOps documented an environment at 100 percent reserved instance coverage and 100 percent utilization that still paid about 16 percent more than running everything on-demand. The two metrics most teams put on the slide can both read perfect while the company loses money. Cast AI measured average CPU utilization across Kubernetes clusters at more than 2,100 organizations: 10 percent, down from 13 percent the year before. Teams reserve roughly ten times the compute they use, and the trend is moving the wrong way. Datadog found that 98 percent of organizations incur cross-AZ data transfer charges, and that those charges account for nearly half of all transfer costs. That is not a pricing trap. That is architecture nobody re-examined after it shipped. None of these are knowledge gaps. Engineers overprovision because an availability incident is a career event and waste is not. That is rational behavior under the incentives they were handed. Nobody has ever been promoted for deleting a snapshot. ## The Missing Name So the dashboards exist, the levers are known, and the waste grows anyway. The missing piece is ownership. A dashboard reports. An owner decides. The difference matters because every one of the six levers requires a decision that crosses a department boundary. Finance can see the bill but cannot change the architecture. Engineering can change the architecture but rarely sees the bill. Cloud waste lives in that gap: the one cost in the company that is simultaneously everyone's fault and no one's responsibility. The data supports the org chart reading. CloudZero's State of Cloud Cost research found only one in four organizations allocate 100 percent of their cloud spend to owning teams, and unallocated spend is exactly where waste hides, because cutting a shared bucket is nobody's job. The same research found that where engineering owns cloud costs, 81 percent of organizations say spend is about where it should be. Same clouds. Same tools. Different org chart. Here is the test I use. Pick any line item on the bill and ask who would notice if it doubled next month. If the answer is a person, you have an owner. If the answer is "the dashboard would catch it," you have a smoke detector in a house where nobody lives. This is why I keep saying the bill needs names, not teams. A team is a place for a number to hide. A name is a person who flinches when the number moves. Every line item that matters should have one, with the authority to change the architecture behind it, not just the duty to report on it. ## Keeping It Taken Ownership sounds abstract until you write down what an owner actually does. Four things, mostly. 1. **They review weekly, not monthly.** A monthly cadence gives every leak up to thirty days of free runway. Weekly caps the blast radius of any misconfiguration or runaway autoscaler at seven days. 2. **They report unit economics, not totals.** Cost per customer, per transaction, per request. Rising spend with falling unit cost is growth, and only the unit number can prove it. Only an owner bothers to compute it. 3. **They set caps, not budgets.** Budgets are alarms, not brakes. Milkie Way's engineering postmortem describes burning $72,000 in a few hours against a $7 budget when a recursive Cloud Run job fanned out and billing data ran a day behind. The alert arrived after the damage, as alerts do. 4. **They sequence the levers.** Rightsize first, commit second, because a three-year reservation bought against an oversized fleet locks the waste in for the full term. An owner thinks in sequences. A dashboard thinks in snapshots. We compressed all six levers into 50 concrete checks in our [Cloud Cost Optimisation Playbook](/reports/cloud-cost-optimization-playbook), with every statistic re-verified against its primary source in June 2026. It covers the commitment math, the spot failure modes, the storage tiering fine print, and the egress traps. But I will tell you now how to read it: run it as an audit, and put two columns next to the 50 rows. Pass or fail, and a name. The second column decides whether the savings survive the quarter. My prediction for any team that does this honestly: the technical fixes will take weeks, and most are not hard. The naming will take one uncomfortable meeting. That meeting is the actual work. Until it happens, every optimization sprint is borrowing against the next one, and [the playbook](/reports/cloud-cost-optimization-playbook) is just a longer to-do list nobody is assigned. The load balancer pointing at nothing gets deleted in the first week the bill has an owner. Not because anyone built a better dashboard. Because someone finally owned the delete button. --- ### Survive the Audit *Published: 2026.06.11 · 6 min read · Tag: METHODOLOGY · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/survive-the-audit **Excerpt**: A VC technical auditor is not grading your code. They are pricing your risk. The founders who survive produce artefacts, not answers, on every front. Day two of technical diligence. The auditor asks when you last restored a backup from cold storage. Not whether you take backups. When you last restored one, how long it took, and where the log lives. You have a backup policy. It is a tidy PDF, approved last year, filed in Notion. The auditor does not want the PDF. They want the timestamped restore log, the measured wall-clock time, and the name of the engineer who ran the drill. If those do not exist, your backup policy is a hope with a cron job. This is the moment most founders misread. They think they are being examined on engineering. They are not. The person across the table is not evaluating your code. They are pricing your risk. Every claim you cannot verify becomes a number somewhere downstream: a valuation discount, an escrow holdback, a tighter condition in the term sheet. The audit is not an exam you pass. It is a negotiation you can lose one unverifiable answer at a time. ## The Pricing Exercise A technical diligence report sorts everything it finds into two buckets: verified and asserted. Verified findings feed the investment model. Asserted findings feed the risk premium. That distinction explains behaviour that otherwise looks strange. Why does the auditor seem uninterested in your cleverest subsystem? Because cleverness is not on the risk register. Why do they keep asking for logs, dashboards, and signed documents instead of letting your CTO explain the design? Because an explanation is testimony, and testimony is only as good as the person giving it, and that person might not be in the building a year after the deal closes. Two companies with near-identical systems can walk out of diligence with very different terms because one founder answered questions while the other produced evidence. The engineering did not differ much. The provability did. An artefact is anything that exists independently of the person describing it: a log with a timestamp, a dashboard with history, a signed agreement, a postmortem with a date and a follow-up commit. Answers describe intent. Artefacts prove behaviour. Auditors price behaviour. ## Artefacts, Not Answers Here is the rule I hold founders to before they walk into that room: if you cannot produce the artefact, treat your answer as no. Not "mostly yes." Not "yes, in spirit." No. That is exactly how the auditor will score it, and you would rather find out six weeks before the data room opens than inside it. The rule sounds brutal until you run your own claims through it. Then it sounds clarifying. - **"Our system scales."** The artefact is a load test that found the ceiling: the bottleneck named, the profiling evidence attached. A system that has never been pushed to failure has an unknown ceiling, and unknown ceilings get priced as low ones. - **"Our architecture is documented."** The artefact is a current diagram that provably matches infrastructure-as-code and runtime topology. A diagram from two years ago is a historical record, not documentation. - **"We made deliberate technical choices."** The artefact is a set of Architecture Decision Records for the irreversible ones: data model, tenancy, authentication and authorisation. Architecture vibes do not transfer to a new owner. ADRs do. - **"The team knows the system."** The artefact is a measured bus factor and proof you raised it. An org chart shows reporting lines; it does not show how many people can disappear before you are incapacitated. Avelino and colleagues studied 133 popular GitHub systems and found 65 percent had a truck factor of two or less. Assume you are not the exception until you have measured it. Every line of questioning works this way. The question is never "do you do X." The question is "show me the thing that could only exist if you did X." ## What They Pull On The questions sound varied in the room, but they cluster into a handful of territories. Here is where the pulling actually happens in each one. **Architecture and scalability.** The restore drill, first and always. An RPO and RTO you can state and then demonstrate. Failover that has been tested, not designed. SLOs with recent p95 and p99 numbers behind them, captured during peak traffic rather than a quiet Tuesday. **Code quality and tech debt.** Nobody reads your code line by line; there is no time. Auditors look at proxies that resist last-minute polish: flake rate, the modules where high churn meets low coverage, and whether technical debt is quantified in a form an investor can price, with remediation costs and visible paydown. They also hunt the hero component, the critical module only one engineer can modify safely. They tend to find that engineer by watching where the room looks when a hard question lands. **Security and compliance.** The first pull is your git history, not your firewall. Automated secret detection across every commit you have ever made, because GitGuardian counted 28.65 million new hardcoded secrets pushed to public GitHub in 2025, up 34 percent in a year, and the auditor's working assumption is that some of yours are in there too. After that: a threat model that is maintained rather than produced once for a workshop, and incident postmortems that demonstrably changed controls. A postmortem that changed nothing is just a record that something went wrong. **Engineering operations.** DORA-style delivery metrics with a two-quarter trend, not a screenshot from your best week. A CI pipeline that is a hard gate rather than a suggestion engineers bypass "temporarily." The ability to tie any production state back to a specific commit. And backup restore automation again, because one artefact that settles two separate lines of questioning is the cheapest evidence you will ever produce. **IP and bus factor.** The quiet territory, and the one that kills deals outright rather than discounting them. Signed IP assignment from every founder, employee, and contractor who ever touched the code. Proof that the company, not somebody's personal account, owns the repos, the domains, the cloud accounts, and the signing keys. Whether a new senior engineer can stand up a working dev environment from documented steps in under a day. And the thirty-day question: if your CTO resigned tomorrow, what would actually break, and what mitigation exists today? ## Run It on Yourself The defence is not complicated. It is early. Almost every artefact above can be produced in the months before diligence, and almost none of them can be produced during it. A restore drill you run today costs an afternoon. One you needed to have run last quarter cannot be run at all. So audit yourself before someone with a term sheet does. Work through [The Technical Due Diligence Defence Kit](/reports/the-technical-due-diligence-defence-kit) the way an auditor would score you: artefact or no, nothing in between. The first pass is usually uncomfortable. Good. Every no you convert before the data room opens is risk you repriced in your own favour, on your own schedule, without a counterparty watching. I hold BearPlex to the standard I am preaching. We are running our own SOC 2 Type II audit right now, and our process is NDA-first, because a studio that scrutinises other people's systems should expect the same scrutiny back. If you want an engineer's eyes on your posture before the real thing, our first conversation is with an engineer, not an account manager, and discovery is not billed. The kind of work that discipline produces is on display in our [case studies](/case-studies). And if you do exactly one thing after closing this tab, schedule the restore drill. This week. The log it generates will say more for you in that room than anything you could say yourself. --- ### 0 Out of 47: What Happened When BearPlex Scanned Japan's Enterprise Security Live on Stage *Published: 2026.04.14 · 7 min read · Tag: COMPANY · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/bearplex-japan-it-week **Excerpt**: At Japan IT Week Spring 2026, BearPlex ran live security diagnostics on 47 Japanese enterprises. None had zero vulnerabilities. Six audit contracts followed. The number on the screen read **0 / 47**. Zero. Out of forty-seven Japanese enterprises scanned live on the exhibition floor at Japan IT Week Spring 2026, not a single one came back clean. Every company had exploitable vulnerabilities. The screen refreshed in real time. The crowd around our booth grew. Nobody was smiling. That was the point. ![BearPlex team at Japan IT Week: Sheroz Pervaiz, Hamad Pervaiz, and Anita at the BearPlex booth in Tokyo Big Sight](/images/media/japan-it-week-booth.jpg) ## The Spectacle We didn't come to Tokyo to hand out brochures. We came to demonstrate something that no slide deck or capability presentation could communicate: that Japan's rapid digital transformation has outpaced its security posture, and the gap is not theoretical. It is measurable, live, in front of an audience. The booth at Tokyo Big Sight carried a sign in Japanese: **先端技術・高度 AI エンジニアリング支援**. Advanced Technology. Advanced AI Engineering Support. Beneath it, a screen cycling through real-time diagnostic results. Forty-seven scans. Forty-seven failures. The booth didn't need a pitch. The numbers did the talking. Six formal security audit engagements were signed before we left Tokyo. Several are now published as case studies, with full Japanese-language reports delivered to each client. ## The Invitation BearPlex didn't apply for a booth. The Japan International Cooperation Agency (JICA) invited us. JICA operates across 150+ countries and identifies high-capability firms for strategic introduction into the Japanese market. Their invitation was institutional recognition: BearPlex's track record in enterprise AI and application security had preceded us. In Japan, credibility is not claimed. It is conferred. JICA's backing meant that when Japanese CTOs approached our booth, the question wasn't "who are you?": it was "what can you do for us?" ## The Calculation Security was a deliberate opening move. AI engineering is our core: autonomous agent frameworks, RAG knowledge systems, enterprise platforms deployed through 90-day War Room sprints. But AI is abstract. Security is visceral. A vulnerability scan that returns red is a problem a CTO can feel in their chest. Once the diagnostic created urgency, the conversation shifted naturally. The same enterprises alarmed by their security gaps were the ones with approved budgets for AI initiatives and no qualified teams to execute them. Japan's Ministry of Economy, Trade and Industry projects a shortfall of 790,000 IT professionals by 2030. The appetite for expert engineering partnerships isn't a market trend: it's a structural reality. We positioned security as the proof of competence. AI engineering as the scope of ambition. The sequence was intentional. ## The Room Every meaningful conversation at our booth lasted over thirty minutes. This is not a market where you close on a handshake. Japanese enterprises evaluate partners on philosophy, on structure, on how you handle failure: before they evaluate your technical capabilities. The technology is table stakes. What they're measuring is whether you'll still be there in five years. This suited us. BearPlex has never competed on speed of sale. We compete on depth of commitment. The War Room model (cross-functional teams embedded for 90-day production deployments, compensation tied to outcomes) translates well to a culture that values long-term orientation over short-term wins. The conversations that mattered most weren't about what we could build. They were about what we refused to build. In a market saturated with vendors who say yes to everything, the ability to say no (to turn down scope that doesn't serve the outcome) was the signal that separated us. ## The Presence The team on the ground was small by design: myself, Sheroz Pervaiz, and Anita, our Japan business development lead operating full-time from Tokyo. Three people, not thirty. The booth was a statement, not a spectacle of headcount. The engineering force (65+ in Lahore) doesn't need to be in the room. It needs to be in the work. Anita is now building our permanent Japan operation. Partnerships with Japanese system integrators are in progress. Security audit reports and case studies are being localized into Japanese. Engagement models are being designed for the Japanese procurement process: longer discovery, deeper relationship building, same outcome-based pricing. The pipeline out of Japan IT Week spans gaming, manufacturing, financial services, and logistics. Several engagements are already in active scoping. ## The Signal There's a reason we chose Japan as our first major expansion into Asia. This is not an emerging market experimenting with AI. This is a mature economy with the world's third-largest GDP, confronting a generational technology transition with a structural talent shortage. The enterprises here don't need education about what AI can do. They need partners who can execute at the standard they demand. The booth sign read **AI エンジニアリングスタジオ**. AI Engineering Studio. Tokyo didn't become our newest market. It became our proving ground. And the first number we showed them was zero. --- ### BearPlex Is Quietly Building the Most Dangerous AI Consultancy in the World *Published: 2025.01.14 · 8 min read · Tag: COMPANY · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/dangerous-ai-consultancy **Excerpt**: An in-depth look at how a stealth consultancy from Pakistan is reshaping enterprise AI deployment with a radical no-prototypes philosophy and war room methodology. In the basement of a nondescript office building in Lahore, Pakistan, a team of 65 engineers is doing something that terrifies their competitors: they're actually shipping. While most AI consultancies are still pitching PowerPoint decks and proof-of-concepts, BearPlex has deployed 47 production AI systems for Fortune 500 clients in the past 18 months. Their secret? They don't do prototypes. "Prototypes are where good ideas go to die," says Hamad Pervaiz, the 28-year-old CEO who founded BearPlex in 2022. "Every prototype is a system that was built to be thrown away. We only build systems meant to run forever." ## The War Room Model The company's methodology is radical by consulting standards. Instead of traditional project phases (discovery, design, development, deployment), BearPlex operates in what they call "War Rooms": 90-day intensive engagements where a cross-functional team embeds with the client and ships production code from day one. "Day one is deployment day," explains Pervaiz. "Not day one of thinking about deployment. Day one of actual users using actual software. Everything else is iteration." This approach requires a different caliber of engineer. BearPlex's hiring process reportedly takes 12 weeks and involves live client work. Only 3% of applicants make it through. ## Outcome-Based Everything Perhaps most unusually for a consulting firm, BearPlex ties their compensation directly to client outcomes. Not just deliverables: actual business metrics. "We have skin in the game," says Pervaiz. "If the system we build doesn't move the numbers the client cares about, we don't get paid. It's that simple." This model has attracted attention from private equity firms and Fortune 100 companies looking to de-risk AI investments. When the consultant only wins if you win, alignment problems disappear. ## The Competitive Threat Traditional consulting firms are watching nervously. BearPlex's combination of low-cost engineering talent, outcome-based pricing, and aggressive timelines makes them formidable competition. "They're doing in 90 days what takes us 18 months," admits one partner at a Big Four firm who requested anonymity. "And they're guaranteeing results. We can't match that model without completely restructuring how we operate." For enterprise clients drowning in AI pilots that never graduate to production, BearPlex offers something refreshing: systems that actually work, built by people who only get paid if they do. The AI consulting industry may never be the same. --- ### Shipping Autonomous Agent Framework v3.0 - Now With Multi-Model Orchestration *Published: 2025.01.08 · 8 min read · Tag: ENGINEERING · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/autonomous-agent-framework-3 **Excerpt**: Autonomous Agent Framework 3.0 introduces multi-model orchestration: different AI models collaborating on complex tasks like a coordinated engineering team. Today we're releasing version 3.0 of the BearPlex Autonomous Agent Framework, and it's the most significant update we've shipped since we first open-sourced the project internally in 2023. The headline feature: multi-model orchestration that allows GPT-4, Claude, Gemini, and other models to work together on complex tasks. Let me explain why this matters and how we built it. ## What is the Autonomous Agent Framework? For those unfamiliar, our Autonomous Agent Framework is the infrastructure layer we use to build AI-powered systems that can actually get work done. Not chatbots. Not simple Q&A interfaces. Systems that can reason through multi-step problems, use tools, maintain context across long-running tasks, and recover gracefully from failures. We started building this in early 2023 when we realized that the gap between "impressive demo" and "production-ready system" was massive. Every client engagement revealed the same pattern: the AI could do the task in isolation, but making it reliable, observable, and maintainable required infrastructure that didn't exist. So we built it. The framework handles agent orchestration, state management, tool execution, memory systems, and all the glue code that turns a language model into a dependable worker. Over the past two years, it's powered everything from automated due diligence systems for PE firms to intelligent document processing pipelines handling millions of pages. ## What's New in v3.0 Version 3.0 introduces **multi-model orchestration**: the ability to compose agents that use different underlying models, each contributing their strengths to a shared task. Here's the problem we were solving: no single model is best at everything. Claude excels at nuanced analysis and following complex instructions. GPT-4 has strong general reasoning and broader tool use. Gemini handles multimodal inputs natively. Smaller models like Mistral are fast and cheap for simple classification tasks. Previously, you picked one model and lived with its limitations. Now, you can design agent workflows where: - A fast, cheap model does initial triage and routing - A reasoning-focused model handles complex analysis - A code-specialized model writes and debugs implementations - A multimodal model processes images and documents - A large context model synthesizes everything at the end All coordinated, all sharing context appropriately, all observable through a single unified interface. ## Technical Architecture The multi-model orchestration layer sits between your agent definitions and the underlying model providers. Here's how it works: ### The Conductor Pattern We implement what we call the "Conductor" pattern. A lightweight orchestration layer (itself optionally powered by a model) manages the flow of work between specialized agents. Each agent declares: 1. **Capabilities**: What tasks it can handle 2. **Model affinity**: Which model(s) it prefers 3. **Context requirements**: What information it needs from other agents 4. **Output schema**: What it produces The Conductor examines incoming tasks, decomposes them when necessary, routes subtasks to appropriate agents, and handles the context passing between them. ### Unified Context Layer The trickiest part of multi-model orchestration is context management. Different models have different context windows, different tokenization, and different strengths in utilizing long context. Our solution is the **Unified Context Layer** (UCL). It maintains a semantic representation of the shared context that can be serialized appropriately for each model. When Agent A produces output that Agent B needs, the UCL: 1. Extracts the semantically relevant portions 2. Compresses or expands based on the target model's context window 3. Formats according to the target model's preferred structure 4. Tracks provenance so outputs can be attributed correctly This means a 200K context Claude agent can pass relevant findings to a 32K context GPT-4 agent without manual intervention. ### Fallback and Redundancy Real production systems need to handle model outages, rate limits, and degraded performance. The framework now supports automatic fallback at the agent level. When the primary model hits issues, the agent automatically fails over while maintaining task continuity. ## Real Client Examples Let me share how this is being used in production. ### Due Diligence Acceleration A private equity client uses multi-model orchestration for deal analysis. The workflow: 1. **Gemini** processes scanned documents and extracts text from images 2. **Mistral** classifies documents and routes them to appropriate analysis tracks 3. **Claude** performs deep analysis of financial statements and legal documents 4. **GPT-4** cross-references findings against market data via function calling 5. **Claude** synthesizes everything into a structured due diligence report What previously took analysts 3 weeks now completes in 2 days with higher consistency. ### Intelligent Customer Support An enterprise SaaS client routes support tickets through a multi-model pipeline: 1. **Small local model** (running on their infrastructure for privacy) does initial classification 2. **Claude** handles complex technical questions requiring deep product knowledge 3. **GPT-4** manages multi-turn conversations requiring tool use (checking order status, etc.) 4. **Specialized fine-tuned model** handles domain-specific regulatory questions Resolution time dropped 60%, and escalation to human agents dropped 45%. ## Performance Improvements Beyond the architectural changes, v3.0 includes significant performance work: - **40% reduction in median latency** through better request batching and parallel execution - **60% reduction in token usage** via smarter context compression - **Near-linear scaling** to 100+ concurrent agents (up from ~30 in v2.x) - **Sub-second agent spin-up** through improved warm pooling We also added comprehensive observability: distributed tracing across model calls, token usage attribution, latency breakdowns, and quality metrics tracking. ## What's Coming Next We're already working on v3.1, targeting Q2 2025: **Adaptive Model Selection**: Instead of static model affinity, agents will dynamically select models based on task complexity, current costs, and observed performance. Early experiments show 30% cost reduction with no quality degradation. **Agent Memory Sharing**: Right now, long-term memory is per-agent. We're building infrastructure for agents to contribute to and query shared knowledge bases, enabling learning across the entire system. **Self-Improving Workflows**: Using execution traces and outcome data to automatically identify bottlenecks and suggest workflow optimizations. **Local Model Integration**: First-class support for running agents on local models (Llama, Mixtral) for privacy-sensitive workloads or cost optimization. ## Getting Started If you're a BearPlex client, your team lead can get you access to the updated framework. We're running hands-on workshops throughout January to help teams migrate from v2.x. If you're not yet working with us but building serious AI systems, let's talk. The problems we've solved building this framework (reliability, observability, multi-model coordination) are the same problems every organization faces once they move past prototypes. The gap between "AI demo" and "AI in production" is real. We've spent two years building the bridge. --- ### How We Deployed 47 AI Agents for a Fortune 100 Client in 90 Days *Published: 2024.12.19 · 6 min read · Tag: CASE STUDY · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/47-agents-90-days **Excerpt**: From initial scoping to production deployment, how outcome-based pricing and parallel development accelerate enterprise AI adoption. When a Fortune 100 logistics company approached BearPlex in September, they had a seemingly impossible ask: automate 47 distinct operational workflows using AI agents, all within 90 days. Most consulting firms would have proposed a 12-month phased rollout. BearPlex said yes. ## The Scale of the Challenge The client's operations involved everything from route optimization and inventory forecasting to customer service automation and compliance monitoring. Each workflow required specialized knowledge, integration with legacy systems, and fail-safe mechanisms for when AI confidence dropped below thresholds. "This wasn't 47 chatbots," clarifies Hamad Pervaiz, BearPlex's CEO. "These were 47 autonomous systems that needed to make real decisions with real consequences. A routing agent that makes a bad call costs the client six figures. A compliance agent that misses something could trigger regulatory action." ## The War Room Approach BearPlex deployed a team of 23 engineers who physically relocated to the client's headquarters for the 90-day engagement. They called it a "War Room": a methodology the firm has refined across dozens of similar engagements. The structure is unusual: no project managers, no status meetings, no documentation sprints. Just engineers paired with client domain experts, shipping code daily. "We had 47 agents in production by day 14," says one BearPlex engineer who worked on the project. "Not all fully featured: some were just monitoring and alerting. But in production, handling real data, with real users watching the outputs." ## The Multi-Model Architecture Technically, the deployment showcases BearPlex's proprietary Autonomous Agent Framework, which allows multiple AI models to collaborate on tasks. A single workflow might use: - Gemini for processing shipping documents and images - Claude for complex reasoning about regulatory compliance - GPT-4 for multi-turn conversations with suppliers - Smaller fine-tuned models for classification tasks "No single model is best at everything," Pervaiz explains. "Our framework lets us compose the right model for each subtask, while maintaining unified observability and failure handling." ## The Results By day 90, all 47 agents were live and processing real workloads. The client reported: - 73% reduction in manual processing time - 94% accuracy across all automated decisions - $14M in annualized cost savings - Zero critical failures in the first 30 days post-deployment ## The Pricing Model BearPlex's fee structure was outcome-based: a base amount for deployment, plus a significant bonus tied to the cost savings actually realized. This meant BearPlex was financially incentivized to optimize for performance, not just delivery. "We turned down about 30% of the original scope," admits Pervaiz. "There were workflows where AI wasn't the right answer. Under hourly billing, we'd have built them anyway. Under outcome billing, it made no sense." For enterprise leaders frustrated by AI pilots that never scale, BearPlex's model offers a compelling alternative: guaranteed outcomes, shared risk, and engineers who only win when you do. The 47-agent deployment is now being cited as a template for enterprise AI adoption across the logistics industry. --- ### Why We Stopped Selling Hours and Started Selling Outcomes *Published: 2024.12.03 · 6 min read · Tag: STRATEGY · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/selling-outcomes-not-hours **Excerpt**: Eighteen months ago we stopped billing by the hour. Uncomfortable, risky, and one of our best decisions. What we learned about aligning consulting incentives. Eighteen months ago, we made a decision that felt risky at the time: we stopped selling hours. No more timesheets. No more billing increments. No more conversations where clients ask why a task took 12 hours instead of 8. Instead, we started selling outcomes. Defined deliverables with fixed prices. Success metrics with our compensation tied to hitting them. Risk shared between us and our clients. It was uncomfortable. It required us to get much better at scoping and estimation. It meant sometimes eating costs when we underestimated complexity. It was also one of the best decisions we've made as a company. ## The Problem With Hours Hourly billing is the default model in consulting for a reason: it's simple. You track time, multiply by rate, send invoice. The client knows what they're paying for. The consultant knows they'll get paid for work done. But simplicity isn't the same as alignment. Here's the uncomfortable truth about hourly billing: the consultant's financial incentive is for the project to take longer. Not deliberately padding hours (most consultants are honest), but there's no structural pressure toward efficiency. If a task takes 20 hours, you bill 20 hours. If you find a clever way to do it in 5 hours, you've just reduced your revenue by 75%. Clients feel this tension even when they can't articulate it. That's why they ask about hours. That's why they want detailed time breakdowns. They're trying to answer the question: "Am I getting value, or am I paying for inefficiency?" The question itself reveals the problem. In a well-structured engagement, the client shouldn't need to audit your timesheets to know they're getting value. ## The Outcome-Based Alternative Outcome-based pricing flips the model. Instead of selling time, you sell results. The conversation changes from "We estimate this will take 400 hours at $200/hour" to "We will deliver X, and it will cost $Y." X is defined clearly. Success criteria are explicit. The client knows exactly what they're getting and what they're paying. We know exactly what we need to deliver and what we'll earn. If we find a way to deliver the outcome faster, we keep the margin. If we underestimate and it takes longer, we absorb the cost. The incentives align: we're motivated to be efficient without sacrificing quality, because quality is what defines the outcome. ## How We Structure Engagements Moving to outcome-based pricing required us to completely rethink how we scope and structure work. Here's the model we've evolved to: ### Phase 0: Discovery (Fixed Price) Every engagement starts with a paid discovery phase. Usually 1-2 weeks, fixed price. The deliverable is a detailed specification: what we're building, how we'll know it's successful, what the risks are, what it will cost. This protects both parties. The client gets a thorough analysis before committing to a large project. We get enough information to price accurately. ### Phase 1+: Outcome Milestones The main engagement is structured as a series of outcome milestones. Each milestone has: - **Clear deliverable**: What artifact or capability will exist - **Success criteria**: Objective measures of completion - **Fixed price**: What the milestone costs - **Timeline**: Expected completion window Payment is tied to milestone completion, not time elapsed. If we deliver early, great. If we need to iterate to meet the success criteria, that's on us. ### The Risk Pool Here's where it gets interesting. On larger engagements, we structure a "risk pool": a portion of the total project value (typically 15-25%) that's tied to overall project success, not just milestone completion. If the project achieves its high-level business objectives (the metrics the client actually cares about), we earn the full risk pool. If it falls short, we don't. This means we have skin in the game beyond just delivering features. We care whether the thing we built actually works in the real world. ## Real Examples Let me share two engagements that illustrate how this works in practice. ### Example 1: Document Processing System A legal services firm needed to automate processing of court filings. Under hourly billing, this might have been scoped as "estimated 600-800 hours of development." Instead, we scoped it as outcomes: - **Milestone 1**: System processes standard filing types with 95%+ accuracy ($X) - **Milestone 2**: System handles edge cases, accuracy reaches 99%+ ($Y) - **Milestone 3**: Full integration with existing case management system ($Z) - **Risk pool**: 20% of total, paid if system processes 10,000+ documents in first 90 days post-launch We delivered all milestones, earned the risk pool. Total project cost was actually lower than the hourly estimate would have been, but our effective rate was higher because we were efficient. Win-win. ### Example 2: AI Assistant Platform A client wanted an AI assistant for their customer service team. We structured it as: - **Milestone 1**: Assistant handles top 10 query types with <5% escalation rate ($X) - **Milestone 2**: Coverage expands to top 25 query types, same escalation threshold ($Y) - **Milestone 3**: Assistant handles 80% of all queries with <10% escalation ($Z) - **Risk pool**: Tied to measured reduction in average handle time Milestone 3 was harder than expected. We spent more time than estimated. But we delivered, and the client got exactly what they needed. Under hourly billing, this would have been a difficult conversation about cost overruns. Under outcome pricing, it was simply us doing what we committed to do. ## Aligning Incentives Changes Everything The shift to outcome-based pricing changed more than our invoicing. It changed how we think about projects. We invest more in upfront discovery because accurate scoping directly impacts our economics. We're more thoughtful about architecture decisions because we'll live with the maintenance burden. We push back on scope creep because the scope is what we're being paid to deliver. Most importantly, our conversations with clients are different. We're not defending time logs. We're collaborating on outcomes. When something changes mid-project (and something always changes), the discussion is about adjusting milestones and pricing, not about who's at fault for the extra hours. ## Advice for Firms Considering This Model If you're running a consulting firm and considering outcome-based pricing, here's what we've learned: **Start with strong clients.** Your first outcome-based engagements should be with clients who trust you and have realistic expectations. Don't try to convert a difficult client relationship to this model. **Invest in discovery.** Your ability to price outcomes depends on your ability to understand scope. Paid discovery isn't overhead: it's the foundation of accurate pricing. **Build in buffers, especially early.** Until you've calibrated your estimation, price with healthy margins. You'll get better at scoping over time. **Be explicit about what's included.** Scope creep kills outcome-based projects. Define boundaries clearly, and have a process for handling change requests. **Track your data.** After every project, compare actual effort to estimates. You need this feedback loop to improve. **Accept that you'll lose some.** Some projects will cost more than you estimated. If you never lose money on a project, you're probably overpricing. The goal is to win in aggregate, not on every engagement. ## The Bigger Picture Outcome-based pricing isn't just a billing model. It's a philosophy about how consulting should work. Clients hire consultants because they want results, not activity. The clearer we can be about what results we're delivering (and the more we tie our compensation to those results), the more honest the relationship becomes. We're not perfect at this. We're still learning, still refining the model. But we're never going back to selling hours. Because at the end of the day, nobody cares how many hours something took. They care whether it worked. --- ### The War Room Model: Inside BearPlex's Radical Approach to Enterprise AI *Published: 2024.11.18 · 7 min read · Tag: METHODOLOGY · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/war-room-model **Excerpt**: Forget traditional consulting. We embed cross-functional teams for 90-day sprints, shipping production systems while competitors are still writing proposals. The whiteboard in BearPlex's Lahore headquarters is covered in red and green dots. Red means blocked. Green means shipped. Today, there's a lot more green than red. "Every dot is a system in production," explains Hamad Pervaiz, walking me through the company's war room methodology. "We don't track tasks or tickets. We track things running in the real world." ## A Different Kind of Consulting Traditional consulting follows a predictable pattern: discover, analyze, recommend, implement. Projects stretch over months or years. Success is measured in deliverables, not outcomes. BearPlex inverts this entirely. Their "War Room" model compresses everything into 90-day engagements where teams ship production code from the first week. "Most consultants are optimized for looking busy," says Pervaiz. "Long proposals, detailed analyses, beautiful slide decks. We're optimized for moving numbers. If the metrics don't move, we failed." ## Inside a War Room A typical War Room engagement begins with what BearPlex calls "Day Zero": a single day where the entire team, including client stakeholders, aligns on exactly three things: 1. What metrics define success 2. What's the first thing that ships 3. Who can make decisions "That's it," says Pervaiz. "Everything else is distraction. We've seen six-month consulting engagements fail because nobody agreed on these three things." From Day One, engineers are writing production code. Not prototypes, not proofs of concept: actual systems that will run in the client's infrastructure. ## The No-Prototype Philosophy "Prototypes teach you that something is possible," Pervaiz explains. "Production teaches you whether it actually works. We'd rather learn the hard lessons early, with real data and real users." This approach requires different engineers. BearPlex's hiring process takes 12 weeks and includes working on actual client projects. The washout rate is 97%. "We're not looking for people who can solve LeetCode problems," says a senior engineer. "We're looking for people who can ship under pressure, communicate with executives, and fix their own mistakes at 2am." ## The Economics of Urgency BearPlex's fee structure reinforces the urgency. A significant portion of each engagement is tied to outcome metrics: if the system doesn't deliver measurable results, BearPlex doesn't get paid in full. "It aligns incentives perfectly," says one client, a CTO at a financial services firm. "They're not incentivized to extend the engagement or find new problems to solve. They're incentivized to make our numbers go up." This model isn't for every client. BearPlex is selective about engagements, turning down companies that can't commit to the War Room intensity. "We need a decision-maker in the room every day," Pervaiz says. "Not someone who needs to 'check with leadership.' If you can't move that fast, we're not the right fit." ## The Competitive Moat For traditional consulting firms, replicating BearPlex's model is nearly impossible. It requires: - Engineers willing to relocate for 90-day sprints - Outcome-based compensation structures - Willingness to turn down scope - A culture that celebrates shipping over billing "Most firms can't do outcome-based pricing because they can't estimate accurately enough," Pervaiz observes. "And they can't estimate accurately because they've never had to. When you bill by the hour, accuracy doesn't matter." The War Room model isn't just a methodology: it's a filter. It selects for engineers who ship, clients who decide, and projects that matter. The whiteboard keeps filling with green dots. --- ### Meet the Firm That Charges for Results, Not Resumes *Published: 2024.11.02 · 5 min read · Tag: BUSINESS · By Hamad Pervaiz* URL: https://www.bearplex.com/feed/results-not-resumes **Excerpt**: Our outcome-based model attracts PE firms tired of consulting projects that promise everything and deliver PowerPoints. Private equity firms have a consulting problem. Every acquisition target comes with a stack of consultant reports: market analyses, technology assessments, operational reviews. Most gather dust. The ones that don't gather dust often recommend changes that never get implemented. Enter BearPlex, a Pakistani AI consultancy that's catching attention in PE circles for a simple reason: they only get paid if their work actually delivers results. ## The PE Opportunity "Due diligence reports are CYA documents," says one managing partner at a mid-market PE firm. "They tell you what could go wrong so you can't blame the consultant later. They don't tell you how to make things go right." BearPlex's model is different. Instead of delivering reports, they deliver systems. Instead of billing for hours worked, they tie compensation to metrics achieved. For PE firms looking to accelerate value creation in portfolio companies, this is compelling. Traditional consulting is a cost center. Outcome-based consulting is an investment with measurable returns. ## The Model in Practice Consider a recent engagement with a PE-backed logistics company. A traditional consulting firm might have billed $2M for a six-month AI strategy and implementation roadmap. BearPlex's proposal: a fixed base plus outcome bonuses tied to specific operational metrics: processing time, error rates, cost per transaction. A meaningful share of the total fee sat in the bonus pool, at risk against results. The result: BearPlex deployed 23 AI agents in 90 days, achieved all bonus thresholds, and earned the full bonus pool. The client saved $14M annually. "We paid more than we would have paid the traditional firm," admits the portfolio company's CEO. "But we actually got something. Not a roadmap: a system. Not a recommendation: results." ## The Trust Factor Outcome-based pricing requires trust on both sides. The consultant trusts the client to measure outcomes fairly. The client trusts the consultant to optimize for long-term value, not short-term metrics gaming. BearPlex addresses this through aggressive transparency. Clients get full access to code repositories, system metrics, and engineering decisions. There's no black box. "We're not protecting IP," says Hamad Pervaiz, BearPlex's CEO. "We're building systems that clients will run forever. They need to understand everything we built." ## The Scaling Question The challenge for BearPlex is scale. Outcome-based engagements require deep trust and close collaboration. Can that model work beyond a few dozen clients per year? "We don't want to be Accenture," Pervaiz says. "We want to be the firm that Accenture's best clients call when they need something to actually work." For PE firms, that positioning is attractive. They don't need a consulting firm that can handle everything. They need one that can handle the things that matter. BearPlex seems to be building exactly that. --- ## AI Model Briefs (Full Briefs) ### Qwen 3: Open-Weights LLM *Publisher: Qwen Team (Alibaba) · Paper date: 2025.05.14 · Brief date: 2026.08.09 · 9 min read* *Parameters: 0.6B / 1.7B / 4B / 8B / 14B / 32B dense + 30B-A3B / 235B-A22B MoE · License: Apache 2.0* URL: https://www.bearplex.com/ai/qwen-3 arXiv: https://arxiv.org/abs/2505.09388 Model card: https://huggingface.co/Qwen/Qwen3-235B-A22B GitHub: https://github.com/QwenLM/Qwen3 **Excerpt**: BearPlex's engineering brief on Qwen 3: the Apache 2.0 license story, the 0.6B to 235B size ladder, hybrid thinking modes, and why it wins on-prem evaluations. **Qwen 3's open-weight models use the Apache 2.0 license.** The [official Qwen3 repository](https://github.com/QwenLM/Qwen3) states that all of its open-weight models use Apache 2.0, and the [flagship model card](https://huggingface.co/Qwen/Qwen3-235B-A22B/blob/main/LICENSE) carries the license text. Commercial use, fine-tuning, modification, and redistribution are allowed under the standard Apache conditions. There is no product badge, user threshold, or derivative naming rule. ### Where to verify the official license Do not take this page's word for it, and do not take an aggregator's word for it either. There are two authoritative locations, both one click away: - **The LICENSE file on the model card.** Every Qwen 3 repository on Hugging Face carries the full Apache 2.0 text, for example the [Qwen3-235B-A22B LICENSE](https://huggingface.co/Qwen/Qwen3-235B-A22B/blob/main/LICENSE). This is the file to attach to a legal review. - **The official repository.** [github.com/QwenLM/Qwen3](https://github.com/QwenLM/Qwen3), maintained by the Qwen team at Alibaba, states the licensing for the open-weight family. Check the specific checkpoint you intend to deploy rather than assuming family-wide terms. It costs thirty seconds, and it is the difference between a license review and a license assumption. ### What Apache 2.0 permits, in plain terms - **Commercial use:** permitted, with no revenue, user, or seat threshold. - **Modification and fine-tuning:** permitted, including on private data you never disclose. - **Redistribution:** permitted, in source or binary form. - **Shipping a derivative under your own brand:** permitted. No naming prefix, no attribution badge in your product UI. - **Private use:** permitted, with no obligation to publish anything. - **Patent grant:** included. Contributors grant a patent license covering their contributions, which MIT does not. - **What you must do:** keep the license text and copyright notices with copies you redistribute, and note significant changes you made to the files. - **What you do not get:** warranty or liability cover. Apache 2.0 disclaims both, which is why regulated deployments pair open weights with their own evaluation and monitoring rather than a vendor contract. When a bank, a health system, or a government body evaluates open-weights models for an on-prem deployment, the shortlist process looks nothing like a leaderboard. Legal reviews the license before engineering benchmarks anything. Infrastructure asks what runs on the hardware already racked. Product asks what happens when the pilot needs to scale down to an edge device or up to a flagship. Qwen 3 is the family that keeps clearing all three gates at once, and this brief is about why. ## What it actually is Qwen 3 ([Qwen Team, April 2025 release; technical report May 2025](https://arxiv.org/abs/2505.09388)) is a family of eight open-weights models spanning three orders of magnitude, per the [official announcement](https://qwenlm.github.io/blog/qwen3/): - **Six dense models**: 0.6B, 1.7B, 4B, 8B, 14B, and 32B. The three smallest carry 32K context; the larger ones 128K. - **Two Mixture-of-Experts models**: Qwen3-30B-A3B (30B total, 3B activated) and the flagship **Qwen3-235B-A22B** (235B total, 22B activated; 128 experts with 8 active, per the [model card](https://huggingface.co/Qwen/Qwen3-235B-A22B), 32,768 context natively and 131,072 with YaRN). Training scale roughly doubled over the prior generation: about **36 trillion tokens** versus Qwen2.5's 18 trillion, with multilingual coverage expanding from 29 to **119 languages and dialects**. The signature feature is the **hybrid thinking mode**: one checkpoint that reasons step-by-step in `` blocks when `enable_thinking` is on (or `/think` is issued in-conversation) and answers immediately when it is off. The [technical report](https://arxiv.org/abs/2505.09388) adds a thinking-budget mechanism, so inference-time reasoning depth becomes a dial rather than a model choice. You deploy one model and choose per-request whether to pay for reasoning, instead of running a [dedicated reasoning model](/ai/deepseek-r1) beside a fast one. ## The license, and what commercial use really permits Every Qwen 3 model, from the 0.6B to the 235B flagship, ships under **[Apache 2.0](https://huggingface.co/Qwen/Qwen3-235B-A22B)**. In enterprise legal review, that one fact does more work than any benchmark: - **It is a known quantity.** Apache 2.0 has two decades of history and pre-existing approval in most corporate open-source policies. The review is a lookup, not an analysis. Compare the [Llama 4 Community License](/ai/llama-4), which is short but bespoke: attribution clauses, derivative naming rules, and a user-threshold gate that each need a lawyer's read. - **It includes an explicit patent grant**, which MIT does not. For patent-sensitive industries this is a real, if rarely decisive, point in Apache 2.0's favor. - **No attribution badge, no naming rules, no usage thresholds.** Fine-tune it, rename it, white-label it, embed it in a product you sell. Standard Apache 2.0 conditions apply (keep the license text and notices in redistributed source), none of which touch your product's UI or brand. - **Uniformity across the ladder matters more than it looks.** The license story does not change when you move from the 4B on an edge box to the 235B in the datacenter. One legal approval covers the entire deployment surface, now and as it grows. ## Real deployment cost The ladder is the cost story. Working from published parameter counts (weights only, before KV cache): - **Qwen3-4B at BF16**: roughly 8GB. Consumer GPUs, small cloud instances, serious edge hardware. - **Qwen3-32B at BF16**: roughly 64GB, a single 80GB-class GPU; at 4-bit quantization roughly 16GB, a single 24GB card. - **Qwen3-30B-A3B**: 30B of memory, but only 3B parameters active per token, so it serves with the per-token compute of a small model. This checkpoint is frequently the price-performance sweet spot in our evals. - **Qwen3-235B-A22B at FP8**: roughly 235GB of weights, a multi-GPU node, with 22B-active MoE economics per token. For teams that want managed capacity before committing GPUs: as of July 15, 2026, [Together AI lists](https://www.together.ai/pricing) its Qwen3-235B-A22B FP8 throughput tier at $0.20 per million input tokens and $0.60 per million output tokens. We treat hosted pricing like that as the rent-before-you-buy phase of a [self-hosted versus managed](/compare/self-hosted-vs-managed-llm) decision, not the end state for regulated data. ## Latency and eval behavior that matters - **Thinking mode changes the cost curve, per request.** With thinking on, the model emits reasoning tokens before the answer: better on hard tasks, slower and more expensive on everything. The engineering win is that this is a request-level flag, so your gateway can route by difficulty without a second deployment. - **Sampling settings are documented and non-optional.** The model card specifies temperature 0.6, top-p 0.95 for thinking mode, temperature 0.7, top-p 0.8 for non-thinking, and warns in capitals against greedy decoding in thinking mode because of repetition loops. Bake these into the serving layer; do not leave them to application defaults. - **Context is 32K native on the flagship, 131K with YaRN.** Long-context work needs the YaRN rope-scaling configuration enabled and validated; do not assume 128K-class behavior out of the box. - **119 languages** makes the family a default candidate for multilingual products, an area where the [Llama 4](/ai/llama-4) instruct models list 12 languages. - **We quote no leaderboard numbers here.** The Qwen team publishes competitive claims against frontier reasoning models; as with every vendor, the numbers that matter are the ones from evals on your task, your documents, your language mix. ## When to use it, and when not **Use Qwen 3 when:** - The deployment is on-prem or sovereign-cloud in a regulated industry and license friction must be near zero. - You need one model family across heterogeneous hardware: edge devices, workstation inference, and datacenter serving, with a single chat template and one legal review. - Your product roadmap includes distributing or white-labeling a fine-tuned model under your own brand, which Apache 2.0 permits and the Llama license complicates. - The workload mixes quick interactive requests and hard reasoning ones, and per-request thinking control beats operating two model fleets. **Do not use it when:** - You need the absolute frontier of reasoning quality regardless of ownership; evaluate [DeepSeek R1](/ai/deepseek-r1) and the closed APIs against it on your tasks. - The workload is natively multimodal document intake; the core Qwen 3 line is a text family, and Llama 4's early-fusion design or dedicated vision-language models fit better. - Your compliance regime requires a contractual counterparty standing behind model behavior. Apache 2.0 weights come with no warranty; that gap is filled by your engineering discipline, not the license. ## How we would architect it for a client The recurring shape is the regulated-SaaS pattern we know from building [PeoplePlus](/case-studies/peopleplus), our own HR platform, where tenant data boundaries and predictable unit economics matter more than peak benchmark scores: 1. **One family, three tiers.** Qwen3-4B for high-volume classification and routing, Qwen3-32B (or 30B-A3B) as the workhorse for generation and extraction, and the 235B flagship reserved for the low-volume hard tier. Because the chat template and tokenizer are shared, promotion between tiers is an eval decision, not a re-integration project. 2. **Thinking mode as a gateway policy.** The routing layer sets `enable_thinking` per request class and caps reasoning budgets, so cost is governed centrally instead of negotiated per feature team. This is the same routing discipline described in our [DeepSeek R1 brief](/ai/deepseek-r1), collapsed into one model family. 3. **Sovereign deployment first.** Weights live in the client's [sovereign cloud](/services/sovereign-cloud) footprint with no vendor in the inference path; hosted endpoints are used only for pre-commitment evaluation, then retired. 4. **Eval harness before model choice.** Our [model engineering](/services/model-engineering) engagements start by building the client's task-level eval set, then let the ladder compete: the smallest Qwen 3 that clears the quality bar wins the tier. That procedure, more than any single launch, is why this family keeps ending up in production: it almost always has a rung that clears the bar at the lowest infrastructure cost in the room. Apache 2.0 at every size is not a marketing detail. It is the property that lets one architecture, one legal review, and one integration serve a product from pilot to scale. That is the whole brief. #### Frequently Asked Questions **Q: Is Qwen 3 officially Apache 2.0, and where is that officially stated?** Yes. The official statement lives in two places maintained by the Qwen team at Alibaba: the LICENSE file inside each model repository on Hugging Face, which carries the verbatim Apache 2.0 text (for example huggingface.co/Qwen/Qwen3-235B-A22B/blob/main/LICENSE), and the official GitHub repository at github.com/QwenLM/Qwen3, which states the licensing for the open-weight family. Those are the authoritative sources; every summary elsewhere, including this page, is secondary. Verify the specific checkpoint you plan to deploy rather than relying on family-wide terms, because vendors do occasionally license individual checkpoints differently from the rest of a family. **Q: What license does Qwen 3 use, and is commercial use allowed?** Yes. Every Qwen 3 model, dense and MoE, from 0.6B to 235B, is released under Apache 2.0, which permits commercial use, modification, distribution, and white-labeling with no user thresholds, no attribution badge in your product, and no naming requirements on fine-tuned derivatives. Standard Apache 2.0 conditions apply, such as retaining license text and notices when you redistribute, and it includes an explicit patent grant that MIT-licensed alternatives lack. **Q: Which Qwen 3 size should we deploy?** Run the ladder against your own eval set and take the smallest model that clears the quality bar. As a starting map from the published parameter counts: 4B for high-volume classification on modest hardware, 32B on a single 80GB GPU (or roughly 16GB at 4-bit) as the general workhorse, 30B-A3B when you want near-small-model serving cost with larger-model quality, and 235B-A22B on a multi-GPU node for the hard tier. The shared chat template makes moving between rungs an eval decision rather than a rebuild. **Q: What is Qwen 3's hybrid thinking mode, and why does it matter operationally?** Each checkpoint can either reason step-by-step in a think block before answering or respond immediately, controlled by the enable_thinking flag or /think and /no_think commands in conversation. Operationally this means one deployment serves both fast interactive traffic and hard reasoning requests, with your gateway deciding per request whether to spend reasoning tokens. Without it, you typically operate a fast model and a separate reasoning model, which doubles the serving, eval, and upgrade surface. **Q: Why does Qwen 3 keep winning regulated on-prem evaluations?** Three properties compound. Apache 2.0 across the whole family makes legal review a lookup instead of a bespoke license analysis. The 0.6B-to-235B ladder means the same family fits whatever hardware the client already has, from edge boxes to datacenter nodes, with one integration. And self-hosted weights keep every token inside the client's infrastructure boundary, which is the core data-residency requirement in finance, healthcare, and government work. Evaluations are won on the intersection of constraints, and Qwen 3 rarely fails any single gate. **Q: What context window does Qwen 3 actually have?** Per the official materials: the three smallest dense models (0.6B, 1.7B, 4B) carry 32K contexts, and the larger models 128K-class contexts. The 235B-A22B flagship is 32,768 tokens natively, extended to 131,072 with YaRN rope scaling, which must be explicitly enabled in the serving configuration. If your workload depends on long context, validate quality at your real document lengths with YaRN configured rather than assuming ceiling behavior. **Q: How does Qwen 3 compare to DeepSeek R1 and Llama 4?** They solve different constraints. DeepSeek R1 (MIT) is a dedicated 671B-scale reasoning model: strongest when maximum reasoning depth inside your own infrastructure is the requirement. Llama 4 is natively multimodal but carries a bespoke community license with attribution and derivative-naming clauses. Qwen 3's edge is breadth under one clean license: eight sizes, hybrid per-request reasoning, 119 languages, Apache 2.0 everywhere. In our evaluations the decision usually falls out of the client's constraints, and we run task-level evals rather than trusting anyone's leaderboard, including the vendors'. **Q: Can we fine-tune Qwen 3 and ship it under our own product name?** Yes. Apache 2.0 places no naming requirements or attribution badges on derivative models, so a fine-tuned Qwen 3 can ship under your brand, embedded in your product, including white-label arrangements. This is a sharp contrast with the Llama 4 Community License, which requires distributed derivatives to carry Llama at the beginning of the model name and the product to display Built with Llama. Keep the standard Apache notices in redistributed source, and the rest is your call. --- ### Kimi K2.6: Open-Weights LLM *Publisher: Moonshot AI · Paper date: 2026.04.21 · Brief date: 2026.08.09 · 11 min read* *Parameters: 1T (32B active) · License: Modified MIT* URL: https://www.bearplex.com/ai/kimi-k2-6 arXiv: https://www.kimi.com/blog/kimi-k2-6 Model card: https://huggingface.co/moonshotai/Kimi-K2.6 GitHub: https://github.com/MoonshotAI/kimi-code **Excerpt**: BearPlex's engineering brief on Kimi K2.6: verified Modified MIT terms, the 300-agent swarm claims, real API and self-hosting costs, and when it wins. > **Update, August 2026:** Moonshot has since shipped [Kimi K3](/ai/kimi-k3), a 2.8T-parameter successor with a 1M-token context. If you are choosing a Kimi model today, read that brief first. One thing carried the wrong way between generations: K2.6's Modified MIT terms described below did **not** carry over. K3 ships under a bespoke Kimi K3 License with revenue and user thresholds attached. K2.6 remains the more permissively licensed and far more deployable of the two. For two years the honest answer on open weights and agentic coding was "close, but you would not bet an unattended production run on it." Kimi K2.6 is the release that forces a re-evaluation: an open-weights model whose vendor-reported numbers edge out closed flagships on the agentic coding benchmarks, wrapped in a license that is one paragraph away from plain MIT, at output-token prices several multiples below the closed frontier. The headline everyone repeats is the agent swarm: 300 sub-agents, 4,000 coordinated steps, coding sessions that run past 12 hours. This brief takes those claims seriously enough to read them like an engineer, because the swarm story is simultaneously the most interesting thing about K2.6 and the easiest place to misbudget a project. ## What it actually is Kimi K2.6 ([Moonshot AI, announced April 21, 2026](https://forum.moonshot.ai/t/meet-kimi-k2-6-advancing-open-source-coding/369); live on Cloudflare Workers AI on April 20 US time) is a Mixture-of-Experts model with **1 trillion total parameters and 32B activated per token**, per the [model card](https://huggingface.co/moonshotai/Kimi-K2.6): 61 layers, 384 experts with 8 selected per token, a 160K vocabulary, and a **262,144-token context window**. It is natively multimodal via a 400M-parameter MoonViT vision encoder, so screenshots and design mocks go straight in, which matters for the "coding-driven design" workflows Moonshot markets. The weights ship **natively in INT4**, produced with quantization-aware training rather than post-hoc compression, and the model card recommends two operating modes: Thinking (temperature 1.0) and Instant (temperature 0.6). The MoE math is the same lesson we flagged on [Llama 4](/ai/llama-4), at a larger scale: per-token compute resembles a 32B dense model, memory resembles a trillion parameters. K2.6 is cheap to serve per token and very expensive to provision, and nothing in the swarm story changes that arithmetic. ## The agent-swarm claims, read like an engineer Moonshot's stated capability: K2.6 scales to **300 parallel sub-agents executing 4,000 coordinated steps** in a single run, up from 100 sub-agents and 1,500 steps on K2.5, per the [announcement](https://forum.moonshot.ai/t/meet-kimi-k2-6-advancing-open-source-coding/369). The long-horizon evidence in the [tech blog](https://www.kimi.com/blog/kimi-k2-6) is two case studies: a 12-plus-hour continuous run making over 4,000 tool calls across 14 iterations to optimize a small-model inference engine (throughput raised from roughly 15 to roughly 193 tokens per second), and a 13-hour run making over 1,000 tool calls to modify more than 4,000 lines of an exchange-core project. Four things a technical buyer should hold onto: 1. **These are vendor case studies, not reproducible benchmarks.** They demonstrate the model can sustain coherence across thousands of tool calls, which is genuinely the hard part of long-horizon coding. They do not tell you the success rate across attempts, and no vendor publishes that. 2. **The swarm lives in the harness, not the weights.** Sub-agent orchestration is a property of Kimi's agent products and of whatever framework you run; the open weights give you a model that is unusually good at being orchestrated. If you self-host, you are building or adopting the orchestration layer yourself. 3. **A 12-hour unattended run is a review problem before it is a capability.** Four thousand tool calls produce a diff no human reviews line by line. The engineering answer is checkpoints, test gates, and scoped write permissions, not longer leashes. 4. **300 sub-agents is a ceiling, not a default.** Parallel agents multiply token burn linearly. The right mental model is "burst capacity for decomposable work," not "300x the output." ## Benchmarks, and how to read them All of the following are Moonshot's own reported numbers from the [K2.6 tech blog](https://www.kimi.com/blog/kimi-k2-6) and [model card](https://huggingface.co/moonshotai/Kimi-K2.6); treat them as the vendor's best case until reproduced on your tasks. | Benchmark | K2.6 | GPT-5.4 | Claude Opus 4.6 | |---|---|---|---| | SWE-Bench Pro | **58.6** | 57.7 | 53.4 | | Humanity's Last Exam (w/ tools) | **54.0** | 52.1 | 53.0 | | Terminal-Bench 2.0 | **66.7** | 65.4 | 65.4 | | BrowseComp | 83.2 | 82.7 | **83.7** | | MathVision (w/ python) | 93.2 | **96.1** | 84.6 | The model card adds SWE-Bench Verified 80.2, LiveCodeBench v6 89.6, and GPQA-Diamond 90.5. Our reading: the correct summary is not "K2.6 beats the closed frontier," it is "K2.6 sits inside the frontier tier on agentic coding, with margins thin enough to be eval noise in both directions." It trails on browsing by half a point, on vision-math by three, and the same model card has it behind both rivals on GPQA-Diamond. That an open-weights model under a near-MIT license is in this table at all is the actual news; which cell is bold will change again within a quarter, and Moonshot already shipped a K2.7 Code variant (more below). ## The license: one clause away from MIT The [LICENSE file](https://huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENSE) is standard MIT plus a single condition: if the software or derivatives are used in a commercial product or service with **more than 100 million monthly active users, or more than 20 million US dollars in monthly revenue**, you must "prominently display 'Kimi K2.6' on the user interface" of that product. The practical read for client work: for effectively every product we would build with it, this is plain MIT. No attribution below the threshold, no derivative-naming rule, no acceptable-use gate riding along with the grant. Contrast that with the [Llama 4 Community License](/ai/llama-4), where attribution is mandatory at any scale and your distributed fine-tune must carry Meta's brand in its name. One caution the threshold deserves: unlike Llama's snapshot-on-release-date test, this clause reads as an ongoing condition, so a product that later crosses 100M MAU or $20M monthly revenue picks up the UI-attribution duty at that point. Put it in the legal file and move on. ## Real cost **Hosted API.** Verified against the [official pricing page](https://platform.kimi.ai/docs/pricing/chat-k26) as of July 2026: **$0.95 per million input tokens ($0.16 on cache hit) and $4.00 per million output tokens**, at the full 262,144-token context. For comparison, the closed frontier flagships we track charge many multiples of that on output; the verified table is in our [GPT-5 brief](/ai/gpt-5). Marketplace routing runs even lower: [OpenRouter](https://openrouter.ai/moonshotai/kimi-k2.6) lists K2.6 across multiple providers around $0.66 input / $3.41 output. The trap is reading cheap tokens as cheap tasks. Swarm-style workloads are token multipliers: parallel sub-agents each carry context, and a 12-hour run compounds thinking tokens for half a day. The metric that matters is **cost per merged, passing change**, not cost per million tokens, and only your own harness can measure it. Budget the eval before the rollout. **Self-hosting.** The [official vLLM recipe](https://recipes.vllm.ai/moonshotai/Kimi-K2.6) is verified on **8x H200, or roughly 640GB of aggregate VRAM for the INT4 weights**, on vLLM 0.19.1+ with tensor parallel 8, before KV cache, which at 262K contexts is substantial. Supported engines per the model card: vLLM, SGLang, and KTransformers. This is the same provisioning cliff as every trillion-parameter MoE: an 8-GPU H200-class node is the entry ticket, and the per-token serving economics only start winning after that ticket is paid. ## Where you can run it - **Moonshot's API** at platform.moonshot.ai, OpenAI-compatible. - **Cloudflare Workers AI**, which added `@cf/moonshotai/kimi-k2.6` on [April 20, 2026](https://developers.cloudflare.com/changelog/post/2026-04-20-kimi-k2-6-workers-ai/) with vision inputs, multi-turn tool calling, and an OpenAI-compatible endpoint. This is the notable one: a trillion-parameter open model served from western edge infrastructure with no Moonshot account in the data path. - **Multi-provider marketplaces** (OpenRouter and the GPU clouds it fronts). - **Self-hosted** in your own VPC on the vLLM/SGLang/KTransformers stack. - **[Kimi Code CLI](https://github.com/MoonshotAI/kimi-code)**, Moonshot's MIT-licensed TypeScript terminal agent (`npm install -g @moonshot-ai/kimi-code`, Node 24.15+), with built-in coder/explore/plan subagents and MCP support. It is the first-party harness tuned for the model, and because the model speaks OpenAI-compatible APIs, third-party harnesses work too. One paragraph procurement will ask about, so say it early: Moonshot AI is a Beijing-based company. The weights are weights; where you run them is entirely your choice, and self-hosting or Workers AI keeps every byte off Moonshot's infrastructure. But if your compliance regime restricts routing client data through a China-based vendor's hosted API, decide the deployment lane first, because it changes the cost model and the integration. For regulated clients we treat this as a [sovereign-cloud](/services/sovereign-cloud) question with a clean answer, not a disqualifier. ## Vendor velocity Moonshot moves at the same cadence we documented for [OpenAI](/ai/gpt-5). K2.5 to K2.6 took roughly a quarter, and on [June 12, 2026](https://developers.cloudflare.com/changelog/post/2026-06-12-kimi-k2-7-code-workers-ai/) a code-optimized **Kimi K2.7 Code** variant landed on Workers AI, claiming 30% fewer reasoning tokens than K2.6 on reasoning-heavy work. The difference from the closed-API world: open weights cannot be retired out from under you. Pin the checkpoint you validated, keep the eval suite runnable, and treat each K2.x release as an optional upgrade with a regression test, not a forced migration. ## When to use it, and when not **Use Kimi K2.6 when:** - Agentic coding volume dominates your bill and cost-per-task is the metric. The output-token economics against closed flagships are the whole argument, and at frontier-tier SWE-Bench results the quality discount is thin or absent. - You need frontier-tier coding inside your own boundary, especially with vision in the loop. The open-weights short list for this lane is K2.6 and the MIT-licensed [GLM-5.2](/ai/glm-5-2), which posts a higher SWE-bench Pro score but is text-only; K2.6 is the one that reads screenshots and mocks natively. Eval both on your repositories. - Brand and legal cleanliness matter: no attribution below the 100M-MAU/$20M threshold, no derivative-naming clause, fine-tunes ship under your name. - You want harness freedom: OpenAI-compatible surface on every lane, from Moonshot's API to Workers AI to your own vLLM cluster. **Do not use it when:** - The workload is browsing-heavy research or vision-math, where Moonshot's own table shows it trailing Claude Opus 4.6 and GPT-5.4 respectively. Route those lanes to the model that wins them. - You cannot fund the provisioning: roughly 640GB of aggregate VRAM for self-hosting, or a hosted-API dependency you must clear with compliance. - Your plan is "turn on 300 agents and ship faster." Without checkpoints, test gates, and cost-per-task measurement, the swarm multiplies spend and review debt, not throughput. ## How we would architect it for a client The same gateway discipline as every model lane in our [model engineering](/services/model-engineering) work, with swarm-specific guardrails: 1. **A model-agnostic gateway** owns the checkpoint and the deployment lane, so moving between Moonshot's API, Workers AI, and a self-hosted cluster is configuration. Data-residency tiering falls out for free: sensitive lanes pin to self-host or edge, everything else routes to the cheapest compliant endpoint. 2. **Swarms behind gates, not leashes.** Long-horizon runs execute in sandboxed branches with test suites as checkpoints and scoped write permissions; a 4,000-tool-call session lands as a stack of reviewable, individually-testable changes. This is standard [autonomous-agents](/services/autonomous-agents) practice regardless of vendor. 3. **Cost-per-merged-change as the eval metric.** We benchmark K2.6 against the incumbent lane on the client's own repositories, counting total tokens per accepted change, not per request. Cheap tokens that triple the retry count are not cheap. 4. **Quarterly re-runs.** K2.7 Code already exists; the eval suite that admitted K2.6 is the same one that decides whether its successor replaces it. K2.6 is strong evidence that the open-weights lane has reached the agentic-coding frontier. The engineering work is exactly where it always was: in the harness, the gates, and the economics. #### Frequently Asked Questions **Q: Is Kimi K2.6 really open source, and what does the Modified MIT license require?** The weights are downloadable from Hugging Face under a Modified MIT license: standard MIT terms plus one added clause. If the software or a derivative is used in a commercial product with more than 100 million monthly active users or more than 20 million US dollars in monthly revenue, you must prominently display 'Kimi K2.6' on that product's user interface. Below those thresholds it behaves like plain MIT: no attribution requirement, no derivative-naming rule, and your fine-tunes ship under your own brand. **Q: How much does Kimi K2.6 cost via the API?** Per Moonshot's official pricing page as of July 2026: $0.95 per million input tokens, $0.16 per million on cache hits, and $4.00 per million output tokens, at the full 262,144-token context window. Marketplace routing via OpenRouter lists it around $0.66 input and $3.41 output across multiple providers. The caveat for agentic workloads: parallel sub-agents and long thinking runs multiply token volume, so we recommend budgeting on measured cost per completed task, not on the per-token rate. **Q: What hardware does Kimi K2.6 need to self-host?** The official vLLM recipe is verified on 8x H200 GPUs, or roughly 640GB of aggregate VRAM for the native INT4 weights, running vLLM 0.19.1+ with tensor parallelism of 8, before KV cache. KV cache at 262K-token contexts adds real headroom on top. Moonshot's model card lists vLLM, SGLang, and KTransformers as the recommended inference engines. Because only 32B parameters activate per token, serving throughput is strong once memory is provisioned; as with every large MoE, the cost cliff is the provisioning. **Q: Is the 300-agent swarm real?** It is Moonshot's stated architectural ceiling: up to 300 parallel sub-agents executing 4,000 coordinated steps per run, up from 100 sub-agents and 1,500 steps in K2.5, with published case studies of 12-hour-plus continuous coding runs making thousands of tool calls. Two engineering qualifiers matter. The orchestration lives in the agent harness, so self-hosters build or adopt that layer separately from the weights. And parallel agents multiply token spend linearly, so the swarm is burst capacity for decomposable work, not a free 300x on throughput. **Q: Is Kimi K2.6 actually better than Claude or GPT-5 at coding?** On Moonshot's own reported benchmarks it leads SWE-Bench Pro (58.6 vs 57.7 for GPT-5.4 and 53.4 for Claude Opus 4.6), Terminal-Bench 2.0, and Humanity's Last Exam with tools, while trailing Claude Opus 4.6 on BrowseComp and GPT-5.4 on vision-math. Our reading is that it sits inside the frontier tier for agentic coding with margins thin enough to be eval noise, which at these prices is the story. As always, vendor tables are the vendor's best case; run the comparison on your own repositories before committing a lane. **Q: Can we use Kimi K2.6 if our compliance rules restrict China-based vendors?** Usually yes, because the model does not require Moonshot's infrastructure. The open weights can be self-hosted entirely inside your VPC, and Cloudflare has served the model on Workers AI since April 20, 2026 with an OpenAI-compatible endpoint, keeping your data off Moonshot's API entirely. The compliance decision is about the deployment lane, not the model: we resolve the data-path question with procurement first, because it determines both the cost model and the integration architecture. **Q: Does BearPlex deploy Kimi K2.6 in client work?** We evaluate it as a standard candidate wherever agentic-coding volume, open-weights requirements, or cost-per-task economics drive the decision. It enters through the same model-agnostic gateway as every other lane: benchmarked on the client's own repositories with cost per merged change as the metric, long-horizon runs constrained by test gates and scoped permissions, and the deployment lane (Moonshot API, Workers AI, or self-hosted) chosen by the client's data-residency requirements rather than by default. --- ### Kimi K3: Open-Weights LLM *Publisher: Moonshot AI · Paper date: 2026.07.16 · Brief date: 2026.08.09 · 10 min read* *Parameters: 2.8T total (104B activated per token) · License: Kimi K3 License (custom)* URL: https://www.bearplex.com/ai/kimi-k3 arXiv: https://github.com/MoonshotAI/Kimi-K3 Model card: https://huggingface.co/moonshotai/Kimi-K3 GitHub: https://github.com/MoonshotAI/Kimi-K3 **Excerpt**: BearPlex's engineering brief on Kimi K3: the verified Kimi K3 License thresholds, why it is not MIT like K2.6 was, the 2.8T architecture, and who it fits. **Kimi K3 is not MIT-licensed, and that is the most important thing to know about it.** Its predecessor [Kimi K2.6](/ai/kimi-k2-6) shipped under Modified MIT, one paragraph away from plain MIT. K3 ships under a bespoke document titled the **Kimi K3 License**, tagged on Hugging Face as license:other. Between one generation and the next, Moonshot tightened terms. Most coverage of this release led with the parameter count and treated the license as a footnote, which is backwards for anyone deciding whether to build on it. The terms are still broadly permissive, and for the large majority of teams reading this they will never bind. But "broadly permissive with conditions" is a different procurement conversation from "Apache 2.0", and the difference shows up in legal review, not in a benchmark table. ## What it actually is Kimi K3 ([Moonshot AI, launched July 16, 2026](https://github.com/MoonshotAI/Kimi-K3) on Kimi Code and the Kimi app, with open weights following later that month) is the largest open-weight model released to date. Per the [official repository](https://github.com/MoonshotAI/Kimi-K3): - **2.8 trillion total parameters, 104 billion activated per token.** A Mixture-of-Experts model with **896 experts, 16 selected per token**, which is a far sparser configuration than the field has been shipping. - **93 layers**, split as 69 Kimi Delta Attention (KDA) layers and 24 Gated MLA layers. - **A 1,048,576-token context window**, a full million tokens. - **SiTU-GLU activation**, with **MXFP4 weights and MXFP8 activations**, so the released checkpoint is natively low-precision rather than a post-hoc quantization of something larger. - Text and image input, with agentic tool use as a first-class target. Moonshot attributes the efficiency story to a Stable LatentMoE framework and reports roughly a 2.5x improvement in scaling efficiency over Kimi K2. Treat that as a vendor claim: it is a statement about their training economics, not a benchmark you can reproduce, and we quote no leaderboard numbers here for the usual reason. The numbers that decide a deployment are the ones from evals on your task. ## The license, read properly This is where the engineering decision actually gets made. The [LICENSE file](https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE) grants the rights you would expect from an open-weights release: use, copy, modify, merge, publish, distribute, sublicense, and sell. Download, self-host, fine-tune, and quantize are all permitted. Then it adds conditions: - **The Model-as-a-Service threshold.** If the aggregate revenue of the licensee and its affiliates exceeds **20 million US dollars over any consecutive 12 months**, offering the model as a service requires a **separate agreement with Moonshot AI**. This is the clause that matters most and gets quoted least. - **The branding threshold.** A product or service with **more than 100 million monthly active users, or more than 20 million US dollars in monthly revenue**, must display "Kimi K3" prominently on its user interface. - **Section 4 exemptions, which are broader than the thresholds suggest.** The restrictions do not apply to internal use, defined as any use that does not make the software, its outputs, or its underlying capabilities available to third parties. They also do not apply to use accessed through Moonshot's official products or certified inference partners. Read those together and a clear map falls out: - **Internal deployment: unconditionally fine.** If you are running K3 inside your own company on your own work, no threshold applies to you at any revenue. This covers most enterprise adoption. - **Building a product on top of it: fine until you are very large.** The branding clause needs 100 million monthly active users or 20 million dollars of *monthly* revenue. If you cross that, a wordmark on your interface is not the constraint that will be keeping you up at night. - **Reselling inference: read carefully.** If your business is serving this model to other people and your group revenue clears 20 million dollars annually, you need a signed agreement before you start, not after. The threshold is on *your* revenue, not on your K3 revenue, which is easy to misread. Compare this to [Qwen 3 under plain Apache 2.0](/ai/qwen-3), where legal review is a lookup rather than an analysis, and the practical trade becomes visible. Apache 2.0 has no thresholds to monitor and no clause that changes behavior as you grow. The Kimi K3 License has two, plus an obligation to notice when you cross them. ## The constraint nobody is talking about: you are not self-hosting this 2.8 trillion parameters is not a number most organizations can serve, and the released weights run to well over a terabyte. Even at MXFP4, this is a multi-node deployment with an interconnect budget, not a model you stand up on a spare box to evaluate. For the overwhelming majority of teams, "open weights" here means *auditable and portable in principle*, not *self-hosted in practice*. That reframes the license question. If you are consuming K3 through an API, the branding clause is almost certainly irrelevant to you and the Model-as-a-Service clause belongs to your provider. The reason to care about the license at all is the reason open weights matter generally: the model cannot be deprecated out from under you, you can move providers, and if the economics ever justify bringing it in-house the door is open. Those are real properties. They are just not the same as running it next week. ## When to use it, and when not **Consider Kimi K3 when:** - The workload is long-horizon agentic work and the million-token context is doing real work rather than padding a spec sheet. - You want frontier-class capability with an exit path from any single API vendor, and you are willing to accept a conditional license to get it. - Your usage is internal, where Section 4 removes the conditions entirely. - Multimodal input matters and you would otherwise be running a separate vision model. **Do not reach for it when:** - Your procurement process requires a standard OSI license with no thresholds. That is a legitimate policy, and [Qwen 3](/ai/qwen-3) or [DeepSeek R1](/ai/deepseek-r1) satisfy it where K3 does not. - You intend to build a Model-as-a-Service business on it and your group revenue is already past 20 million dollars, unless you have the separate agreement signed first. - You need on-premise deployment on hardware you already own. At this scale, that is a datacenter project, and a smaller model that clears your quality bar is the better engineering answer. - The task is well-served by a 30B-class model. Sparsity helps the serving cost, but 104B activated parameters per token is still 104B activated parameters per token. ## How we would evaluate it for a client The procedure does not change because the model got bigger: 1. **Build the eval set before touching the model.** Our [model engineering](/services/model-engineering) engagements start with the client's task-level evals, because a million-token context and a trillion-parameter count tell you nothing about whether the thing answers your questions correctly. 2. **Run the license question in parallel with the technical one.** Legal review and evaluation should finish in the same week. Discovering a threshold clause after a successful pilot is how a project loses a quarter. 3. **Assume API access first.** Rent before you buy is always the right sequencing, and at 2.8T it is the only sequencing. Reserve the self-hosting conversation for the point where volume, latency, or [data residency](/services/sovereign-cloud) makes it unavoidable. 4. **Check whether a smaller model wins.** In our evaluations the frontier model usually is not the answer. It sets the ceiling, and then something four rungs down clears the client's bar at a fraction of the cost. The genuinely notable thing about Kimi K3 is not the parameter count, which will be beaten. It is that the largest open-weight release to date arrived with more license conditions than the smaller one before it. If that becomes the pattern as open-weight models approach frontier capability, "open weights" stops being a single category and starts being a spectrum that needs reading every time. #### Frequently Asked Questions **Q: What license is Kimi K3 released under, and is it open source?** Kimi K3 is released under a bespoke document titled the Kimi K3 License, tagged on Hugging Face as license:other. It is not MIT and it is not Apache 2.0, which makes open-weight the accurate label rather than open source. The license grants the rights you would expect (use, copy, modify, merge, publish, distribute, sublicense, sell, plus self-hosting, fine-tuning, and quantization) and then attaches two conditions tied to scale: a Model-as-a-Service clause and a branding clause. Notably this is a tightening relative to Kimi K2.6, which shipped under Modified MIT. **Q: Can we use Kimi K3 commercially without signing anything?** In most cases yes. Section 4 of the license exempts internal use entirely, defined as any use that does not make the software, its outputs, or its underlying capabilities available to third parties, and that exemption has no revenue ceiling. Use accessed through Moonshot's official products or certified inference partners is also exempt. A separate agreement with Moonshot is required specifically for offering the model as a service when the aggregate revenue of the licensee and its affiliates exceeds 20 million US dollars over any consecutive 12 months. Verify the current license text against your own situation before production deployment rather than relying on this summary. **Q: When does the Kimi K3 branding requirement apply?** Only at very large scale. The license requires that Kimi K3 be prominently displayed on the user interface of a product or service with more than 100 million monthly active users, or more than 20 million US dollars in monthly revenue. Those are monthly figures, not annual, so the threshold sits far above where most companies operate. If you cross it, a wordmark in your interface is unlikely to be the binding constraint on your business. **Q: How big is Kimi K3 and can we self-host it?** Kimi K3 has 2.8 trillion total parameters with 104 billion activated per token, using a Mixture-of-Experts configuration of 896 experts with 16 selected per token across 93 layers, and a 1,048,576-token context window. Weights ship natively at MXFP4 precision with MXFP8 activations. Realistically, self-hosting is a multi-node datacenter project rather than something a team stands up to evaluate, and for most organizations open weights here means auditable and portable rather than actually self-hosted. Plan on API access first and treat in-house serving as a decision driven by volume, latency, or data residency. **Q: How does Kimi K3 compare to Kimi K2.6?** K3 is substantially larger (2.8T total and 104B activated versus K2.6's 1T total and 32B activated), extends context from 262,144 tokens to 1,048,576, and moves to a much sparser Mixture-of-Experts design with 896 experts selecting 16 per token. Architecturally it introduces Kimi Delta Attention layers alongside Gated MLA. The change that matters for procurement runs the other way: K2.6 was Modified MIT, and K3 is a custom license with revenue and user thresholds attached. Bigger model, tighter terms. **Q: Should we choose Kimi K3 or Qwen 3 for a regulated deployment?** For regulated on-premise work, Qwen 3 is usually the better fit, and it is not close. Qwen 3 is Apache 2.0 across its entire family, so legal review is a lookup with no thresholds to monitor as you grow, and its 0.6B to 235B ladder means something in the family actually fits the hardware a client already has. Kimi K3 is a frontier-scale model with a conditional license and a serving footprint most regulated environments cannot host. Choose K3 when you need its capability and can consume it through an API; choose Qwen 3 when sovereignty, license cleanliness, and deployability on existing hardware are the constraints. **Q: Does the 1 million token context window change how we build?** Less than the number suggests. A long context window raises the ceiling on what fits in a single request, but it does not make retrieval unnecessary, and it does not guarantee the model attends well across the whole window. In practice we still build retrieval pipelines, still measure quality at the document lengths a client actually uses rather than at the advertised maximum, and still find that carefully selected context beats dumped context on both accuracy and cost. Treat a million tokens as headroom for genuinely long agentic sessions, not as a replacement for a well-built RAG layer. --- ### Llama 4: Open-Weights LLM *Publisher: Meta · Paper date: 2025.04.05 · Brief date: 2026.07.15 · 10 min read* *Parameters: Scout 109B (17B active) / Maverick 400B (17B active) · License: Llama 4 Community License* URL: https://www.bearplex.com/ai/llama-4 arXiv: https://ai.meta.com/blog/llama-4-multimodal-intelligence/ Model card: https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct GitHub: https://github.com/meta-llama/llama-models **Excerpt**: BearPlex's engineering brief on Llama 4: Scout and Maverick specs, what the Llama 4 Community License really requires (700M MAU, Built with Llama), and deployment cost. **Llama 4 is free for commercial use for most organizations, but it is not permissively licensed like Apache 2.0 or MIT software.** The Llama 4 Community License has a 700 million monthly-active-user gate, visible attribution, derivative naming rules, redistribution paperwork, and an EU restriction for its multimodal models. Those conditions are the decision, not a footnote. Every open-weights model brief has a license section. This one *is* the license section, with a model attached. Llama 4 is a genuinely capable multimodal family, but in our client evaluations the technical comparison is rarely what decides the engagement. The Llama 4 Community License is. It is short, readable, and full of obligations that product teams routinely discover after the architecture is committed. Read it first; the model specs will still be there afterward. ## What it actually is Llama 4 ([Meta, April 5, 2025](https://ai.meta.com/blog/llama-4-multimodal-intelligence/)) is Meta's first natively multimodal, Mixture-of-Experts open-weights generation. Text and image tokens are fused early into one backbone rather than bolted together with an adapter. Two models shipped with downloadable weights: - **Llama 4 Scout**: 17B active parameters, 16 experts, **109B total**, with a claimed **10 million token** context window. Meta states it fits on a single H100 with Int4 quantization. - **Llama 4 Maverick**: 17B active parameters, 128 experts, **400B total**, with a **1M token** context window per the [model card](https://huggingface.co/meta-llama/Llama-4-Maverick-17B-128E-Instruct), released in BF16 and FP8 (Meta: FP8 "fits on a single H100 DGX host"). Both are distilled from **Llama 4 Behemoth** (288B active, roughly 2T total), which was still training at announcement time. The instruct models cover 12 languages, with a knowledge cutoff of August 2024. The MoE math is the part to internalize: per-token compute resembles a 17B dense model while memory requirements resemble the total parameter count. Llama 4 is therefore cheap to *run* per token and expensive to *hold* in memory, the exact inverse of what most capacity planning assumes. ## The license, and what commercial use really permits The [Llama 4 Community License](https://github.com/meta-llama/llama-models/blob/main/models/llama4/LICENSE) (effective April 5, 2025) grants a royalty-free, worldwide license to use, modify, and redistribute. It is not open source in the OSI sense, and four clauses matter for client products: **1. The 700M-MAU gate.** Quoting the license: "If, on the Llama 4 version release date, the monthly active users of the products or services made available by or for Licensee, or Licensee's affiliates, is greater than 700 million monthly active users in the preceding calendar month, you must request a license from Meta." For nearly every company on earth this clause is irrelevant; it exists to exclude Meta's direct rivals. But note the measurement date: your MAU *on the Llama 4 release date*. Crossing 700M later does not retroactively strip the license. **2. "Built with Llama" is mandatory and visible.** If you distribute Llama Materials, or a product or service (including another AI model) that contains them, you must "prominently display 'Built with Llama' on a related website, user interface, blogpost, about page, or product documentation." For an internal tool nobody outside the company sees, this is trivial. For a white-label product your client resells under their own brand, it is a real conversation: the attribution requirement travels with the model, and "prominently" is Meta's word, not ours. **3. Derivative models must carry the name.** If you use Llama 4, or its outputs, to build a fine-tuned or improved model that you distribute, the license requires you to "include Llama at the beginning of any such AI model name." Your fine-tune cannot ship as "AcmeMed-8B"; it ships as "Llama-AcmeMed" or it stays in-house. Purely internal models you never distribute are not caught by this. **4. Redistribution paperwork.** Distributing the weights or a derivative means shipping a copy of the license agreement and the notice file text: "Llama 4 is licensed under the Llama 4 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved." Usage must also comply with the [Acceptable Use Policy](https://www.llama.com/llama4/use-policy). **5. The EU multimodal restriction sits in the use policy.** Meta's [official Llama 4 use policy](https://github.com/meta-llama/llama-models/blob/main/models/llama4/USE_POLICY.md), which the license incorporates, says the Section 1(a) rights are not granted for Llama 4 multimodal models to an individual domiciled in the European Union or a company whose principal place of business is there. Scout and Maverick are natively multimodal. The policy separately says this restriction does not apply to end users of a product or service that incorporates those models. An EU organization licensing or deploying the model itself should therefore treat direct use as blocked unless Meta provides a separate authorization path; an EU customer merely using an integrated product falls under the carve-out. This is easy to miss because it appears in the use policy rather than the main license text. The practical read for client work outside the EU restriction: for internal deployments, the license costs you a line in the docs. For customer-facing products, budget a short legal review and a product decision about the attribution badge. For products whose whole value is a proprietary fine-tuned model you distribute to customers, the naming clause can be disqualifying on brand grounds alone, and [Apache 2.0 alternatives](/ai/qwen-3) or [MIT alternatives](/ai/deepseek-r1) start winning the evaluation before benchmarks are even discussed. ## Real deployment cost Meta's own hardware claims set the floor: Scout on a single H100 at Int4; Maverick FP8 on a single 8x H100 DGX host. Working from the published parameter counts (weights only, before KV cache): - **Scout at BF16**: 109B parameters is roughly 218GB, a multi-GPU node. - **Scout at Int4**: roughly 55GB, which is how Meta's single-H100 claim works. Quantization to 4-bit is a quality tradeoff you eval, not a free lunch. - **Maverick at FP8**: roughly 400GB, hence the 8x H100 host sizing. Two cost notes from practice. First, the 17B-active MoE design means throughput per GPU-hour is strong once the memory is provisioned; the cost cliff is the provisioning, not the serving. Second, that giant context window is not free: KV cache grows with tokens actually in context, so "10M context" workloads carry memory and latency costs that have nothing to do with the weights. Treat the 10M figure as Meta's stated architectural ceiling, and eval long-context quality on your own documents before designing a product around it. ## Latency and eval behavior that matters - **Interactive-grade latency.** Unlike [reasoning models](/ai/deepseek-r1), Llama 4 does not spend thousands of thinking tokens before answering. For chat, extraction, and summarization the latency profile is that of a 17B dense model. - **Multimodality is native, not bolted on.** Early fusion means image-plus-text prompts (screenshots, scanned forms, photos with captions) go through one model, simplifying pipelines that previously chained a vision encoder into a text LLM. - **Language coverage is thinner than rivals.** Twelve supported languages, versus 119 claimed by Qwen 3. For multilingual products, check your language list first. - **Benchmark claims deserve skepticism.** We are deliberately quoting no launch benchmark numbers in this brief. Vendor comparisons, every vendor's, are marketing until reproduced on your own tasks; run task-level evals before committing. ## When to use it, and when not **Use Llama 4 when:** - The workload is multimodal document intake (forms, IDs, invoices, screenshots) and you want one open model, self-hosted, instead of an OCR-plus-LLM chain. - You need long-context review over large document sets inside your own VPC and have validated quality at your actual context lengths. - Attribution is a non-issue (internal tools) and the Meta ecosystem tooling matters to your team. **Do not use it when:** - Your product distributes a fine-tuned model under your own brand; the naming clause forces "Llama" into the name. - White-label constraints make "Built with Llama" undisplayable. - You need small edge deployments; the family has no small dense models, which is exactly where the [Qwen 3 ladder](/ai/qwen-3) is strong. ## How we would architect it for a client The fit we see most often is document-heavy regulated platforms, the same shape as the NDIS provider-management platform we build as a long-term development partner ([Vertex360](/case-studies/vertex360)): high volumes of scanned forms, compliance documents, and mixed-media evidence that cannot leave the client's environment. The pattern: 1. **Scout as the multimodal intake layer** in the client's [sovereign cloud](/services/sovereign-cloud): single-GPU Int4 deployment handling classification, extraction, and structured summarization of image-plus-text documents, with outputs validated against schemas in code. 2. **License compliance as a deliverable, not an afterthought.** The engagement checklist includes the attribution placement decision, the notice file in every distributed artifact, and a naming review for any fine-tuned derivative before it gets a product name. 3. **A fallback lane to a permissively licensed model.** Because the license terms are product-shaping, we architect the model interface so Llama 4 is swappable; if the client's product strategy later collides with the attribution or naming clauses, the migration is a config change, not a rebuild. This is standard practice in our [model engineering](/services/model-engineering) work regardless of vendor. Llama 4 is a good model wrapped in a license that is fine for most and fatal for some. Which one you are is a twenty-minute legal read. Do it first. #### Frequently Asked Questions **Q: Can we use Llama 4 in a commercial product?** Yes, for most organizations outside the EU multimodal restriction. The Llama 4 Community License grants royalty-free commercial use unless your products exceeded 700 million monthly active users on the Llama 4 release date. The obligations that affect normal companies are the Built with Llama attribution, the license and notice file for redistribution, and the naming rule for distributed fine-tuned derivatives. Meta's incorporated use policy separately withholds the multimodal license grant from EU-domiciled individuals and companies whose principal place of business is in the EU. **Q: Can an EU company use Llama 4 Scout or Maverick?** Not as a direct model licensee under the standard grant as written. Meta's Llama 4 use policy says the Section 1(a) rights are not granted for multimodal Llama 4 models to individuals domiciled in the European Union or companies whose principal place of business is there. Scout and Maverick are natively multimodal. The policy carves out end users of products or services that incorporate the models, so an EU customer can use an integrated product, but an EU organization building or operating directly on the weights should obtain legal advice and separate written authorization from Meta or choose a permissively licensed alternative. **Q: Do we really have to display 'Built with Llama' in our product?** If you distribute a product or service containing Llama Materials, yes: the license requires prominently displaying 'Built with Llama' on a related website, user interface, blog post, about page, or product documentation. You have flexibility on placement (documentation counts), but not on existence. For internal-only tools that are never distributed, the requirement has no practical bite. For white-label products, resolve this with your client before committing the architecture. **Q: Can we fine-tune Llama 4 and release the model under our own name?** Not under a name of your choosing. The license states that if you use Llama Materials or their outputs to create and distribute an improved AI model, you must include 'Llama' at the beginning of the model name. Internal fine-tunes you never distribute are unaffected. If distributing a branded proprietary model is core to your product, this clause is usually the reason evaluations shift to Apache 2.0 (Qwen 3) or MIT (DeepSeek R1) alternatives. **Q: Is Llama 4 open source?** No, not in the OSI sense, and Meta's own license title says 'Community License' rather than claiming otherwise. The weights are downloadable and free for most commercial use, but the license carries field-of-use conditions (Acceptable Use Policy), attribution obligations, naming requirements on derivatives, and a user-threshold gate. 'Open weights under a conditional license' is the accurate description, and the difference matters mostly when you distribute models or ship white-label products. **Q: What hardware does Llama 4 actually need?** Per Meta's published claims: Scout (109B total, 17B active) fits a single H100 with Int4 quantization, and Maverick's FP8 weights (400B total) fit a single 8x H100 DGX host. Working from parameter counts, Scout at BF16 is roughly 218GB of weights and Maverick at FP8 roughly 400GB, before KV cache, which grows with the context you actually use. Because only 17B parameters are active per token, throughput per GPU-hour is strong once memory is provisioned; the cost cliff is provisioning, not serving. **Q: Is the 10 million token context window real?** It is Meta's stated architectural ceiling for Scout, and it is genuinely one of the largest published context lengths for open weights. Whether quality holds for your use case at extreme lengths is a separate empirical question, and KV-cache memory and latency scale with the tokens you actually load regardless of what the ceiling allows. Our advice is unchanged from every long-context model we evaluate: design for the context you validated on your own documents, not the number on the launch slide. **Q: Does BearPlex deploy Llama 4 in client work?** We evaluate it as a standard candidate wherever multimodal document intake meets data-residency requirements, the pattern common across our compliance-heavy platform work. Two things are always in the engagement: a license-compliance checklist (attribution placement, notice files, derivative naming) treated as a deliverable, and a model-swappable interface so the client is never architecturally locked to the license terms if their product strategy changes. --- ### Claude Opus 4.8: Frontier LLM *Publisher: Anthropic · Paper date: 2026.05.28 · Brief date: 2026.07.07 · 10 min read* *Parameters: Undisclosed · License: Proprietary (API)* URL: https://www.bearplex.com/ai/claude-opus-4-8 arXiv: https://www.anthropic.com/news/claude-opus-4-8 **Excerpt**: BearPlex's engineering brief on Claude Opus 4.8: verified $5/$25 pricing, fast mode and cache economics, dynamic workflows, and when it beats GPT-5.5. Most Claude Opus 4.8 coverage is leaderboard news. This brief is the deployment story, because that is what decides whether it belongs in your stack. Released [May 28, 2026](https://www.anthropic.com/news/claude-opus-4-8), just [41 days after Opus 4.7](https://techcrunch.com/2026/05/28/anthropic-releases-opus-4-8-with-new-dynamic-workflow-tool/), it displaced GPT-5.5 at the top of the [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/articles/claude-opus-4-8-analysis-and-benchmarks) on launch day, which means every "which frontier model" evaluation in mid-2026 now has to price this model first. The interesting parts for a technical buyer are not the index points. They are the pricing modifiers, the honesty re-tuning, and what a 41-day release cadence does to your architecture. ## What it actually is Claude Opus 4.8 (API ID `claude-opus-4-8`) is Anthropic's flagship Opus-tier model. Verified specs from the [official model docs](https://platform.claude.com/docs/en/about-claude/models/overview): a **1M-token context window**, **128K max output tokens** (up to 300K on the Batch API behind the `output-300k-2026-03-24` beta header), and a January 2026 knowledge cutoff. It runs adaptive thinking with an `effort` parameter (five levels, `low` through `max`, including `xhigh`) that **defaults to `high` on every surface**. Manual thinking budgets and the classic sampling parameters (`temperature`, `top_p`, `top_k`) were removed from the Opus line starting with 4.7 and return a 400; steering is done through prompts and the effort dial, per Anthropic's [migration guide](https://platform.claude.com/docs/en/about-claude/models/migration-guide). On the third-party scoreboard: Artificial Analysis measured **61.4** on its Intelligence Index, up 4.1 points from Opus 4.7 and 1.2 points ahead of GPT-5.5 at xhigh effort, the previous leader. Anthropic's own framing was unusually modest, calling the release "a modest but tangible improvement," a candor [Simon Willison praised](https://simonwillison.net/2026/May/28/claude-opus-4-8/) in his day-one review. Both things are true: the capability delta over 4.7 is incremental, and it was still enough to take the #1 spot. The behavioral headline matters more than the score. Anthropic states Opus 4.8 is **around four times less likely than its predecessor to allow flaws in code it has written to pass unremarked**, and its lower hallucination rate comes primarily from abstaining on questions it is uncertain about rather than answering more of them correctly. For production agents, that trades silent failures for explicit "I am not sure" outputs, which is the failure mode you actually want, provided your pipeline handles abstention instead of treating any answer as final. ## Commercial terms Hosted API only; there are no weights. Procurement-relevant facts, all verified against [Anthropic's docs](https://platform.claude.com/docs/en/about-claude/models/overview) as of July 2026: - **Multi-cloud availability is real, not aspirational.** Opus 4.8 is live on the Claude API, Claude Platform on AWS, Amazon Bedrock (`anthropic.claude-opus-4-8`), Google Vertex AI (`claude-opus-4-8`), and Microsoft Foundry. If your procurement runs through an existing AWS or GCP commit, this model is reachable without a new vendor contract. - **Model IDs are pinned snapshots.** Since the 4.6 generation, the dateless ID format (`claude-opus-4-8`) is itself a pinned snapshot, not an evergreen pointer. Behavior does not shift under you without a model-ID change on your side. - **Old versions stay serveable, peripherals do not.** Opus 4.5, 4.6, and 4.7 all remain active alongside 4.8; the only Opus retirement on the calendar is the 2025-era Opus 4.1 (August 5, 2026). But peripheral features move fast: Opus 4.7's fast mode is already deprecated and is removed on July 24, 2026, less than two months after 4.8 shipped. Pin models freely; do not pin preview features. ## Real API cost Verified against the [official pricing page](https://platform.claude.com/docs/en/about-claude/pricing) as of July 2026, per million tokens: | Mode | Input | Output | |---|---|---| | Standard | $5.00 | $25.00 | | Batch (50% off) | $2.50 | $12.50 | | Fast mode (research preview) | $10.00 | $50.00 | | Cached input (read) | $0.50 | n/a | Pricing is unchanged from Opus 4.7, and four modifiers change the real bill more than the headline rates: 1. **The 1M context window has no long-context premium.** Anthropic bills a 900K-token request at the same per-token rate as a 9K one. Compare with [gpt-5.5](/ai/gpt-5), where prompts beyond 272K input tokens are billed at 2x input and 1.5x output. For genuinely long-context workloads (repository-scale analysis, large document sets), this is the single biggest line-item difference between the two frontier options: same $5.00 input rate, $25.00 vs $30.00 output, and no context surcharge. 2. **Fast mode is a 2x-price latency lever, not a different model.** It serves the same Opus 4.8 at up to 2.5x output speed for $10/$50, which is 3x cheaper than fast mode was on Opus 4.7 ($30/$150). It is a research preview, unavailable with the Batch API and on Claude Platform on AWS. Use it for latency-critical lanes you have measured, not as a default. 3. **The cache math favors agent loops.** Cache reads bill at $0.50 (10% of input), 5-minute cache writes at $6.25, 1-hour writes at $10. Agents with a stable prompt prefix see most input tokens at the read rate, and Opus 4.8's new mid-conversation system messages (below) exist specifically to keep that prefix intact. 4. **Tokenizer inflation distorts cross-model comparisons.** The tokenizer introduced with Opus 4.7 produces roughly 30% more tokens for the same text than pre-4.7 Claude models. Per-token price comparisons against [Sonnet 4.5](/ai/claude-sonnet-4-5)-era baselines mislead; compare cost per completed task, not per megatoken. On efficiency, Artificial Analysis measured Opus 4.8 completing GDPval-AA tasks with 15% fewer turns and 35% fewer output tokens than Opus 4.7, while across the full index it spent roughly the same output tokens as 4.7 for materially higher scores. It still used about 30% more turns than GPT-5.5 on agentic tasks, so per-task cost between the two is workload-specific: run your own traces before believing either vendor's efficiency story. And budget for effort: Willison reports a single max-effort request costing him 43 cents. ## Eval behavior that matters in production - **Effort is the reliability and cost dial.** The default is `high` everywhere, which is the right call for agent steps but expensive for routine extraction. Sweep `medium`/`high`/`xhigh` on your own evals per route; do not ship the default unexamined. - **The honesty re-tuning changes failure handling.** A model that flags flaws in its own code four times more often (Anthropic's number, vendor-reported) produces more caveats and more abstentions. Pipelines that regex-parse confident answers will see more "unparseable" outputs; pipelines that route uncertainty to review get exactly what they wanted. - **Mid-conversation system messages are the sleeper API feature.** Opus 4.8 accepts `role: "system"` entries inside the messages array, so an agent harness can change instructions mid-task without editing the top-level system prompt and invalidating the prompt cache. Willison called it "really powerful," and for long agent loops the cache savings are structural. - **Dynamic workflows are a Claude Code product feature, not an API primitive.** The research preview lets Claude plan a task and then run parallel subagents in one session, capped at [16 concurrent agents and 1,000 per run](https://www.marktechpost.com/2026/05/28/anthropic-ships-claude-opus-4-8-alongside-dynamic-workflows-and-cheaper-fast-mode-with-workflows-capped-at-1000-subagents/), on Enterprise, Team, and Max plans. Anthropic pitches it at codebase-scale migrations. It is genuinely useful for internal engineering velocity; it is not something to build a product dependency on while it carries the research-preview label. - **Known weak spots exist.** Artificial Analysis notes Claude still trails GPT-5.4 and GPT-5.5 on CritPt, a frontier physics benchmark, even while leading on other scientific reasoning evals. As with every launch, treat vendor benchmark tables as marketing until reproduced on your tasks. ## When to use it, and when not **Use Claude Opus 4.8 when:** - Long-horizon agentic coding and multi-step enterprise workflows are the core workload; this is the exact regime the honesty re-tuning and turn-efficiency gains target. - Your workload actually uses long context. The flat-rate 1M window makes repository-scale and document-corpus work cheaper than the equivalent on GPT-5.5's premium-priced long context. - Procurement or data-residency policy routes through AWS, GCP, or Azure; day-one availability on all three is a real operational advantage. **Do not use it when:** - Data cannot leave your infrastructure at all. No weights exist; that constraint points to [open-weights models](/ai/deepseek-v3), not a different hosted vendor. - The workload is high-volume and simple. At $5/$25 with a default-high effort setting, Opus 4.8 is the wrong tier for classification and extraction volume; route that to a Sonnet- or Haiku-class model and reserve Opus for the steps that measurably need it. - You need a stable feature surface more than peak capability. Fast mode and dynamic workflows are both research previews, and the 4.7 fast-mode removal shows how quickly preview-tier features get retired. ## How we would architect it for a client The same gateway discipline we apply to [every frontier API](/ai/gpt-5), tuned to this vendor's specifics: 1. **A model-agnostic gateway** owns model IDs and effort settings, so the next 41-day Opus release is an eval run plus a config change, not a code change. Anthropic's cadence in 2026 (4.7 in April, 4.8 in May, Fable 5 above it on June 9) makes a standing quarterly re-evaluation the minimum, and the [Anthropic vs OpenAI comparison](/compare/openai-vs-anthropic) gets rerun with it. 2. **Effort-tiered routing**: cheaper Claude tiers or a rival's mini-class models for volume lanes, Opus 4.8 at `high` for standard agent steps, `xhigh` reserved for requests that demonstrably need it, fast mode only behind a measured latency SLO. 3. **Abstention-aware pipelines.** We treat "the model declined to answer" as a first-class output with its own routing (human review, retrieval retry, or a second model), because Opus 4.8 produces more of these by design and that is where its reliability gains live. This is standard in our [model engineering](/services/model-engineering) work. 4. **Cache-first agent harnesses**: stable prompt prefixes, mid-conversation system messages for in-flight instruction changes, and batch pricing for every non-interactive lane ($2.50/$12.50 is frontier capability at commodity rates). Opus 4.8 is the strongest default frontier choice as of July 2026, and the reasons are mostly unglamorous: flat long-context pricing, multi-cloud availability, pinned snapshots, and a failure mode that announces itself. Those are deployment virtues, not launch-day ones, which is why the launch coverage mostly missed them. #### Frequently Asked Questions **Q: What is Claude Opus 4.8 and when was it released?** Claude Opus 4.8 (API ID claude-opus-4-8) is Anthropic's flagship Opus-tier model, released on May 28, 2026, 41 days after Opus 4.7. It has a 1M-token context window, 128K max output tokens, a January 2026 knowledge cutoff, and adaptive thinking controlled by an effort parameter that defaults to high. On launch day it took the #1 position on the Artificial Analysis Intelligence Index with a score of 61.4, displacing GPT-5.5. **Q: How much does Claude Opus 4.8 cost via the API?** As of July 2026, standard pricing is $5.00 per million input tokens and $25.00 per million output tokens, unchanged from Opus 4.7. Batch processing halves that to $2.50/$12.50, cache reads bill at $0.50 per million, and fast mode (a research preview serving the same model at up to 2.5x output speed) costs $10.00/$50.00. Notably, the full 1M context window bills at standard rates with no long-context premium, unlike GPT-5.5, which bills 2x input and 1.5x output once a prompt exceeds 272K input tokens. **Q: Is Claude Opus 4.8 better than GPT-5.5?** On the third-party Artificial Analysis Intelligence Index, yes by a small margin: 61.4 versus GPT-5.5 at xhigh effort, a 1.2-point lead. The fuller picture is workload-specific. Opus 4.8 wins on flat long-context pricing and cheaper output tokens ($25 vs $30 per million), while Artificial Analysis measured it using about 30% more turns than GPT-5.5 on agentic tasks, and it still trails both GPT-5.4 and GPT-5.5 on CritPt, a frontier physics benchmark. We recommend deciding with task-level evals on your own workload, not the index. **Q: What are dynamic workflows in Claude Opus 4.8?** Dynamic workflows is a research-preview feature in Claude Code (Enterprise, Team, and Max plans) where Claude plans a large task and then executes it with parallel subagents in a single session, up to 16 concurrent agents and 1,000 agents per run. Anthropic pitches it at codebase-scale migrations across hundreds of thousands of lines. It is a product feature of Claude Code rather than an API primitive, so we treat it as an engineering-productivity tool, not an architecture dependency, while it remains in preview. **Q: What is fast mode and when is it worth paying for?** Fast mode serves the identical Opus 4.8 model at up to 2.5x output token speed for double the price: $10/$50 per million tokens, which is 3x cheaper than fast mode was on Opus 4.7. It is a research preview, unavailable with the Batch API and on Claude Platform on AWS, and Anthropic has already scheduled the removal of the 4.7 version, so treat it as a tactical lever. It is worth it only for lanes with a measured, user-facing latency SLO; everything else should run standard or batch. **Q: Can Claude Opus 4.8 run on-premises or in our VPC?** No. It is a hosted model with no downloadable weights, and that does not change at any pricing tier. What Anthropic does offer is unusually broad hosted placement: the first-party Claude API, Claude Platform on AWS, Amazon Bedrock, Google Vertex AI, and Microsoft Foundry, plus a US-only inference option (inference_geo) at a 1.1x price multiplier. If your constraint is that data must never leave infrastructure you control, the evaluation shifts to open-weights models like DeepSeek V3 or Llama 4 rather than to a different hosted vendor. **Q: How reliable is Claude Opus 4.8 for production agents?** Its distinguishing trait is calibrated honesty: Anthropic reports it is around four times less likely than Opus 4.7 to let flaws in its own code pass unremarked, and its hallucination reduction comes mainly from abstaining when uncertain. In production, that means fewer silent failures and more explicit uncertainty, which is a net win only if your pipeline routes abstentions somewhere useful. We design agent harnesses around that behavior: schema-validated outputs, abstention-aware routing, and task-level evals per effort setting before committing an architecture. **Q: Does BearPlex build on Claude Opus 4.8?** Yes, where client evals support it, and as of mid-2026 it is a frequent winner for long-horizon agent lanes and long-context analysis. We run it as one candidate inside a model-agnostic gateway: effort-tiered routing, cache-first prompt architecture, batch pricing for non-interactive work, and a quarterly re-evaluation cadence matched to Anthropic's release speed. Whether Opus 4.8 keeps a lane is decided by the client's own task-level evals, not by the leaderboard. --- ### GLM-5.2: Open-Weights LLM *Publisher: Z.ai (Zhipu AI) · Paper date: 2026.06.16 · Brief date: 2026.07.07 · 10 min read* *Parameters: 753B MoE (~40B active) · License: MIT* URL: https://www.bearplex.com/ai/glm-5-2 arXiv: https://huggingface.co/blog/zai-org/glm-52-blog Model card: https://huggingface.co/zai-org/GLM-5.2 **Excerpt**: BearPlex's engineering brief on GLM-5.2: MIT open weights, verified Z.ai pricing, Terminal-Bench and FrontierSWE results, the verbosity tax, and self-hosting. Every few months an open-weights release forces the same board-level question: are we still paying frontier-API prices for work an MIT-licensed model can do? GLM-5.2 is the strongest version of that question yet asked, because for the first time the gap on long-horizon coding benchmarks is inside the noise floor: Z.ai's published FrontierSWE number sits 0.7 points behind Claude Opus 4.8 at output-token prices roughly a sixth of Anthropic's. This brief is the evaluation we would run before believing that, written down. ## What it actually is GLM-5.2 went to Z.ai's Coding Plan subscribers on June 13, 2026, and the full weights landed on Hugging Face under an **MIT license** on June 16 ([announcement](https://huggingface.co/blog/zai-org/glm-52-blog), [model card](https://huggingface.co/zai-org/GLM-5.2)). It is a Mixture-of-Experts transformer with **753B total parameters** and roughly **40B active per token** (Artificial Analysis tallies the total at 744B; both counts agree on the 40B active figure, which is what your serving economics care about). The BF16 weights are about **1.51TB on disk** per [Simon Willison](https://simonwillison.net/2026/jun/17/glm-52/), who called it "probably the most powerful text-only open weights LLM" the day after release. Three architectural details matter for buyers: - **A 1M-token context window**, up from GLM-5.1's 200K. The announcement credits a sparse-attention scheme called IndexShare, where every four attention layers share a lightweight indexer, for cutting per-token FLOPs by a claimed 2.9x at the 1M length. - **Text only.** There is no vision input. More on why that bites in practice below. - **A multi-token-prediction layer** for speculative decoding, with Z.ai claiming up to 20% better acceptance length, which is throughput you get for free on supported inference stacks. ## The license, and the question behind it MIT, full stop. No user thresholds, no attribution badges, no derivative naming rules, and the announcement goes out of its way to say "no regional limits." Against the [Llama 4 Community License](/ai/llama-4) and its conditions, this is the shortest legal conversation in the open-weights market, the same posture as [DeepSeek's MIT line](/ai/deepseek-v3). The question clients actually ask is the China question, and the answer has two halves. **The weights are inspectable files.** Self-hosted in your VPC, there is no vendor in the request path and no telemetry; provenance is a supply-chain review, not a data-flow review. **The first-party API is a different matter**: Z.ai is a Chinese vendor, and if your data-governance posture excludes that, the practical middle path is US-based inference providers. [OpenRouter](https://openrouter.ai/z-ai/glm-5.2) routes GLM-5.2 across multiple hosts (Willison counted nine providers within a day of release, and Interconnects names Fireworks and Together among them), so you can buy the model without buying the vendor's data path. Decide which half of this you are procuring before the pilot, not after. ## The benchmark story, read carefully Z.ai's own published table is unusually specific, so quote it with attribution and the caveat that vendors choose their tables: | Benchmark | GLM-5.2 | Claude Opus 4.8 | GPT-5.5 | |---|---|---|---| | Terminal-Bench 2.1 | 81.0 | 85.0 | 84.0 | | SWE-bench Pro | 62.1 | 69.2 | 58.6 | | FrontierSWE | 74.4% | 75.1% | 72.6% | | MCP-Atlas | 76.8 | 77.8 | 75.3 | Read the second row as carefully as the third. On FrontierSWE (multi-hour, open-ended engineering tasks) GLM-5.2 is effectively tied with Opus 4.8 and ahead of GPT-5.5. On SWE-bench Pro, Opus 4.8 is still **seven points clear**. Both are true at once: this model closed most of the gap, not all of it, and which gap matters depends on your workload. The third-party signal is what elevates this release above the usual launch-table skepticism. [Artificial Analysis](https://artificialanalysis.ai/articles/glm-5-2-is-the-new-leading-open-weights-model-on-the-artificial-analysis-intelligence-index) independently scored it 51 on their Intelligence Index v4.1, the **highest of any open-weights model**, ahead of MiniMax-M3 (44), DeepSeek V4 Pro (44), and Kimi K2.6 (43). [Cline](https://x.com/cline/status/2066951439793242193) called it "the first open-weights model to cross 80% on Terminal-Bench." And Nathan Lambert's [Interconnects write-up](https://www.interconnects.ai/p/glm-52-is-the-step-change-for-open) makes the point benchmarks cannot: "GLM-5.2 is the open weight model that feels right in coding harnesses as a general agent. It's the first one." That practitioner sentence, not the table, is why this brief exists. Our standing advice is unchanged: these numbers earn the model a seat in your evaluation, and nothing more. Run your own task-level evals before moving traffic. ## Real cost, including the verbosity tax Verified against [Z.ai's pricing page](https://docs.z.ai/guides/overview/pricing) as of July 2026: **$1.40 per million input tokens, $0.26 cached input, $4.40 per million output tokens**. Via OpenRouter, a 35% promotional discount had it near $0.91/$2.86 at the time of writing. Against [Anthropic's published Opus 4.8 rates](https://platform.claude.com/docs/en/about-claude/pricing) of $5.00 input and $25.00 output, GLM-5.2's output tokens cost roughly a sixth and input just over a quarter. Now the correction most cost models miss: **GLM-5.2 spends more tokens to do the same work.** Artificial Analysis measured about 43K output tokens per Intelligence Index task, 37K of it reasoning, versus 24K for MiniMax-M3 and 35K for Kimi K2.6, and roughly $0.46 per task versus $0.18 for MiniMax-M3. The per-token sticker overstates the savings and understates the latency: those thinking tokens arrive before your user sees anything. Price the model per completed task on your own workload, never per million tokens. One more commercial tell from the announcement: Z.ai's Coding Plan meters subscription quota by time of day, with peak hours (14:00 to 18:00 UTC+8) drawing 3x quota and off-peak 2x, promotionally 1x through the end of September. Rationing is a capacity signal; demand for this model is real. **Self-hosting** is a datacenter conversation. The BF16 weights are 1.51TB before KV cache, and a 1M-token context makes KV cache its own capacity line. Unsloth's community quantizations put 4-bit builds between 365GB and 467GB on disk depending on variant, and the 2-bit dynamic builds at 238GB to 254GB, the smallest of which fits a 256GB unified-memory Mac at single-digit tokens per second: a proof of ownership, not a production deployment. Production self-hosting means a multi-GPU node in the 8x 96GB-to-141GB class or larger, served through the stacks the model card lists (vLLM, SGLang, transformers, KTransformers, plus Ascend NPU support). Most teams should start on rented endpoints and let measured volume justify the hardware; MIT means that door never closes. ## What breaks in production - **Text-only is an integration hazard, not just a missing feature.** Modern coding harnesses casually attach screenshots; Lambert reported that his harness sending images "would brick Fireworks API for the session." Audit every image path in your agent stack before cutover. - **Verbosity is a budget and latency problem.** 43K output tokens per task compounds across a multi-step agent. Cap thinking budgets per step and measure. - **Hosted-provider variance.** Different hosts serve different precisions and context configurations. Pin one provider, record the precision you evaluated, and re-evaluate on any provider change. - **Ecosystem velocity cuts both ways.** Interconnects flags the regulatory tail risk on Chinese open weights explicitly. Weights you hold cannot be revoked, but a hosted-only dependency on this model deserves the same fallback lane we build for [every proprietary API](/ai/gpt-5). ## When to use it, and when not **Use GLM-5.2 when:** - Your dominant spend is agentic coding or long-horizon terminal work, the exact lanes where its published and third-party results are strongest. - You want frontier-adjacent capability with an exit ramp to [self-hosting](/services/sovereign-cloud) that a hosted-only frontier model can never offer. - Cost per task, measured honestly against your incumbent, shows the verbosity-adjusted savings are real for your workload. **Do not use it when:** - Your agents need vision. There is no GLM-5.2 image input, and no bolt-on fixes that cleanly. - Your hardest tickets live where the SWE-bench Pro gap lives; keep a frontier escalation lane rather than forcing one model to do everything. - Tight interactive latency budgets cannot absorb 30K-plus reasoning tokens per step. ## How we would architect it for a client The engagement shape is a shadow-lane evaluation inside the same model-agnostic gateway we use for [every model engineering](/services/model-engineering) build: 1. **Shadow the incumbent.** Route a copy of real coding-agent traffic to GLM-5.2 on a pinned US-hosted endpoint, score both lanes on task completion, review burden, and cost per completed task over two to four weeks. 2. **Exploit the cache line.** At $0.26 per million cached input tokens, stable prompt prefixes do the same quiet work here as on every frontier API. 3. **Route by difficulty, not by loyalty.** The realistic end state is GLM-5.2 owning volume coding-agent traffic with a proprietary frontier model held for the tail, the same [routing discipline](/ai/deepseek-r1) we apply across vendors. 4. **Keep the self-host option priced.** Because the license is MIT, a move into the client's VPC is a procurement decision, not a re-architecture. We keep the GPU sizing sheet current so the trigger is a spreadsheet threshold, not a research project. GLM-5.2 is the first open-weights model where "replace the frontier API for coding" is a serious engineering question rather than a cost fantasy. Serious questions deserve evals, not vibes. Run them. #### Frequently Asked Questions **Q: Is GLM-5.2 really MIT licensed, with no strings attached?** Yes. The weights on Hugging Face carry a plain MIT license: no monthly-active-user thresholds, no attribution requirements, no naming rules for fine-tuned derivatives, and Z.ai's announcement states there are no regional limits. That makes the legal review materially shorter than for conditionally licensed open models like Llama 4. The usual diligence still applies to what a license cannot cover: model provenance, your own output-liability posture, and the data path of whichever hosted API you use to serve it. **Q: Can GLM-5.2 actually replace Claude Opus 4.8 or GPT-5.5 for coding?** For part of the lane, plausibly; for all of it, not yet on the published evidence. Z.ai's table has it within 0.7 points of Opus 4.8 on FrontierSWE and ahead of GPT-5.5 there and on SWE-bench Pro, but Opus 4.8 stays seven points clear on SWE-bench Pro and four ahead on Terminal-Bench 2.1. Our recommendation is a shadow evaluation on your own repositories: route volume agentic-coding traffic to GLM-5.2, keep a frontier escalation lane for the hardest tasks, and let cost per completed task decide the split. **Q: What does GLM-5.2 cost via API?** Direct from Z.ai, as of July 2026: $1.40 per million input tokens, $0.26 for cached input, and $4.40 per million output tokens, with OpenRouter listing promotional rates around $0.91/$2.86 across multiple hosting providers. Against Anthropic's published Opus 4.8 pricing of $5.00/$25.00, output tokens cost roughly a sixth. The caveat is verbosity: Artificial Analysis measured about 43K output tokens per benchmark task, more than other leading open models, so compare cost per completed task rather than per million tokens. **Q: What hardware does self-hosting GLM-5.2 require?** It is a 753B-parameter MoE whose BF16 weights occupy about 1.51TB before KV cache, so full-precision serving means a large multi-GPU datacenter node, and the 1M-token context window adds significant KV-cache memory on top. Unsloth's community quantizations bring 4-bit builds to 365GB to 467GB on disk depending on variant and 2-bit dynamic builds to 238GB to 254GB, the smallest of which fits a 256GB unified-memory Mac at single-digit tokens per second, useful as a demo, not production. Most teams should serve it via rented GPU endpoints first; the MIT license keeps in-VPC hosting open whenever volume justifies the hardware. **Q: Is using a Chinese model a data risk?** Separate the weights from the API. Self-hosted weights are inspectable files running in your infrastructure with no vendor in the request path, which is a stronger data posture than any hosted API from any country. Z.ai's first-party API does route your prompts to a Chinese vendor, and if your governance posture excludes that, US-based inference providers reachable through OpenRouter, such as Fireworks and Together, serve the same open weights under US jurisdiction. We treat this as a procurement decision made explicitly at the start of an engagement, not a detail discovered during security review. **Q: Does GLM-5.2 support images or other modalities?** No. It is text-only, with no vision input, which Simon Willison flagged in the first day of coverage and which matters more than it sounds for agents: modern coding harnesses routinely attach screenshots, and Interconnects reported that image attachments could break an API session on one hosting provider. If your agent workflows depend on screenshot reading, UI verification, or document images, either keep a multimodal model in the loop for those steps or choose a different primary model. **Q: Does BearPlex deploy GLM-5.2 in client work?** We evaluate it the way this brief describes: as the current strongest open-weights candidate for the agentic-coding lane, run in shadow against the client's incumbent frontier model inside a model-agnostic gateway. The deciding metrics are task completion, review burden, and verbosity-adjusted cost per completed task on the client's own repositories, never launch benchmarks. Where it wins, it takes the volume lane with a frontier escalation path retained; and because the license is MIT, the migration from hosted endpoints to the client's own cloud stays a config change. --- ### MiniMax M3: Open-Weights LLM *Publisher: MiniMax · Paper date: 2026.06.01 · Brief date: 2026.07.07 · 11 min read* *Parameters: 428B MoE (23B active) · License: MiniMax Community License* URL: https://www.bearplex.com/ai/minimax-m3 arXiv: https://www.minimax.io/blog/minimax-m3 Model card: https://huggingface.co/MiniMaxAI/MiniMax-M3 **Excerpt**: BearPlex's engineering brief on MiniMax M3: verified API pricing, the 1M-token context, the Community License revenue gate, and when cost justifies a switch. The June story everyone repeated about MiniMax M3 was the price: launch coverage ([VentureBeat's headline](https://venturebeat.com/technology/minimax-m3-debuts-eclipsing-gpt-5-5-and-gemini-3-1-pro-on-key-benchmark-performance-for-just-5-10-of-the-cost) put it at 5 to 10 percent of the cost of GPT-5.5 and Gemini 3.1 Pro on key benchmarks). The story that actually matters to a buyer is different: under what conditions does that price gap justify a migration, what does the license really permit, and where does the model earn or lose its place in a production stack. That is this brief. ## What it actually is MiniMax M3 ([announced June 1, 2026](https://www.minimax.io/blog/minimax-m3)) is Shanghai-based MiniMax's frontier open-weights release: a Mixture-of-Experts model with roughly **428B total parameters and 23B active per token** per the [Hugging Face model card](https://huggingface.co/MiniMaxAI/MiniMax-M3), a **1M-token context window**, and native multimodality (text, image, and video input, trained mixed-modality from the first step rather than through a bolted-on encoder). MiniMax also demonstrates it operating a desktop computer for agentic tasks. The architectural headline is **MiniMax Sparse Attention (MSA)**, a new sparse-attention design. MiniMax's stated numbers: per-token compute at 1M context is about **1/20th of its previous generation**, with prefill speedups of **more than 9x** and decode speedups of **more than 15x** at the 1M length. That efficiency claim is what makes the pricing below structurally plausible rather than a loss-leader mystery. MiniMax promised weights within ten days of the announcement and delivered inside the first week: the weights are live on Hugging Face (June 7 per launch coverage) under a license tagged `minimax-community`, with official serving support in [vLLM](https://recipes.vllm.ai/MiniMaxAI/MiniMax-M3) and SGLang. Third-party hosting followed immediately: [Fireworks announced day-0 support](https://fireworks.ai/blog/minimax-m3-launch) (initially at 500K context, with the full 1M to follow), and the model is listed on OpenRouter. So the deployment menu is real: MiniMax's own API, a Western inference host, or your own GPUs. ## The license: open weights with a revenue gate "Open weights" is doing careful work in that sentence. The [MiniMax Community License](https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE) is not Apache 2.0 or MIT, and three clauses shape the commercial decision: **1. Non-commercial use is genuinely free.** Research, evaluation, and personal use carry broad rights: use, copy, modify, distribute, sublicense. **2. Commercial use has a two-tier gate.** If your annual revenue is **under $20M**, commercial use requires a one-time notice email to api@minimax.io. If your revenue is **at or above $20M**, the license requires **separate, prior written authorization** from MiniMax before commercial use. That second tier is the clause procurement needs to see early: it is not a formality you self-certify, it is a permission you request, with whatever timeline and terms MiniMax attaches. **3. Attribution travels with commercial deployments.** Commercial users must prominently display **"Built with MiniMax M3"** on a related website, UI, blog post, about page, or product documentation. Fine-tuned and post-trained derivatives used commercially trigger the same notification and authorization requirements. The practical read: for a startup or mid-market product team, the license costs an email and a badge. For an enterprise above the revenue line that wants to self-host, the license is a negotiation, and you should not commit architecture to it before the authorization exists in writing. Note also the split that matters: this license governs the **weights**. Consuming MiniMax's hosted API is governed by their platform terms, like any other API vendor. If the license friction is disqualifying but the economics are attractive, that usually resolves to using a hosted endpoint, or to the permissively licensed alternatives ([Apache 2.0 Qwen 3](/ai/qwen-3), [MIT DeepSeek V3](/ai/deepseek-v3)) accepting a capability tradeoff you measure yourself. ## Real API cost Verified against MiniMax's [official pay-as-you-go pricing](https://platform.minimax.io/docs/guides/pricing-paygo) as of July 2026, per million tokens: | Tier | Input | Output | Cache read | |---|---|---|---| | Standard, up to 512K input | $0.30 | $1.20 | $0.06 | | Standard, above 512K input | $0.60 | $2.40 | $0.12 | | Priority, up to 512K input | $0.45 | $1.80 | $0.09 | | Priority, above 512K input | $0.90 | $3.60 | $0.18 | Three things to internalize before putting these numbers in a business case: 1. **The headline rate is a discount MiniMax controls.** The pricing page labels the standard rates "Permanent 50% off", which means the list price is $0.60/$2.40 and the billed price today is $0.30/$1.20. "Permanent" is a pricing-page word, not a contract term. Model the list price as your risk case; the economics below survive it. 2. **Long context is a doubled tier, not free.** The 1M window exists, but requests beyond 512K input tokens bill at 2x. Same conclusion we reach on every long-context model: retrieval discipline still pays. 3. **Subscriptions exist for coding-agent workloads.** MiniMax sells token plans at [$20/month (~1.7B tokens), $50 (~5.1B), and $120 (~9.8B)](https://www.minimax.io/blog/minimax-m3), aimed at its MiniMax Code agent. For individual-developer and small-team coding use these are the cheapest way in; production systems should stay on metered API pricing you can forecast. Now the switch math, against [OpenAI's verified pricing](https://developers.openai.com/api/docs/pricing) for gpt-5.5 ($5.00 input / $30.00 output per million): take a workload of 10B input and 1B output tokens per month. On gpt-5.5 that is $80,000/month. On M3 at the billed rate it is $4,200/month, about 5 percent. Even at M3's undiscounted list price it is $8,400, about 10 percent. OpenAI's batch tier (50% off) and $0.50 cached-input rate narrow the gap for batch-heavy, prefix-stable workloads, and M3's own $0.06 cache-read rate widens it again. There is no realistic modifier stack that closes a 20x headline gap; the question is never whether M3 is cheaper, it is whether the migration cost and the risk profile are worth the delta on your volume. ## Self-hosting cost The [vLLM recipe](https://recipes.vllm.ai/MiniMaxAI/MiniMax-M3) sizes it honestly: vLLM 0.24.0+, **8x H200 or H20 for a tight single-node BF16 fit** at tensor parallel 8, with multi-node TP for long-context headroom. Working from the parameter count, 428B at BF16 is roughly 856GB of weights before KV cache, which is why the single-node fit needs 141GB-class GPUs. An MXFP8 quantization is published alongside the BF16 weights for smaller footprints, and MSA brings deployment quirks of its own: `--block-size 128` is mandatory for the sparse-attention index cache, and an fp8 KV cache buys roughly 1.5x KV pool capacity for long-context serving. The strategic point: at $0.30/$1.20 on the hosted side, self-hosting M3 almost never wins on cost alone. An 8x H200 node is only cheaper than the API at sustained high utilization, and most teams overestimate their utilization. Self-hosting M3 is a **data-governance decision**: it is what you do when the workload cannot leave your infrastructure and you still want frontier-adjacent capability with a 1M window. That it is merely affordable, rather than profitable, is fine; that is what the option is for. ## Eval behavior that matters - **Vendor benchmarks are vendor benchmarks.** MiniMax's [launch numbers](https://www.minimax.io/blog/minimax-m3): 59.0% on SWE-Bench Pro, 66.0% on Terminal-Bench 2.1, 74.2% on MCP Atlas, with the model card adding 80.5 on SWE-bench Verified, 78.1 on MMMU Pro, and 85.4 on Video-MME v2. MiniMax's own comparisons place M3 ahead of GPT-5.5 and Gemini 3.1 Pro on SWE-Bench Pro and behind Claude Opus 4.7, per [The Decoder's coverage](https://the-decoder.com/minimax-m3-open-weight-model-with-a-million-token-context-challenges-proprietary-leaders/), which also notes these are internal tests. Treat them as marketing until reproduced on your tasks. - **Independent signal exists and is strong.** [Artificial Analysis](https://artificialanalysis.ai/models/minimax-m3) scores M3 at 44 on its Intelligence Index, ranking it #2 of the 93 open-weights models it tracks, measuring 95.3 output tokens/second and 1.87s time to first token, and notes it is comparatively concise for a reasoning model in token usage. - **It is a reasoning model, and output tokens are where reasoning lives.** At $1.20/M output the thinking is cheap, but latency budgets still need to account for it; this is not the model for a 200ms autocomplete path. - **The proven lanes are agentic.** The launch evidence, and the benchmark selection itself (SWE-Bench, Terminal-Bench, MCP Atlas), point at agentic coding, tool-heavy agents, and long-horizon autonomous tasks as the lanes MiniMax optimized for. Multimodal long-context intake (video plus documents in one window) is the differentiated capability nothing at this price matches on paper; it is also the least independently validated, so eval it first. ## When you actually switch for cost The decision framework we use with clients, stated plainly: **Switch, or add M3 as a routing lane, when all of these hold:** - Model spend is a real budget line. If your monthly token bill is a rounding error, the engineering time for migration evals exceeds years of savings; do nothing. - Your workload sits in M3's proven lanes: agentic coding, tool-calling agents, long-context repo and document work, or multimodal intake at volume. - You have your own eval suite, so switching is a golden-set run plus a canary period, not a research project. - Governance clears one of the two paths: your data can go to MiniMax's hosted API (a Chinese vendor's platform), or a Western host like Fireworks serving the open weights satisfies your requirements, or you self-host. **Do not switch when:** - You are at or above $20M revenue and the plan involves self-hosted weights without MiniMax's written authorization in hand. Get the authorization first or stay on hosted endpoints. - Your compliance posture cannot accept the available hosting paths for the data in question. - Your architecture leans on a specific vendor's first-party tool surface (hosted shell, computer use, file search). M3 speaks standard tool calling and MCP, but platform tools do not migrate. - The workload is latency-critical interactive chat where reasoning tokens hurt more than the price helps. ## How we would architect it for a client The same gateway discipline we apply to [every model lane](/services/model-engineering), with two M3-specific additions: 1. **M3 enters as a cost lane, not a replacement.** A model-agnostic gateway routes the high-volume agentic and long-context work to M3 while the incumbent keeps the lanes it wins. Golden-set evals decide the split, rerun quarterly; both the "permanent" discount and the benchmark story get re-verified on that cadence, because both are MiniMax's to change. 2. **License compliance as a deliverable.** The engagement checklist covers the tier determination against the $20M line, the notice email or written authorization, the "Built with MiniMax M3" attribution placement, and the same review for any fine-tuned derivative before it ships. 3. **A rollback lane by construction.** Because the price advantage is the whole thesis, we keep the interface swappable and the previous lane warm. If pricing, terms, or hosted-API behavior shift, reverting is a config change plus an eval run, the same posture we take with [US frontier vendors' deprecation velocity](/ai/gpt-5). M3 is the strongest version yet of a question the market keeps asking: what is the last 10 percent of capability worth to your specific workload? For a growing share of production token volume, the honest answer in mid-2026 is: not twentyfold. Run your evals and find out which share is yours. #### Frequently Asked Questions **Q: How much does MiniMax M3 cost via the API?** As of July 2026, MiniMax's official pay-as-you-go pricing bills M3 at $0.30 per million input tokens and $1.20 per million output tokens for requests up to 512K input tokens, with cache reads at $0.06. Above 512K input tokens the rates double to $0.60/$2.40. A priority tier costs 1.5x for faster admission. Note the standard rates are labeled a 'permanent 50% off' discount from a $0.60/$2.40 list price, so model the list price as your risk case. Subscription token plans start at $20/month for roughly 1.7B tokens. **Q: Is MiniMax M3 really 10 to 20 times cheaper than GPT-5.5?** On headline rates, yes. Against OpenAI's verified gpt-5.5 pricing of $5.00 input and $30.00 output per million tokens, M3's billed rate of $0.30/$1.20 is roughly 6 percent on input and 4 percent on output. OpenAI's batch tier and cached-input pricing narrow the gap for specific workload shapes, and M3's own cache-read rate widens it again. The real question is not whether M3 is cheaper but whether your monthly volume makes the savings worth the migration and evaluation work, and whether the capability holds on your tasks. **Q: Can we use MiniMax M3 commercially?** Yes, with conditions that depend on your size. The MiniMax Community License makes non-commercial use free. Commercial use requires prominently displaying 'Built with MiniMax M3' and, if your annual revenue is under $20M, a one-time notice email to api@minimax.io. At or above $20M annual revenue, the license requires separate prior written authorization from MiniMax before commercial use, and commercially deployed fine-tunes trigger the same requirements. Consuming MiniMax's hosted API is governed by their platform terms rather than the weights license. **Q: What hardware does self-hosting MiniMax M3 need?** The official vLLM recipe calls for vLLM 0.24.0+ and 8x H200 or H20 GPUs for a tight single-node BF16 fit at tensor parallel 8, with multi-node setups for long-context headroom. Working from the 428B parameter count, BF16 weights are roughly 856GB before KV cache; an MXFP8 quantization is published for smaller footprints. The sparse-attention architecture requires a mandatory block size of 128, and an fp8 KV cache buys about 1.5x cache capacity. At M3's API prices, self-hosting is a data-governance decision, not a cost play. **Q: Is the 1M-token context window real?** It is the published window on both the model card and the API, backed by an architectural argument: MiniMax Sparse Attention, which MiniMax says cuts per-token compute at 1M context to about a twentieth of its previous generation, with prefill and decode speedups it puts at more than 9x and 15x. Two caveats: API requests beyond 512K input tokens bill at double rates, and long-context quality at the extremes is exactly the kind of claim you validate on your own documents before designing a product around it. Fireworks initially served it at 500K context at launch. **Q: What about data residency, given MiniMax is a Chinese company?** This is the constraint to resolve before any pricing conversation. Using MiniMax's first-party API means your prompts and data transit the platform of a Shanghai-based vendor, which some compliance postures cannot accept. The open weights create two alternatives: Western inference hosts onboarded M3 at launch (Fireworks announced day-0 support, and it is listed on OpenRouter), and self-hosting inside your own VPC removes the vendor from the data path entirely. We treat the hosting path as a governance decision made per data classification, not per model. **Q: How does MiniMax M3 compare to other open-weights models?** On independent measurement it is at or near the top: Artificial Analysis ranks it #2 of the 93 open-weights models it tracks on its Intelligence Index, and it is the only one at that tier combining a 1M-token window with native image and video input. The tradeoffs versus alternatives are licensing and scale: Qwen 3 is Apache 2.0 and DeepSeek V3 is MIT, with no revenue gates or attribution requirements, and both families offer smaller deployment footprints. M3's vendor benchmarks against US frontier models are internal MiniMax tests; validate on your own workload. **Q: Does BearPlex deploy MiniMax M3 in client work?** We evaluate it as a cost lane behind a model-agnostic gateway: high-volume agentic coding, tool-calling, and long-context workloads route to M3 where the client's own golden-set evals support it, while incumbent models keep the lanes they win. The engagement includes a license-compliance checklist (revenue-tier determination, notice or authorization, attribution placement) and a warm rollback lane, so a pricing or terms change from the vendor is a configuration change rather than an incident. The discount label and benchmark claims get re-verified quarterly. --- ### Gemma 4: Open-Weights LLM *Publisher: Google · Paper date: 2026.04.02 · Brief date: 2026.07.07 · 10 min read* *Parameters: E2B / E4B / 26B MoE (3.8B active) / 31B dense · License: Apache 2.0* URL: https://www.bearplex.com/ai/gemma-4 arXiv: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ Model card: https://huggingface.co/google/gemma-4-31B-it **Excerpt**: BearPlex's engineering brief on Gemma 4: what Apache 2.0 changes for client products, verified memory numbers for every size, and when local beats a hosted API. Every AI platform we architect eventually hits the same question: what do we run when the model cannot be an API call? Data that must stay in the building, apps that must work offline, per-request economics that a hosted frontier model cannot survive. For 2026, Gemma 4 is our default first answer to that question, and this brief is about why, what it actually costs to run, and where that answer stops being right. ## What it actually is Gemma 4 ([Google, April 2, 2026](https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/)) is Google's open-weights family built from the same research line as Gemini 3, and the first Gemma generation released under a plain **Apache 2.0 license**. Four models launched: - **E2B and E4B**: edge models that activate an effective 2 billion and 4 billion parameter footprint, built for phones, browsers, and embedded hardware. **128K context window**, and the only launch variants with **native audio input**. - **26B A4B**: a Mixture-of-Experts model, roughly 26B total parameters with only **3.8B active per token**. Context up to **256K** (262,144 tokens). - **31B dense**: the flagship. All parameters active, 256K context, and the size Google says fits on a **single 80GB NVIDIA H100**. All four process text and images at variable resolutions, and Google states all models handle **video input** as well. The family is trained on **over 140 languages**, ships with configurable thinking modes, and has native support for function calling, structured JSON output, and system instructions per the [model card](https://huggingface.co/google/gemma-4-31B-it). The family has kept moving since launch, per the [official release log](https://ai.google.dev/gemma/docs/releases): Multi-Token Prediction (MTP) variants of all four launch models landed on April 16, 2026, and a fifth model, **Gemma 4 12B Unified**, arrived June 3, 2026 with text, image, and audio input at 256K context, closing the audio gap in the mid-range. ## The license is the headline Previous Gemma generations shipped under Google's bespoke Gemma Terms of Use. Gemma 4 ships under Apache 2.0, and for client work that is not a footnote, it is the differentiator: - **No user-count gates, no attribution badges, no derivative naming rules.** Compare the [Llama 4 Community License](/ai/llama-4), where "Built with Llama" attribution and the derivative naming clause are product-shaping obligations. Apache 2.0 has none of that: fine-tune it, white-label it, ship it under your own brand. - **An explicit patent grant.** Standard Apache 2.0 machinery, and exactly the clause enterprise counsel asks about first. - **Legal review shrinks from a memo to a checkbox.** Our healthcare and fintech clients run procurement processes where "bespoke AI license" triggers weeks of review. Apache 2.0 is already on every approved-license list on earth. The practical effect: in evaluations where Gemma 4 and a conditionally licensed model are within noise of each other on task quality, the license decides it, before benchmarks are discussed. ## Real deployment cost Google publishes an official memory table in the [Gemma 4 docs](https://ai.google.dev/gemma/docs/core), and it is the real sizing document. Weights only, before KV cache: | Model | BF16 | SFP8 | Q4_0 | |---|---|---|---| | E2B | 11.4GB | 5.7GB | 2.9GB | | E4B | 17.9GB | 8.9GB | 4.5GB | | 12B | 26.7GB | 13.4GB | 6.7GB | | 26B A4B | 57.7GB | 28.8GB | 14.4GB | | 31B | 69.9GB | 34.9GB | 17.5GB | Read that table carefully, because the model names mislead in both directions: - **The "effective" models are bigger on disk than their names suggest.** E2B needs 11.4GB at BF16. The 2B refers to active compute, not weights. On-device deployments get it far smaller: Google's [edge stack](https://developers.googleblog.com/bring-state-of-the-art-agentic-skills-to-the-edge-with-gemma-4/) runs E2B in **under 1.5GB** on supported devices using LiteRT's 2-bit and 4-bit weights. - **The MoE saves compute, not memory.** The 26B A4B holds nearly as much memory as the 31B dense while activating 3.8B parameters per token. You provision for 26B and pay per-token compute like a 4B. If your bottleneck is VRAM rather than throughput, the MoE buys you nothing. - **Quantization-aware training is the production default.** Google ships QAT checkpoints (q4_0 GGUFs for the llama.cpp ecosystem, plus unquantized QAT weights for custom pipelines) trained to hold quality at 4-bit, which is a materially better starting point than post-hoc quantization. Still eval the delta on your own tasks. What this maps to in hardware: the 31B at Q4 fits a 24GB consumer GPU or a 24GB Apple Silicon Mac; community measurements on [Apple hardware](https://sudoall.com/gemma-4-31b-apple-silicon-local-guide/) put the Q4 31B at roughly 40 to 50 tokens/second on an M4 Max and 15 to 25 on a 24GB M2/M3 Pro. At the bottom of the ladder, Google's own edge numbers: E2B on a Raspberry Pi 5 CPU does 133 prefill / 7.6 decode tokens per second, and 3,700 prefill / 31 decode on a Qualcomm Dragonwing IQ8 NPU. That spread, one family from a Pi to an H100, is the operational argument for Gemma 4: one license, one chat template, one eval harness across every deployment tier. For managed deployment, [Google Cloud](https://cloud.google.com/blog/products/ai-machine-learning/gemma-4-available-on-google-cloud) runs the family on Vertex AI Model Garden, Cloud Run GPUs (RTX PRO 6000 Blackwell, 96GB, scale-to-zero), and GKE with vLLM, with TPU serving via vLLM TPU announced alongside. The option that matters for regulated clients: Google has committed Gemma 4 across its sovereign offerings, up to air-gapped on-premises Google Distributed Cloud. ## The Arena story, read honestly Google's launch claim is that the **31B dense ranks #3 among all open models on the Arena text leaderboard, with the 26B MoE at #6**, both ahead of models with many times their parameter count; the [DeepMind model page](https://deepmind.google/models/gemma/gemma-4/) lists the 31B thinking variant at an Arena score of 1452. The rank is real and the achievement is real: a 31B dense model in that neighborhood changes what "small" means. Three caveats before you repeat it in a business case. Arena measures human preference on chat, not your workload. The margins over the adjacent [Qwen models](/ai/qwen-3) are thin enough that community reads of the same data saw a lead a single leaderboard update could erase. And leaderboards move monthly; the ranks quoted here are the April 2026 launch snapshot. Our standing advice is unchanged: vendor benchmarks shortlist models, your own task evals pick them. ## Launch-week reality, and behavior that matters Gemma 4's first weeks were rougher than the launch post, and the failure modes are instructive ([community postmortem](https://letsdatascience.com/blog/google-gemma-4-open-source-apache-community-found-catches)): - **Day-one ecosystem lag.** The new architecture broke Hugging Face Transformers and PEFT support at release, and teams with fine-tuning plans waited on upstream fixes. By July 2026 the ecosystem has caught up (official QAT GGUFs, llama.cpp, Ollama, vLLM, Unsloth), but the lesson generalizes: for any new open-weights family, pin runtime versions and budget slack between release day and production day. - **MoE throughput lagged in early local runtimes.** One early community test measured the 26B MoE at 11 tokens/second where a comparable Qwen model did 60+ on identical hardware. Runtime-level, not architectural, and improving, but it is why we benchmark the actual runtime you will ship, not the architecture diagram. - **Audio is edge-only at the top.** The 26B and 31B do not take audio input; E2B, E4B, and the June 12B Unified do. A voice product on the flagship needs a separate transcription stage. - **Video input has practical limits.** Local runtimes process short clips at low frame rates (roughly a minute at about 1 frame/second per [Unsloth's deployment docs](https://unsloth.ai/docs/models/gemma-4)). Treat it as sampled-frames understanding, not long-form video analysis. - **Mind the sampling defaults and thinking hygiene.** Recommended defaults are temperature 1.0, top_p 0.95, top_k 64, and in multi-turn use you keep only final answers in history, never prior thought blocks. Small details, real quality deltas. ## When to use it, and when not **Use Gemma 4 when:** - Data cannot leave the device or the VPC and you want one Apache 2.0 family covering every tier, phone-class intake to server-side reasoning, under a license procurement will not fight. - The product is genuinely edge: field apps that work offline, kiosks, NPU-equipped hardware, on-device inference where E2B/E4B latency and footprint are the feature. - You are replacing a mid-tier hosted model on cost: a Q4 31B on a 24GB GPU serves a large class of extraction, summarization, and agent workloads with zero per-token vendor spend. - White-label or fine-tune-and-rebrand plans make [conditionally licensed alternatives](/ai/llama-4) legally awkward. **Do not use it when:** - The workload needs frontier-grade reasoning. The family tops out at 31B dense; the hardest problems still belong to a [hosted frontier model](/ai/gpt-5) or a much larger open model, ideally behind a router so only those requests pay for it. - You need audio input at the large sizes, or long-form video understanding, today. - Serving throughput is the whole game and your stack's MoE support is immature; measure your runtime before committing to the 26B A4B. - Your context requirements genuinely exceed 256K. ## How we would architect it for a client The fit we see most clearly is the compliance-heavy field-workforce pattern, the same shape as the NDIS provider-management platform we build as a long-term development partner ([Vertex360](/case-studies/vertex360)): mobile workers capturing forms, photos, and notes in environments with unreliable connectivity, feeding a regulated back office. 1. **E4B on the device** for offline intake: classification, structured extraction to JSON via the native constrained output support, running in the app's own process so nothing leaves the handset unprocessed. 2. **31B (QAT q4_0) in the client's [sovereign cloud](/services/sovereign-cloud)** for the heavy work: cross-document summarization, compliance narrative drafting, agentic workflows over the case record. One GPU class, predictable capex, no per-token bill. 3. **One eval harness across both tiers.** Same task suite, run against E4B, 31B, and the incumbent hosted model, so every routing decision is a measured tradeoff. This is standard [model engineering](/services/model-engineering) discipline, not Gemma-specific. 4. **A swappable model interface anyway.** Apache 2.0 removes the license reasons to leave, not the technical ones. The open-weights field re-ranks every quarter, and the gateway keeps the exit cost at a config change. Gemma 4 is not the best model in the world. It is something more useful: the best-licensed, best-laddered family for the deployments where the model has to live where the data lives. For that slot, in mid-2026, it is the first name on our shortlist. #### Frequently Asked Questions **Q: Can we use Gemma 4 in a commercial product?** Yes, and with less friction than most open-weights alternatives. Gemma 4 is the first Gemma generation released under plain Apache 2.0, which means no user-count thresholds, no attribution requirements, no derivative naming rules, and an explicit patent grant. You can fine-tune it and ship the result under your own brand. That contrasts directly with the Llama 4 Community License (mandatory 'Built with Llama' attribution, naming rules on distributed fine-tunes) and with earlier Gemma generations' bespoke Gemma Terms of Use. **Q: What hardware does each Gemma 4 size need?** Google's published memory table (weights only): E2B 2.9GB at Q4_0 (11.4GB BF16), E4B 4.5GB (17.9GB), 12B 6.7GB (26.7GB), 26B A4B 14.4GB (57.7GB), 31B 17.5GB (69.9GB). In practice: E2B runs on phones and a Raspberry Pi 5, and under 1.5GB on supported devices with LiteRT 2-bit and 4-bit weights; E4B is comfortable on a 16GB laptop; the Q4 31B fits a 24GB consumer GPU or 24GB Mac; and the BF16 31B fits a single 80GB H100 per Google. KV cache comes on top and grows with the context you actually use. **Q: Which Gemma 4 model should we pick?** Start from the deployment constraint, not the benchmark. On-device or browser: E2B or E4B (also the only launch models with native audio input). Single consumer GPU or a capable workstation: the 31B at Q4, which is the strongest quality per dollar in the family. High-throughput server workloads where VRAM is plentiful but per-token compute matters: the 26B A4B MoE. The 12B Unified (added June 2026) covers the mid-range and adds audio. Then confirm with your own task evals; the right answer is workload-specific. **Q: Does the 26B MoE need less memory than the 31B dense?** No, and this is the most common Gemma 4 sizing mistake. Mixture-of-Experts reduces compute per token, not weights in memory: the 26B A4B loads roughly 26B parameters (57.7GB BF16, 14.4GB Q4_0 per Google's table) while activating only 3.8B per token. You provision memory for the full model and get throughput economics closer to a 4B. If your constraint is VRAM, the dense 31B at Q4 is nearly the same footprint with all parameters active, which is why many local deployments prefer it. **Q: Is Gemma 4 really better than models 20x its size?** On the specific measure Google cites, yes at launch: the 31B dense ranked #3 among open models on the Arena text leaderboard (the DeepMind page lists a 1452 score for the thinking variant) and the 26B MoE ranked #6, ahead of much larger models. The honest caveats: Arena measures human preference on chat rather than your workload, the margins over adjacent Qwen releases were thin enough that community reads saw no decisive winner, and leaderboard ranks move monthly. Treat it as evidence the 31B is in the top open-model tier, then run task-level evals. **Q: Can Gemma 4 run on phones and edge devices?** Yes, that is the family's strongest card. E2B and E4B are built for edge deployment: Google's stack runs E2B in under 1.5GB of memory on supported devices using LiteRT 2-bit and 4-bit weights, with published figures of 133 prefill / 7.6 decode tokens per second on a Raspberry Pi 5 CPU and 3,700 prefill / 31 decode on a Qualcomm Dragonwing IQ8 NPU. Google also announced Gemma 4 as the foundation for the next generation of Gemini Nano on Android. Both edge models take image, video, and audio input with a 128K context window. **Q: What are Gemma 4's real weaknesses?** Four we flag in evaluations. The family tops out at 31B dense, so frontier-hard reasoning still routes to a hosted frontier model or a larger open model. Audio input is missing on the 26B and 31B (only E2B, E4B, and the later 12B Unified have it). Launch-week ecosystem support lagged, with Transformers and PEFT breakage and slow MoE throughput in early local runtimes; that has largely resolved, but it argues for pinning runtime versions. And video input in practice means short, low-frame-rate clips, not long-form video analysis. **Q: Does BearPlex deploy Gemma 4 in client work?** It is our default first candidate for the local and edge slot: workloads where data residency, offline operation, or per-token economics rule out a hosted API. The pattern we deploy is tiered, an edge model (E4B) for on-device intake and the Q4 31B in the client's own cloud for heavy work, with one eval harness across tiers and a model-swappable gateway so the choice stays reversible. The Apache 2.0 license is a genuine factor in regulated engagements because it removes the bespoke-license review from procurement. --- ### MAI-Code-1-Flash: Coding LLM *Publisher: Microsoft · Paper date: 2026.06.02 · Brief date: 2026.07.07 · 10 min read* *Parameters: 137B total (5B active) · License: Proprietary (Copilot service terms)* URL: https://www.bearplex.com/ai/mai-code-1-flash arXiv: https://microsoft.ai/news/introducingmai-code-1-flash/ Model card: https://microsoft.ai/pdf/MAI-Code-1-Flash-Model-Card.PDF **Excerpt**: BearPlex's engineering brief on MAI-Code-1-Flash: verified specs (137B MoE, 256K context), Copilot pricing at $0.75/$4.50, and when to flip the org policy. Most coverage of MAI-Code-1-Flash is a Build 2026 keynote recap. This brief is about the decision that actually landed on engineering leaders' desks on June 26, 2026, when the model went [generally available for Copilot Business and Copilot Enterprise](https://github.blog/changelog/2026-06-26-mai-code-1-flash-for-copilot-business-and-copilot-enterprise/): whether to enable the org policy and let this model carry part of your Copilot fleet's workload. That framing matters because, unlike almost every other model we brief, you cannot deploy MAI-Code-1-Flash anywhere else. There are no weights and, as of this writing, no generally available standalone API. The product surface is the decision surface. ## What it actually is MAI-Code-1-Flash was announced on [June 2, 2026 at Microsoft Build](https://microsoft.ai/news/introducingmai-code-1-flash/) as the coding member of a seven-model in-house MAI family (alongside MAI-Thinking-1, the image, voice, and transcription lines). Per the official [model card](https://microsoft.ai/pdf/MAI-Code-1-Flash-Model-Card.PDF): - **Architecture**: transformer with sparse Mixture-of-Experts layers, **137B total parameters, 5B active** per token. - **Context**: **256K tokens**. Text in, text out only. - **Training window**: March to May 2026, with a **December 2025 knowledge cutoff**. - **Lineage**: built from **MAI-Thinking-1's mid-training checkpoint**, then taken through supervised fine-tuning, a "mid2" phase of roughly 2 million synthetic agentic tasks, and a final large-scale reinforcement learning stage spanning more than 150,000 environments. - **Supported language**: English, per the model card. The strategic subtext is the headline in most coverage: Microsoft states the model was trained end to end by Microsoft "on clean, traceable and enterprise-grade data, without distillation from third-party models," part of an explicit push toward what the company calls long-term self-sufficiency in models. For a company whose flagship developer product has run on OpenAI models since launch, with Anthropic models added to the picker later, shipping its own coding model into that product is a supply-chain move as much as a research result. GitHub's own changelog calls it "the first in a new wave of purpose-built coding models from Microsoft," and the Flash suffix telegraphs that a bigger sibling is coming. The genuinely novel engineering claim is the training environment: the model was **trained directly with the GitHub Copilot harness used in production**, meaning the file-editing tools, terminal integrations, and multi-step task loops it saw in RL are the same ones that serve users. Microsoft's evaluations ran in that same harness. Read both directions of that fact. It is a real reason to expect benchmark transfer inside Copilot to be better than typical vendor numbers, and it is also a reason the numbers tell you little about the model anywhere else, which is currently nowhere. ## The benchmark table, and how to read a vendor-run comparison All published comparisons are Microsoft's own, run against exactly one competitor, [Claude Haiku 4.5](/ai/claude-sonnet-4-5), with the Haiku numbers footnoted as coming from Microsoft's internal benchmark system. From the model card: | Benchmark | MAI-Code-1-Flash | avg tokens | Claude Haiku 4.5 | avg tokens | |---|---|---|---|---| | SWE-Bench Verified | 71.6% | 10.8K | 66.6% | 27.3K | | SWE-Bench Pro | 51.2% | 28.0K | 35.2% | 29.8K | | SWE-Bench Multilingual | 65.5% | 15.3K | 62.7% | 17.2K | | Terminal Bench 2 | 54.8% | 21.6K | 41.6% | 25.0K | | IF Bench (instruction following) | 75.0 | | 46.1 | | | Artifacts Bench (visual coding) | 36.4 | 12.0K | 36.6 | 23.6K | Three honest readings. First, the token column is the real story: Microsoft's headline claim of solving harder problems with **up to 60% fewer tokens** is visible in the SWE-Bench Verified row (10.8K vs 27.3K), and under token-metered billing that is a cost claim, not a bragging right. The model card credits "adaptive solution length control," spending reasoning budget on hard problems and staying terse on easy ones. Second, the comparison set is conspicuously narrow. [Hacker News practitioners](https://news.ycombinator.com/item?id=48374466) noted immediately that Microsoft benchmarked only against Haiku, not against Qwen, Gemini Flash, or the strong open coding models, and questioned the size-to-score ratio against much smaller open competitors. Third, the one row Microsoft published where it does not win, visual coding, aligns with the model's text-only design. Our standing advice applies with extra force when every number is vendor-run in the vendor's own harness: treat this table as a hypothesis to test on your repositories, not a result. ## Commercial terms: a policy toggle, not a procurement There is no license to review because there is nothing to take delivery of. The model card lists the license as the product and service terms of wherever the model is deployed. Concretely, as of this brief: - **Copilot is the deployment.** Rollout went: VS Code model picker for individual plans on [June 2](https://github.blog/changelog/2026-06-02-mai-code-1-flash-is-now-available-for-github-copilot/); Copilot CLI, cloud agent, the Copilot app, Copilot Chat on github.com, Visual Studio, JetBrains, Eclipse, Xcode, and GitHub Mobile on [June 18](https://github.blog/changelog/2026-06-18-mai-code-1-flash-available-on-more-copilot-surfaces/); Business and Enterprise GA on June 26, gated behind an **admin-enabled policy** in Copilot settings. - **No standalone API today.** Microsoft's [MAI family announcement](https://microsoft.ai/news/building-a-hillclimbing-machine-launching-seven-new-mai-models/) says the models are "going to be widely available" on OpenRouter, Fireworks, and Baseten, alongside distribution on Foundry. Future tense. When we checked on July 7, 2026, the OpenRouter listing was not live. If your architecture needs this model outside Copilot, the honest status is "announced, not shipped," and the model card commits only to updating documentation if an API release happens. - **The Copilot auto picker can route to it.** Even before you deliberately select it, the model card notes that Copilot may route tasks to MAI-Code-1-Flash through the Auto picker. Fleet admins should assume some traffic lands on it once the policy is on. ## Real cost under Copilot's new billing The timing of this launch is not an accident. On **June 1, 2026**, GitHub [moved Copilot from premium request units to usage-based billing](https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/): AI credits (1 credit = $0.01) consumed by actual input, cached, and output tokens at per-model list rates. Plans keep their seat prices with included credits: Pro $10/month with $10 in credits, Business $19/user with $19 (promotional $30 through August), Enterprise $39/user with $39 (promotional $70 through August). Overage is billed at list. Verified against [GitHub's models-and-pricing page](https://docs.github.com/en/copilot/reference/copilot-billing/models-and-pricing) as of July 2026, per million tokens: | Model | Input | Cached input | Output | |---|---|---|---| | MAI-Code-1-Flash | $0.75 | $0.075 | $4.50 | | Claude Haiku 4.5 | $1.00 | $0.10 | $5.00 | | GPT-5.4 mini | $0.75 | $0.075 | $4.50 | | Gemini 3 Flash | $0.50 | $0.05 | $3.00 | | Claude Sonnet 4.6 | $3.00 | $0.30 | $15.00 | | GPT-5.5 | $5.00 | $0.50 | $30.00 | Two observations that should anchor the business case. The list-price edge over Haiku 4.5 is modest: 25% on input, 10% on output, and Gemini 3 Flash undercuts both. The pricing table alone does not justify a switch. The claim that would justify it is token efficiency: if the roughly 60% reduction in tokens per completed hard task holds on your workloads, the effective cost per task, which is the only unit that matters, drops far more than the list-price gap suggests, and your per-seat included credits stretch across proportionally more agent runs. That "if" is the entire evaluation. Measure tokens per completed task, not price per million. ## Behavior that matters in production - **Harness-native agentic behavior.** Trained and evaluated in the production Copilot harness, the model's strongest published deltas are exactly the agentic ones: Terminal Bench 2 (54.8 vs 41.6) and instruction following (75.0 vs 46.1, vendor-run). For high-volume, multi-step Copilot agent loops, this is the profile you want. - **Latency is the design center.** GitHub positions it for "high-volume, iterative agentic coding workflows where speed and efficiency matter most." A 5B-active MoE is cheap to serve, and low latency plus low serving cost are the goals the model card leads with. - **Text-only, English-only.** No screenshots, no design-to-code from images, and the model card lists English as the supported language. Mixed-language teams and multimodal workflows should look at [Gemini 3.5 Flash](/ai/gemini-3-5-flash) or frontier options in the same picker. - **The small-model failure mode still applies.** Practitioner skepticism in the HN thread matches our experience with every model in this class: on genuinely hard problems, cheap models can cost you senior-engineer review time that dwarfs the token savings. The escalation path matters more than the default. ## When to use it, and when not **Enable and route to it when:** - You run Copilot Business or Enterprise and your credit burn is dominated by high-volume agentic work: repo Q&A, refactors, test generation, CLI tasks, iterative agent loops. - Your own golden-set evals confirm the tokens-per-task advantage on your repositories, making it the cheapest adequate lane in the picker. - Model provenance is a procurement factor: a single-vendor stack with Microsoft-owned training data lineage simplifies some enterprise reviews. **Do not build around it when:** - The workload lives outside Copilot. There is no GA API; wait for the Foundry and OpenRouter listings to be real before designing against them. - The work is multimodal (screenshots, designs, diagrams) or substantially non-English. - The task class is hard architecture and planning work, where the published gap to frontier models in the same picker ([GPT-5.5](/ai/gpt-5), Claude Sonnet 4.6) is worth paying for, and where a failed cheap attempt costs more than the delta. - You need vendor portability. A model that exists only inside one product is the opposite of a portable dependency. ## How we would run the evaluation for a client This is a fleet-routing decision under metered billing, the same discipline we apply in our [model engineering](/services/model-engineering) work, pointed at Copilot's picker instead of an API gateway: 1. **Pilot behind the policy, not fleet-wide.** Enable the MAI-Code-1-Flash policy for one org or team. Note that the auto picker may start routing traffic to it immediately; that is part of what you are measuring. 2. **Golden-set evals on your own repos.** Define your top task classes (bug fix, refactor, test scaffold, repo Q&A, agent runs) and measure pass rate and tokens per completed task per model. The vendor's 60% token claim is your primary hypothesis to confirm or reject. 3. **Route by task class, escalate deliberately.** The winning pattern we see across [agentic deployments](/services/autonomous-agents) is a cheap default lane with an explicit escalation lane, and review discipline on everything the cheap lane merges. 4. **Re-run quarterly.** GitHub's model catalog and prices moved three times in June alone. The Flash suffix, the announced-but-unshipped API story, and the promotional credits all say this landscape is mid-shift; whatever you decide in July 2026 is a snapshot, not a settlement. MAI-Code-1-Flash is a credible first coding model and an unmistakable strategic statement. Whether it belongs in your fleet is a two-week measurement exercise, not a keynote takeaway. #### Frequently Asked Questions **Q: What is MAI-Code-1-Flash?** MAI-Code-1-Flash is Microsoft's first in-house coding model, announced at Build on June 2, 2026 as part of a seven-model MAI family. Per the official model card it is a sparse Mixture-of-Experts transformer with 137B total parameters and 5B active per token, a 256K context window, text-only input and output, and a December 2025 knowledge cutoff. It was trained from MAI-Thinking-1's mid-training checkpoint and reinforcement-learned directly inside the GitHub Copilot production harness, and it ships exclusively as a model option inside GitHub Copilot. **Q: Was MAI-Code-1-Flash really built without OpenAI?** Microsoft states the model was trained end to end by Microsoft on curated data 'without distillation from third-party models,' using its own infrastructure and data pipelines, and frames the MAI family as part of a long-term self-sufficiency strategy. GitHub's changelog calls it 'the first in a new wave of purpose-built coding models from Microsoft.' The training disclosure in the model card describes the full pipeline: pretraining, mid-training, supervised fine-tuning, roughly 2 million synthetic agentic tasks, and large-scale RL across more than 150,000 environments. **Q: How much does MAI-Code-1-Flash cost?** On GitHub Copilot's published price list as of July 2026: $0.75 per million input tokens, $0.075 cached input, $4.50 output. That is 25% below Claude Haiku 4.5's input rate and 10% below its output rate on the same list, and identical to GPT-5.4 mini. Under Copilot's usage-based billing (live June 1, 2026), those tokens consume AI credits at 1 credit = $0.01, against included credits per seat. The bigger cost lever is Microsoft's claim of up to 60% fewer tokens on hard coding tasks; if it holds on your workloads, cost per completed task drops much further than the list prices suggest. **Q: Is MAI-Code-1-Flash better than Claude Haiku 4.5?** On Microsoft's published numbers, yes across coding and instruction-following benchmarks: 71.6 vs 66.6 on SWE-Bench Verified, 51.2 vs 35.2 on SWE-Bench Pro, 54.8 vs 41.6 on Terminal Bench 2, while using substantially fewer tokens per task. But every published comparison is vendor-run, in Microsoft's own harness, against exactly one competitor, and the Haiku figures are footnoted as coming from Microsoft's internal benchmark system. Practitioner threads flagged the narrow comparison set immediately. Treat the table as a hypothesis and verify pass rates and tokens per task on your own repositories. **Q: Can we use MAI-Code-1-Flash outside GitHub Copilot, via API?** Not as of July 2026. There are no downloadable weights, and the model card describes distribution only through GitHub Copilot surfaces, with an API release framed as a possible future format. Microsoft's MAI family announcement says the models are 'going to be widely available' on OpenRouter, Fireworks, and Baseten alongside Foundry, but the OpenRouter listing was not live when we checked on July 7, 2026. If your architecture needs this model outside Copilot, wait until an endpoint actually exists before designing against it. **Q: Should we enable MAI-Code-1-Flash for our Copilot Business or Enterprise fleet?** Run it as a measured pilot, not a fleet-wide flip. The model went GA for Business and Enterprise on June 26, 2026 behind an admin policy in Copilot settings, and once enabled the auto picker may route tasks to it on its own. Enable it for one team, measure pass rate and tokens per completed task against your current default on your own repositories, and confirm the token-efficiency claim before routing volume. The economics are attractive under usage-based billing; whether the quality holds on your codebase is the empirical question the vendor numbers cannot answer for you. **Q: Does BearPlex work with MAI-Code-1-Flash?** We evaluate it the way we evaluate any fleet-routing option: golden-set evals per task class on the client's own repositories, tokens-per-completed-task as the primary cost metric, and an explicit escalation lane to frontier models for hard planning work. Because the model currently exists only inside GitHub Copilot, engagements center on Copilot fleet policy, routing, and measurement rather than deployment architecture. We rerun the evaluation quarterly; June 2026 alone saw three catalog changes, and the Flash naming signals more family members coming. --- ### DeepSeek R1: Open-Weights LLM *Publisher: DeepSeek-AI · Paper date: 2025.01.22 · Brief date: 2026.07.03 · 10 min read* *Parameters: 671B MoE (37B active) + distills 1.5B-70B · License: MIT · Base: DeepSeek-V3-Base* URL: https://www.bearplex.com/ai/deepseek-r1 arXiv: https://arxiv.org/abs/2501.12948 Model card: https://huggingface.co/deepseek-ai/DeepSeek-R1 GitHub: https://github.com/deepseek-ai/DeepSeek-R1 **Excerpt**: BearPlex's engineering brief on DeepSeek R1: what the MIT license really permits, self-hosting cost per quant level, latency behavior, and when to choose it. Most model briefs about DeepSeek R1 were written in the week it broke the app-store charts. This one is written for a different moment: the model is eighteen months old, the hype has moved on, and the engineering question has gotten sharper. R1 is no longer the model you evaluate because it is in the news. It is the model you evaluate because it is the most permissively licensed serious reasoning model you can put on your own hardware. That framing changes what matters. So this brief covers what a build team actually needs: the license, the real deployment cost, the latency behavior, and where R1 wins or loses against the alternatives in mid-2026. ## What it actually is DeepSeek R1 ([DeepSeek-AI, January 2025](https://arxiv.org/abs/2501.12948), later published in *Nature*) is a Mixture-of-Experts reasoning model: **671B total parameters with 37B activated per token** and a **128K context window**, per the [official model card](https://huggingface.co/deepseek-ai/DeepSeek-R1). The research contribution was showing that large-scale reinforcement learning, without human-annotated reasoning traces, is enough to make long chain-of-thought behavior emerge: self-checking, backtracking, strategy switching. Two release details matter more for engineering than the headline model: 1. **The R1-0528 refresh.** The May 2025 update ([model card](https://huggingface.co/deepseek-ai/DeepSeek-R1-0528)) is the checkpoint you should be evaluating, not the January original. Per DeepSeek's own published numbers, AIME 2025 accuracy went from 70% to 87.5%, LiveCodeBench from 63.5% to 73.3%, hallucination rate dropped, function calling improved, and system prompts became properly supported. 2. **The distill ladder.** DeepSeek released six distilled models (Qwen-based 1.5B, 7B, 14B, 32B and Llama-based 8B, 70B) that inherit R1's reasoning style at hardware budgets normal companies have. The Qwen-based distills carry Apache 2.0 from their base models. ## The license, and what commercial use really permits The weights are released under the **[MIT License](https://huggingface.co/deepseek-ai/DeepSeek-R1)**, and the model card is explicit that this "supports commercial use" and allows "any modifications and derivative works, including, but not limited to, distillation and fine-tuning." For client products, MIT is about as clean as model licensing gets. Concretely, compared to the community licenses attached to [Llama 4](/ai/llama-4): - No monthly-active-user threshold and no revenue gate. - No attribution badge required in your product UI. - No naming requirements on fine-tuned derivatives. - Distilling R1's outputs into your own smaller model is expressly permitted, which is exactly the clause that makes R1 attractive as a teacher model. One nuance worth counsel's five minutes: the *Llama-based distills* are not pure MIT. They are derived from Llama 3.1 and 3.3 checkpoints and remain subject to those Llama license terms. If license simplicity is the point, use the Qwen-based distills. The other conversation that comes up in every regulated-industry evaluation: DeepSeek is a Chinese lab, and some procurement teams stop there. The engineering answer is that with self-hosted open weights there is no telemetry and no data path to the vendor at all. The weights are inspectable files running in your VPC. That is a materially different risk posture from sending prompts to any hosted API, DeepSeek's or anyone else's. ## Real deployment cost As of July 2026, the access landscape has shifted in a way most older writeups miss: DeepSeek's own [API platform](https://api-docs.deepseek.com/quick_start/pricing) has moved to the V4 generation, and the legacy `deepseek-reasoner` endpoint that served the R1 line is scheduled for deprecation on July 24, 2026. First-party per-token access to R1 is effectively ending. What remains is what the MIT license always made possible: running the weights yourself, or renting them from a GPU cloud. Self-hosting footprints, derived from the published parameter counts (weights only, before KV cache): - **Full R1 at native FP8**: 671B parameters at one byte per parameter is roughly 670GB of weights. That is a multi-GPU node in the 8x 96GB-141GB class. This is a serious infrastructure commitment, not a pilot-project footprint. - **Full R1 at ~4-bit quantization**: roughly 340GB, which still means several datacenter-class GPUs, and reasoning quality at aggressive quantization needs eval before you commit. - **R1-Distill-Qwen-32B at BF16**: roughly 64GB, a single H100/A100-80GB with headroom, or two 48GB cards. - **R1-Distill-Qwen-7B / Llama-8B at 4-bit**: workstation and even laptop territory. The honest cost conclusion: for most mid-market deployments the full 671B model is the wrong first choice. The 32B distill inside your own cloud, promoted to the full model only if evals prove the gap matters, is the pattern that survives contact with a budget. ## Latency and eval behavior that matters R1 spends tokens to think, and you pay for that in both latency and output-token cost. DeepSeek's own 0528 notes show average reasoning depth on AIME questions rising from roughly 12K to 23K tokens per question. Product implications: - **Time-to-first-answer is a product decision.** A visible answer can arrive minutes after the request on hard prompts. If the workload is interactive chat, R1-class reasoning is the wrong default path; route only the hard cases to it. - **Budget output tokens, not requests.** Cost models that assume a few hundred output tokens per call will be wrong by an order of magnitude on reasoning workloads. - **Decoding settings are not optional.** The model card recommends temperature in the 0.5 to 0.7 range (0.6 with top-p 0.95 for 0528); greedy decoding degrades output and invites repetition loops. - **Benchmarks are self-reported.** The numbers above are DeepSeek's published figures on the model cards. Treat them as directional and run your own task-level evals; that is the standard we apply in every [model engineering](/services/model-engineering) engagement. ## When to use it, and when not **Use DeepSeek R1 when:** - You need strong multi-step reasoning (math-adjacent logic, code synthesis, structured analysis) inside your own infrastructure boundary, with data that cannot leave. - You want a teacher model for distillation and the license must permit it. - Latency tolerance is measured in tens of seconds and correctness is worth the wait: batch analysis, agent planning steps, review pipelines. **Do not use it when:** - The workload is fast interactive chat or high-volume extraction. A non-reasoning model, or a small distill, is cheaper and quicker. - You have no GPU story and were relying on DeepSeek's own API for the long term; the first-party R1 endpoint is going away as of July 2026, so plan around self-hosting or a GPU cloud from day one. - Your compliance regime requires a vendor to stand behind the model contractually. MIT weights come with no warranty and no counterparty. ## How we would architect it for a client The pattern we reach for mirrors the privilege-aware architecture we published in our [SaulLM-7B brief](/ai/saullm-7b) and built for [Letti AI](/case-studies/letti-ai): the open model runs inside the client's VPC as a specialized capability, never as the whole system. Concretely, for a regulated-data reasoning workload: 1. **A routing gateway** classifies each request. High-volume, low-difficulty traffic goes to a small model (an R1 distill or a [Qwen 3](/ai/qwen-3) mid-size); only genuinely hard reasoning requests hit the large model. This single decision usually dominates the economics. 2. **R1-Distill-Qwen-32B as the sovereign workhorse** on a single 80GB-class GPU in the client's [sovereign cloud](/services/sovereign-cloud) footprint, with the full R1 reserved for an offline eval track until the quality gap on the client's own tasks justifies the cluster. 3. **Reasoning-trace hygiene.** R1 emits its chain of thought; we treat those traces as sensitive intermediate data (they restate the input), so they are logged under the same access controls as the source documents and never surfaced raw to end users. 4. **Deterministic verification after generation**, the same discipline as citation checking in legal work: schema validation, unit-testable claims checked in code, and a bounded re-prompt loop. Reasoning models reduce error rates; they do not remove the need for verification. That architecture is why the license section of this brief matters more than the benchmark section. MIT is what makes every one of those four decisions available to you at all. #### Frequently Asked Questions **Q: Is DeepSeek R1 free for commercial use?** Yes. The weights are released under the MIT License, and the official Hugging Face model card states explicitly that the series supports commercial use, modification, derivative works, distillation, and fine-tuning. There is no user threshold, no attribution requirement, and no naming rule. One caveat: the Llama-based distills (8B and 70B) inherit Llama license terms from their base models, so use the Qwen-based distills if you want a fully permissive stack. **Q: Can DeepSeek R1 run on-prem?** Yes, and on-prem or private-VPC is now the primary way to run it, since DeepSeek's first-party API has moved to its V4 models (the legacy reasoner endpoint is deprecated as of July 2026). The full 671B model needs roughly 670GB of weights memory at native FP8, which means a multi-GPU datacenter node. The distilled models are the practical on-prem ladder: the 32B distill fits a single 80GB GPU at BF16, and the 7B/8B distills run on workstation hardware at 4-bit. **Q: Which DeepSeek R1 variant should we actually deploy?** Start with R1-Distill-Qwen-32B unless evals prove otherwise. It fits on one 80GB-class GPU, carries Apache 2.0 from its Qwen base, and inherits most of the reasoning behavior that matters for typical enterprise workloads. Reserve the full 671B model for an offline evaluation track, and promote it only if the measured quality gap on your own tasks justifies an 8-GPU cluster. If you do run the full model, use the R1-0528 checkpoint, not the January original. **Q: Does using DeepSeek R1 send data to China?** Not if you self-host. Open weights are static files; there is no telemetry, no phone-home path, and no vendor in the request loop. Inference happens entirely inside your infrastructure boundary, which is a stronger data-residency posture than any hosted API. The China question is real for DeepSeek's hosted services, where their terms apply, but it does not apply to the MIT weights running in your own VPC. **Q: Why is DeepSeek R1 slow, and can we fix that?** It is slow by design: R1 spends thousands of reasoning tokens before answering, and DeepSeek's own 0528 notes show roughly 23K tokens of reasoning per AIME-class question. You manage this architecturally, not by tuning: route only genuinely hard requests to R1, stream intermediate status to users, cap reasoning budgets per request class, and serve everything else from a smaller, faster model. Treat output-token spend, not request count, as the unit of cost. **Q: How does DeepSeek R1 compare to closed reasoning models like OpenAI's o-series?** The closed frontier models still hold an edge on some hard reasoning tasks, and they come with a vendor, an SLA, and zero infrastructure burden. R1's case is different: it is the reasoning model you can own. If your constraint is data residency, distillation rights, cost control at scale, or independence from a vendor roadmap, R1 wins on the constraint that actually decides the project. We run this as a task-level eval in every engagement rather than trusting leaderboards, ours or anyone's. **Q: Does BearPlex deploy DeepSeek R1 in client work?** We evaluate it as a standard candidate in model engineering engagements wherever reasoning-heavy workloads meet data-residency constraints, using the routing-gateway pattern described in this brief: a small model for volume traffic, an R1 distill in the client's cloud for hard cases, and deterministic verification after generation. Whether it ships depends on the client's eval results and infrastructure budget, not on the leaderboard of the month. --- ### GPT-5: Frontier LLM *Publisher: OpenAI · Paper date: 2025.08.07 · Brief date: 2026.07.03 · 10 min read* *Parameters: Undisclosed · License: Proprietary (API)* URL: https://www.bearplex.com/ai/gpt-5 arXiv: https://arxiv.org/abs/2601.03267 **Excerpt**: BearPlex's engineering brief on the GPT-5 platform: verified gpt-5.5 pricing tiers, routing and reasoning-effort behavior, tool-use in agents, and deprecation risk. Most GPT-5 coverage is a benchmark story. This brief is a procurement and architecture story, because that is what actually decides whether GPT-5 belongs in your stack. As of July 2026, "GPT-5" is not one model. It is a fast-moving platform: the original August 2025 release has already been deprecated, four point-releases have shipped and been retired behind it, and the current flagship is **gpt-5.5**. If you are evaluating OpenAI for a production system, the deprecation section of this brief matters at least as much as the capability section. ## What it actually is GPT-5 launched on August 7, 2025 (the retired API snapshot ID, `gpt-5-2025-08-07`, carries the date). The [GPT-5 System Card](https://arxiv.org/abs/2601.03267) described the consumer product as a unified system: a fast model for most questions, a deeper reasoning model for harder problems, and "a real-time router that quickly decides which model to use based on conversation type, complexity, tool needs, and explicit intent." The router was trained on live signals, including when users switched models and measured correctness, and OpenAI stated the plan was to eventually integrate everything into a single model. That integration is essentially what the 5.x line delivered. The current flagship, [gpt-5.5](https://developers.openai.com/api/docs/models/gpt-5.5) (snapshot `gpt-5.5-2026-04-23`), exposes the routing decision to you as a parameter: `reasoning_effort` accepts **none, low, medium (default), high, and xhigh**. There is no hidden router in the API path. You choose, per request, how much thinking you buy. Specs that matter: a **1,050,000-token context window**, **128,000 max output tokens**, and a knowledge cutoff of December 1, 2025. The engineering consequence of the router history is worth stating plainly: in the API, routing is your job. The consumer product made "which model, how much reasoning" an invisible platform decision; the API hands it back to you as the single biggest cost and latency lever you control. ## Commercial terms There is no license to negotiate in the open-weights sense; this is a hosted API under OpenAI's business terms. What you are actually agreeing to, structurally, is a dependency on OpenAI's model lifecycle. Their published policy commits to **at least 6 months of notice** before retiring generally available models and at least 3 months for specialized variants, per the [deprecations page](https://developers.openai.com/api/docs/deprecations). Six months is the contractual floor you should plan around, because it is the number OpenAI plans around. ## Real API cost Verified against the [official pricing page](https://developers.openai.com/api/docs/pricing) as of July 2026, per million tokens: | Model | Input | Cached input | Output | |---|---|---|---| | gpt-5.5 | $5.00 | $0.50 | $30.00 | | gpt-5.5-pro | $30.00 | n/a | $180.00 | | gpt-5.4 | $2.50 | $0.25 | $15.00 | | gpt-5.4-mini | $0.75 | $0.075 | $4.50 | | gpt-5.4-nano | $0.20 | $0.02 | $1.25 | Three modifiers change the real bill more than the headline rates: 1. **Batch is a flat 50% off** (gpt-5.5 at $2.50/$15.00). Any workload that tolerates asynchronous turnaround should be there by default. 2. **Priority processing costs roughly 2.5x** (gpt-5.5 at $12.50/$75.00). If your latency SLO forces the priority tier, your model is 2.5x more expensive than the number in your business case. 3. **Long context is a premium product.** Prompts beyond 272K input tokens are billed at 2x input and 1.5x output on gpt-5.5. The million-token window exists, but the economics push you toward retrieval and context discipline rather than stuffing the window. The cached-input rate (10% of base) is the quiet workhorse: agent systems that keep a stable prompt prefix routinely see the majority of their input tokens billed at that rate. ## Tool use and eval behavior that matters gpt-5.5's tool surface is broad and first-party: web search, file search, code interpreter, hosted shell, apply-patch, computer use, MCP, and structured outputs are all supported per the [model page](https://developers.openai.com/api/docs/models/gpt-5.5). For production agents, the details we test for in every [model engineering](/services/model-engineering) evaluation: - **Reasoning effort is the reliability dial, not just a cost dial.** Tool-selection and argument-construction errors drop as effort rises, and `none` turns the model into a fast, cheaper executor that is only appropriate for well-constrained tool schemas. Run your own task-level evals per effort level; the deltas are workload-specific and OpenAI publishes no per-effort reliability numbers. - **Structured outputs are the contract.** Schema-constrained generation is the single most effective control we know for keeping multi-step agents parseable at step forty. - **The platform tools create soft lock-in.** Hosted shell, file search, and computer use are excellent, and every one you adopt makes the eventual portability conversation harder. Wrap them behind your own interfaces from day one. ## Deprecation history as platform risk This is the section buyers skip and regret. The verified record from OpenAI's own [deprecations page](https://developers.openai.com/api/docs/deprecations), as of July 2026: - The **original GPT-5** (`gpt-5-2025-08-07`), plus its mini and nano variants, was deprecated on June 11, 2026 and **shuts down December 11, 2026**. Sixteen months from launch to shutdown. - **gpt-5-chat-latest, gpt-5-codex, and the entire 5.1 codex family** were deprecated April 22, 2026 and shut down **July 23, 2026**. - **gpt-5.2-chat-latest and gpt-5.3-chat-latest** shut down August 10, 2026. The 5.2 and 5.3 generations lived well under a year. - Even the GPT-4 era finally ends: `gpt-4-0613`, `gpt-4-turbo`, and the original `gpt-4o` snapshot shut down October 23, 2026. The pattern is consistent: OpenAI ships fast and retires fast, honoring the 6-month floor and rarely much more. Practical implications: pin snapshots, budget a **re-evaluation cycle roughly every two quarters**, keep your eval suite runnable on demand so a forced migration is a regression test rather than a research project, and never hardcode a model ID deeper than one config file. Prompt behavior shifts between point releases; the deprecation notice is also a behavior-change notice. ## When to use it, and when not **Use the GPT-5 platform when:** - You want the broadest first-party tool ecosystem (hosted shell, computer use, MCP) with one vendor and one bill. - Your workload spans wildly different difficulty levels and you can exploit `reasoning_effort` plus the 5.4 mini/nano ladder to route cost. - Batch-eligible volume work dominates: $2.50/$15.00 for frontier-class capability is genuinely hard to beat. **Do not use it when:** - Data cannot leave your infrastructure boundary. There are no weights; an [open-weights model](/ai/deepseek-v3) is the answer to that constraint, not a different API vendor. - Your organization cannot absorb a forced model migration every 12 to 18 months. The deprecation record above is the base rate, not a worst case. - You need multi-year behavioral stability for a regulated, validated workflow; pinned snapshots still retire. ## How we would architect it for a client The same routing-gateway discipline we apply to [open-weights deployments](/ai/deepseek-r1) applies here, with vendor risk added to the design inputs: 1. **A model-agnostic gateway** owns model IDs, so a deprecation is a config change plus an eval run, not a code change. Every prompt lives in version control with a golden-set eval attached. 2. **Effort-tiered routing**: gpt-5.4-nano or mini at low effort for extraction and classification volume, gpt-5.5 at medium for standard agent steps, high or xhigh reserved for the requests that measurably need it. The [comparison with Anthropic's lineup](/compare/openai-vs-anthropic) is rerun quarterly, because both vendors move. 3. **Caching and batch by default**: stable prompt prefixes to exploit the $0.50 cached rate, and every non-interactive pipeline on the batch tier. 4. **A deprecation playbook**, not a deprecation panic: subscribe to the deprecations feed, hold the previous snapshot and the new one in A/B during migration windows, and treat the 6-month notice as the start of a scheduled project. GPT-5 is a strong default choice in mid-2026. It stays a strong choice only for teams that engineer for the platform's velocity instead of pretending it is infrastructure. #### Frequently Asked Questions **Q: What is the current GPT-5 model in the API?** As of July 2026, the flagship is gpt-5.5 (snapshot gpt-5.5-2026-04-23), with a 1,050,000-token context window, 128K max output tokens, and a December 2025 knowledge cutoff. The workhorse tier is the gpt-5.4 family (standard, mini, nano). The original gpt-5 from August 2025 was deprecated in June 2026 and shuts down on December 11, 2026, so new builds should not target it. **Q: How much does GPT-5.5 cost via the API?** As of July 2026, gpt-5.5 is $5.00 per million input tokens, $0.50 for cached input, and $30.00 per million output tokens. Batch processing halves that ($2.50/$15.00), priority processing raises it to $12.50/$75.00, and prompts beyond 272K input tokens are billed at 2x input and 1.5x output. gpt-5.4-mini ($0.75/$4.50) and gpt-5.4-nano ($0.20/$1.25) cover the volume tiers. **Q: Does GPT-5 still use a router?** The router was a ChatGPT product feature, not an API feature. At launch, OpenAI's system card described a real-time router choosing between fast and reasoning models based on conversation type, complexity, and tool needs. In the API, that decision is yours: gpt-5.5 exposes a reasoning_effort parameter with five levels (none, low, medium, high, xhigh), which is the main cost, latency, and reliability lever in any production design. **Q: How reliable is GPT-5 for tool calling in production agents?** gpt-5.5 supports structured outputs, MCP, hosted shell, computer use, and the rest of OpenAI's first-party tool surface, and in our engagements it is one of the strongest tool-calling platforms available. But OpenAI publishes no per-effort-level reliability numbers, and tool-call accuracy varies with reasoning effort and schema design. We treat it as an empirical question: run task-level evals at each effort level on your own tools before committing an architecture. **Q: How big a risk is OpenAI's deprecation policy?** It is the defining operational characteristic of the platform. OpenAI guarantees at least 6 months of notice for GA models, and the record shows they use roughly that: the original GPT-5 goes from launch to shutdown in 16 months, and the 5.2 and 5.3 generations lived under a year. Plan for a forced migration every 12 to 18 months, pin snapshots, keep golden-set evals runnable on demand, and route all model access through a gateway so a migration is configuration, not code. **Q: Can GPT-5 run on-premises or in our VPC?** No. GPT-5 is API-only; there are no weights to host, and that does not change with any pricing tier. If your constraint is data residency or infrastructure control, the decision is not GPT-5 versus another hosted frontier model, it is hosted-API versus open weights. That is the evaluation where models like DeepSeek V3 or Mistral Large 3 enter, and we run it as a routing decision rather than an either-or in most client architectures. **Q: Does BearPlex build on GPT-5?** Yes, where the evaluation supports it. We treat OpenAI as one candidate lane in a model-agnostic gateway: effort-tiered routing across the 5.4 and 5.5 families, cached prefixes and batch pricing exploited by default, and a standing deprecation playbook so vendor velocity never becomes a production incident. Whether GPT-5 wins a given lane is decided by the client's own task evals, rerun quarterly. --- ### Claude Sonnet 4.5: Frontier LLM *Publisher: Anthropic · Paper date: 2025.09.29 · Brief date: 2026.07.03 · 10 min read* *Parameters: Undisclosed · License: Proprietary (API)* URL: https://www.bearplex.com/ai/claude-sonnet-4-5 arXiv: https://www.anthropic.com/news/claude-sonnet-4-5 **Excerpt**: BearPlex's engineering brief on Claude Sonnet 4.5: verified $3/$15 pricing, 200K context economics, agent behavior, and the Sonnet 4.6 / Sonnet 5 migration math. Claude Sonnet 4.5 is the model a large share of today's production agents were built on, and that is exactly why it deserves a sober brief in mid-2026. It now sits in Anthropic's legacy tier: still fully available, still not deprecated, but with two successors above it. The engineering question is no longer "is Sonnet 4.5 good," it is "do we pin it, or do we take the migration." This brief gives you the verified numbers for both sides of that decision. ## What it actually is Claude Sonnet 4.5 shipped on [September 29, 2025](https://www.anthropic.com/news/claude-sonnet-4-5), and Anthropic's launch positioning was unusually direct: "the best coding model in the world" and "the strongest model for building complex agents." The self-reported launch numbers were **77.2% on SWE-bench Verified** and **61.4% on OSWorld** (up from Sonnet 4's 42.2% four months earlier), and Anthropic reported observing it "maintaining focus for more than 30 hours on complex, multi-step tasks." Treat all of those as vendor-reported; the durable claim they support is the design intent, which the market then validated: Sonnet 4.5 became the default engine for long-horizon agent loops. The current, verified spec sheet from the [official model docs](https://platform.claude.com/docs/en/about-claude/models/overview), as of July 2026: API ID `claude-sonnet-4-5-20250929` (a pinned snapshot), **200K-token context window**, **64K max output tokens**, extended thinking supported, reliable knowledge cutoff of January 2025 (training data through July 2025). It is available on the Claude API, Amazon Bedrock, and Google Cloud, and the 4.5 generation is where Bedrock's global-versus-regional endpoint split begins. One under-appreciated production feature: **context awareness**. Per the [context-window docs](https://platform.claude.com/docs/en/build-with-claude/context-windows), the API injects Sonnet 4.5's remaining token budget into the system prompt and updates it after tool calls. The model paces long tasks against real remaining capacity instead of guessing, which is one concrete reason its long-horizon agent behavior holds up. ## Commercial terms Hosted API under Anthropic's commercial terms; there are no weights. The platform-risk picture is the mirror image of [OpenAI's](/ai/gpt-5): Anthropic's published record moves slower. As of July 2026, Sonnet 4.5 is a legacy model but **not deprecated**; in the current lineup only Opus 4.1 carries a deprecation notice. Model IDs are pinned snapshots, so behavior does not shift under you. What you are pricing in is not sudden retirement but gradual staleness: a January 2025 knowledge cutoff gets more expensive to compensate for every quarter. ## Real API cost Verified against the [official pricing page](https://platform.claude.com/docs/en/about-claude/pricing) as of July 2026, per million tokens: - **Base**: $3 input / $15 output. - **Prompt caching**: 5-minute cache writes $3.75 (1.25x), 1-hour writes $6 (2x), cache reads $0.30 (0.1x). - **Batch API**: $1.50 / $7.50 (a flat 50% off). - **Cloud endpoints**: regional or multi-region routing on Bedrock and Google Cloud adds a 10% premium over global endpoints, a 4.5-generation change worth catching in cloud cost reviews. The cache-read rate is the whole economics of agent loops at this tier. A production agent re-sends its system prompt, tool definitions, and history on every step; with a stable prefix, the bulk of that input bills at $0.30 instead of $3. In our [model engineering](/services/model-engineering) work, getting cache discipline right on a Sonnet-class agent routinely moves input cost by high double-digit percentages, which is more than most model-swap decisions move it. ## Context economics, and the successor math Sonnet 4.5's **200K window** is the honest constraint in 2026. The 1M-token context window is a Sonnet 4.6 and Sonnet 5 feature, and on those models it is the default at standard pricing per the [long-context pricing docs](https://platform.claude.com/docs/en/about-claude/pricing#long-context-pricing). Server-side compaction, Anthropic's managed answer to conversations that outgrow the window, is also a 4.6-and-later feature. On Sonnet 4.5 you manage context yourself: retrieval, summarization checkpoints, and tool-result pruning. The one mercy is graceful overflow, since the 4.5 generation returns a `model_context_window_exceeded` stop reason rather than hard-failing the request. Now the migration math, which is where this brief earns its keep. The successors are **Sonnet 4.6** ($3/$15, 1M context, 128K output, same tokenizer) and **Sonnet 5** (1M context, 128K output, and introductory pricing of **$2/$10 through August 31, 2026**, returning to $3/$15 from September 1). The intro price looks like a straight 33% cut. It is not, for one verified reason: per Anthropic's pricing docs, **Sonnet 5 uses a newer tokenizer that produces roughly 30% more tokens for the same text**, while Sonnet 4.6 and earlier keep the old tokenizer. Through August, those two effects roughly cancel and Sonnet 5 is approximately cost-neutral against 4.5 on the same workload. From September 1, the same workload on Sonnet 5 costs on the order of 30% more than it did on Sonnet 4.5, unless the newer model's quality lets you cut tokens elsewhere. Sonnet 4.6 is the migration that buys the 1M window and compaction with no tokenizer inflation and no price change. ## When to use it, and when not **Pin Claude Sonnet 4.5 when:** - You have a validated production agent already running on it. A pinned snapshot that passes your evals is an asset; do not migrate on marketing cadence. - Your workload fits comfortably in 200K with good context hygiene, and your token bill is tokenizer-sensitive. - You need the proven cost/latency point: "Fast" latency class at $3/$15 with 0.1x cache reads is still, in July 2026, one of the best-understood price/performance positions in the market. **Do not choose it when:** - You are starting a new build. Start evals at Sonnet 4.6 and Sonnet 5; only land on 4.5 if your task evals say so, which occasionally they do. - Your documents or agent sessions genuinely need the 1M window or server-side compaction rather than better retrieval. See our [RAG vs long-context](/compare/rag-vs-fine-tuning) discussion before deciding they do. - A January 2025 knowledge cutoff is a real liability for your domain and you are not grounding with retrieval or search. ## How we would architect it for a client 1. **Pin, but instrument.** `claude-sonnet-4-5-20250929` in one config file, golden-set evals in CI, and a standing quarterly bake-off against the current Sonnet line so the migration decision is data, not vibes. 2. **Cache-first prompt architecture.** Stable system prompt and tool definitions at the prefix, volatile context at the tail, 1-hour cache for long-running agents. This is the highest-ROI engineering hour on any Sonnet-class deployment. 3. **Context discipline over context size.** Tool-result pruning, summarization checkpoints around the 150K mark, and retrieval instead of window-stuffing. Teams that build this muscle on 200K get materially cheaper 1M-window behavior if they later migrate, because a full million-token prompt is never the cheap path. 4. **Route by difficulty.** Haiku 4.5 ($1/$5) underneath for classification and extraction volume, Sonnet 4.5 as the agent workhorse, Opus-class or [a reasoning model](/ai/deepseek-r1) above it for the rare hard cases, with the [OpenAI comparison](/compare/openai-vs-anthropic) rerun on your own tasks each quarter. Sonnet 4.5's brief is ultimately about discipline: it rewards teams that treat a good model as infrastructure to be measured and defended, and it quietly punishes teams that chase every successor by default. #### Frequently Asked Questions **Q: Is Claude Sonnet 4.5 still available, or is it deprecated?** As of July 2026 it is fully available but sits in Anthropic's legacy-models tier. It is not deprecated: in the current lineup only Opus 4.1 carries a deprecation notice. The model ID claude-sonnet-4-5-20250929 is a pinned snapshot on the Claude API, Amazon Bedrock, and Google Cloud, so existing deployments keep running with stable behavior. The practical risk is staleness (a January 2025 knowledge cutoff), not sudden retirement. **Q: What does Claude Sonnet 4.5 cost?** Verified against Anthropic's pricing docs as of July 2026: $3 per million input tokens and $15 per million output tokens. Prompt cache reads are $0.30 (10% of base input), 5-minute cache writes $3.75, 1-hour writes $6, and the Batch API halves everything to $1.50/$7.50. On Bedrock and Google Cloud, regional or multi-region endpoints add a 10% premium over global routing. **Q: Does Claude Sonnet 4.5 have a 1M token context window?** No. Per Anthropic's current documentation it has a 200K-token context window and 64K max output. The 1M window, at standard pricing and with server-side compaction support, belongs to Claude Sonnet 4.6 and Sonnet 5. If your workload genuinely needs more than 200K, the answer is usually Sonnet 4.6 (same $3/$15 price, same tokenizer) rather than window-stuffing on 4.5, though better retrieval beats a bigger window more often than teams expect. **Q: Should we migrate from Sonnet 4.5 to Sonnet 5?** Run the math before the upgrade reflex. Sonnet 5's introductory pricing is $2/$10 through August 31, 2026, but it uses a newer tokenizer that produces roughly 30% more tokens for the same text, so the intro period is closer to cost-neutral than it looks, and from September 1 the same workload costs about 30% more than on 4.5 at the restored $3/$15. Migrate for capability (1M context, compaction, higher output limits) if your evals show it pays, and consider Sonnet 4.6 as the no-inflation middle path. **Q: Why did Sonnet 4.5 become the default model for production agents?** It hit a specific combination first: fast-class latency at $3/$15, strong tool use, extended thinking, and long-horizon stability (Anthropic reported it holding focus for over 30 hours on multi-step tasks at launch). It also has context awareness: the API injects the remaining token budget and updates it after each tool call, so the model paces work against real capacity. Those properties, plus 10x-cheaper cache reads on stable prompt prefixes, match exactly how agent loops spend money. **Q: Are the launch benchmark numbers (77.2% SWE-bench Verified) trustworthy?** They are Anthropic's self-reported figures from the September 2025 announcement, and we treat all vendor-reported benchmarks the same way regardless of vendor: directionally useful, never decision-grade. Nine months on, the more relevant fact is deployment history, since a large base of production agents was built and validated on this model. For your decision, run task-level evals on your own workload; that is the standard we apply in every engagement. **Q: Does BearPlex deploy Claude Sonnet 4.5 in client work?** Yes. It has been a standard workhorse lane in our agent architectures: cache-first prompt design, difficulty-based routing with Haiku 4.5 underneath, and pinned snapshots defended by golden-set evals in CI. For new builds we now start evaluations at Sonnet 4.6 and Sonnet 5, but we do not force-migrate healthy 4.5 deployments; a validated pinned model is an asset, and migration is a data decision, not a calendar one. --- ### Gemini 3.5 Flash: Frontier LLM *Publisher: Google DeepMind · Paper date: 2026.05.19 · Brief date: 2026.07.03 · 10 min read* *Parameters: Undisclosed · License: Proprietary (API)* URL: https://www.bearplex.com/ai/gemini-3-5-flash arXiv: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/ **Excerpt**: BearPlex's engineering brief on Gemini 3.5 Flash: verified pricing, the 1M context window economics, Vertex AI vs AI Studio access paths, and when to choose it. The first thing to verify about Gemini in mid-2026 is the name, because Google's lineup inverted the usual hierarchy. The current flagship is not a Pro model. Per [ai.google.dev](https://ai.google.dev/gemini-api/docs/models), the top stable model is **Gemini 3.5 Flash**, which Google describes as its "most intelligent model for sustained frontier performance on agentic and coding tasks." The strongest Pro-branded model, **Gemini 3.1 Pro**, is still in preview, and the announced Gemini 3.5 Pro had not reached the API model list as of July 2026. If your evaluation matrix still says "Gemini 2.5 Pro" in the flagship slot, it is a year out of date; if it says "Flash means the cheap tier," it is two months out of date. ## What it actually is Gemini 3.5 Flash ([announced May 19, 2026 at Google I/O](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/)) is a generally available multimodal model: text, image, video, audio, and PDF in, text out. The verified spec sheet from the [official model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-flash): model ID `gemini-3.5-flash`, **1,048,576-token input window**, **65,536-token output limit**, knowledge cutoff January 2025, with thinking, function calling, structured outputs, context caching, code execution, computer use (preview), search grounding, and a batch API all supported. Google's own launch numbers position it above its Pro sibling: 76.2% on Terminal-Bench 2.1, 83.6% on MCP Atlas, and an "outperforms Gemini 3.1 Pro on challenging coding and agentic benchmarks" framing, plus a claim of roughly 4x the output speed of other frontier models. Those are vendor-reported; we cite them as Google's claims, not as findings. The credible core is the strategy: Google collapsed "frontier intelligence" and "high throughput" into one SKU and priced it accordingly. ## Commercial terms Hosted API, no weights, two distinct front doors, and choosing the door is a real architectural decision: - **AI Studio / Gemini Developer API**: API-key access, a genuinely useful free tier (verified on the [pricing page](https://ai.google.dev/gemini-api/docs/pricing)), minimal setup. This is the prototyping path, and for many products it remains the production path. - **Vertex AI**: the same model family behind Google Cloud's enterprise controls. Per the [Vertex model docs](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/learn/models), this is where you get IAM-based access control, VPC Service Controls, customer-managed encryption keys, Private Service Connect, audit logging, and **Provisioned Throughput** for reserved capacity. The rule we apply in client work: if the workload touches regulated data, needs contractual capacity, or must live inside an existing GCP governance perimeter, build on Vertex from day one. Migrating from an AI-Studio-keyed codebase to Vertex auth and quotas later is not hard, but it is never free, and it always lands in the week you least want it. ## Real API cost Verified against the [Gemini API pricing page](https://ai.google.dev/gemini-api/docs/pricing) as of July 2026, paid tier, per million tokens: | Model | Input | Output | Status | |---|---|---|---| | Gemini 3.5 Flash | $1.50 | $9.00 | GA | | Gemini 3.1 Pro (≤200K / >200K) | $2.00 / $4.00 | $12.00 / $18.00 | Preview | | Gemini 3.1 Flash-Lite | $0.25 | $1.50 | GA | | Gemini 2.5 Flash | $0.30 | $2.50 | GA | | Gemini 2.5 Pro (≤200K / >200K) | $1.25 / $2.50 | $10.00 / $15.00 | GA | Batch runs at 50% off across the line (3.5 Flash at $0.75/$4.50), and context caching bills cached 3.5 Flash tokens at $0.15 with $1.00/hour storage. Two readings matter. First, the flagship price point is aggressive: $1.50/$9.00 sits well under [gpt-5.5's $5/$30](/ai/gpt-5) and under Sonnet-class $3/$15, which is exactly the pressure Google intends. Second, "Flash" no longer means cheap: 3.5 Flash costs 5x the input of 2.5 Flash. Teams that treated Flash as the bargain tier need to re-point that role at 3.1 Flash-Lite ($0.25/$1.50) and re-run the math. ## The long-context story The entire current Gemini line runs a **1,048,576-token input window**, and Gemini has held the 1M position longer than any competitor, since early 2024. The 2026 nuance is in the billing: on the Pro models, tokens beyond 200K are billed at a higher tier (3.1 Pro doubles to $4.00 input above 200K), while **3.5 Flash has no published long-context surcharge**, making it the economically interesting choice for genuinely long inputs. Our standing engineering advice survives the big window: a million tokens of context is a capability, not an architecture. Retrieval that puts the right 30K tokens in front of the model beats 800K tokens of haystack on cost, latency, and usually accuracy. Where the 1M window genuinely earns its keep is whole-corpus reasoning you cannot pre-chunk well: full codebases, long video, discovery-style document sets, and as the fallback lane when [RAG](/services/rag-knowledge) misses. Use context caching aggressively either way; at $0.15 per million cached tokens, a stable long prefix costs 10% of resending it. ## When to use it, and when not **Use Gemini 3.5 Flash when:** - Long-input multimodality is the workload: video, audio, PDFs, and mixed document sets at 1M-token scale with no long-context price cliff. - You want frontier-class agent capability at $1.50/$9.00, the lowest verified flagship price point of July 2026, and you will validate quality on your own evals. - You are already a GCP shop: Vertex governance, Provisioned Throughput, and existing IAM make it the path of least organizational resistance. **Do not use it when:** - Weights-on-your-hardware is the constraint; that conversation belongs to [DeepSeek V3](/ai/deepseek-v3) and [Mistral Large 3](/ai/mistral-large-3). - You need the Pro-branded top end today: 3.1 Pro is still preview-stage, and preview models are not something we let clients build load-bearing systems on. - Your budget tier was actually 2.5-Flash-shaped. The new bargain lane is 3.1 Flash-Lite; benchmark it before assuming 3.5 Flash is the drop-in upgrade. ## How we would architect it for a client 1. **Decide the front door first.** Regulated or capacity-sensitive workloads go Vertex (VPC-SC, CMEK, Provisioned Throughput); everything else can start on the Developer API with the free tier doing evaluation duty. Wrap the client SDK so the door can change without touching product code. 2. **Two-lane routing inside the family**: 3.1 Flash-Lite for extraction and classification volume, 3.5 Flash for agentic and long-context work. The 6x price gap between lanes dominates most optimization you could do inside either lane. 3. **Long-context with a budget guard.** Even without a surcharge, a full-window 3.5 Flash call is roughly $1.57 of input; we cap per-request context in the gateway, cache stable prefixes, and route to retrieval when the requested window exceeds the cap. 4. **Quarterly cross-vendor bake-off.** Google's lineup moved three times in the twelve months to July 2026 (2.5 to 3 to 3.1 to 3.5); the [managed-versus-self-hosted question](/compare/self-hosted-vs-managed-llm) and the vendor question get rerun on your tasks, on schedule, in every [model engineering](/services/model-engineering) engagement. The naming will keep churning. The engineering posture that survives it: verify the lineup at the docs page, price the tiers from the pricing page, and let your own evals, not the brand suffix, assign each model its lane. #### Frequently Asked Questions **Q: What is the current flagship Gemini model?** As of July 2026, the flagship stable model on ai.google.dev is Gemini 3.5 Flash (model ID gemini-3.5-flash), which Google describes as its most intelligent model for agentic and coding tasks. The strongest Pro-branded model, Gemini 3.1 Pro, is still in preview, and Gemini 3.5 Pro was announced as in progress at I/O 2026 but had not reached the API model list. The Gemini 2.5 line remains available as the previous stable generation. **Q: What does Gemini 3.5 Flash cost?** Verified against the Gemini API pricing page as of July 2026: $1.50 per million input tokens and $9.00 per million output tokens on the paid tier, with batch processing at half price ($0.75/$4.50) and context-cache reads at $0.15 per million plus $1.00 per hour of storage. A free tier exists on the Developer API. Note the tier shift: 3.5 Flash costs about 5x the input rate of Gemini 2.5 Flash, so Flash no longer means the budget lane. **Q: How big is Gemini 3.5 Flash's context window?** 1,048,576 input tokens with a 65,536-token output limit, per the official model page. Unlike the Pro models (Gemini 3.1 Pro bills input at $4.00 instead of $2.00 beyond 200K tokens), 3.5 Flash has no published long-context price tier, which makes it the economical choice for genuinely long inputs. We still cap context per request and prefer retrieval where it works: a full-window call is roughly $1.57 of input before output costs. **Q: Should we access Gemini through AI Studio or Vertex AI?** AI Studio and the Gemini Developer API give you API-key access and a free tier, which is ideal for prototyping and fine for many production products. Vertex AI serves the same model family with enterprise controls: IAM, VPC Service Controls, customer-managed encryption keys, Private Service Connect, audit logging, and Provisioned Throughput for reserved capacity. Our rule: regulated data, contractual capacity needs, or an existing GCP governance perimeter mean you build on Vertex from day one. **Q: Is Gemini 3.5 Flash really better than Gemini 3.1 Pro?** That is Google's claim, and the numbers behind it (76.2% Terminal-Bench 2.1, 83.6% MCP Atlas, outperforming 3.1 Pro on agentic coding benchmarks) are Google's own launch figures, so we treat them as directional. Two facts are independently verifiable: 3.5 Flash is GA while 3.1 Pro remains preview, and 3.5 Flash is cheaper ($1.50/$9.00 versus $2.00/$12.00). For production work in July 2026, the GA flagship is the sensible default and your own task evals settle the rest. **Q: How does Gemini 3.5 Flash pricing compare to GPT-5.5 and Claude?** On July 2026 list prices, aggressively: $1.50/$9.00 per million tokens versus $5.00/$30.00 for gpt-5.5 and $3/$15 for Sonnet-class Claude models. List price is not cost of ownership: token consumption per task, caching behavior, output verbosity, and failure-retry rates differ by model, so we compare cost per completed task on the client's own workload rather than per-token rates. But the headline gap is real and is exactly the pressure Google priced for. **Q: Does BearPlex build on Gemini?** Yes, where evals support it, most often for long-input multimodal workloads (video, audio, large document sets) and for clients already inside the GCP perimeter where Vertex governance is the path of least resistance. We run it as one lane in a model-agnostic gateway with 3.1 Flash-Lite underneath for volume traffic, context budgets enforced at the gateway, and a quarterly cross-vendor bake-off, because Google's lineup changed three times in the year to July 2026 and stale assumptions are the real cost driver. --- ### DeepSeek V3: Open-Weights LLM *Publisher: DeepSeek-AI · Paper date: 2024.12.27 · Brief date: 2026.07.03 · 10 min read* *Parameters: 671B MoE (37B active); V3.2 685B · License: MIT (V3-0324 onward)* URL: https://www.bearplex.com/ai/deepseek-v3 arXiv: https://arxiv.org/abs/2412.19437 Model card: https://huggingface.co/deepseek-ai/DeepSeek-V3-0324 GitHub: https://github.com/deepseek-ai/DeepSeek-V3 **Excerpt**: BearPlex's engineering brief on DeepSeek V3: the V3 vs R1 routing decision, MIT open-weights status through V3.2, self-hosting cost, and the V4-era API reality. DeepSeek R1 got the headlines; V3 does the work. In every architecture we have built around the DeepSeek family, the V3 line handles the overwhelming majority of traffic, because most production requests do not need a reasoning model, they need a fast, competent generalist. This brief covers the V3 line as an engineering asset in mid-2026: what the weights are, what the license ladder actually says, what happened to the first-party API, and, most importantly, the V3-versus-R1 routing decision that determines your economics. ## What it actually is DeepSeek-V3 ([technical report, December 2024](https://arxiv.org/abs/2412.19437)) is a Mixture-of-Experts model with **671B total parameters and 37B activated per token**, built on Multi-head Latent Attention and the DeepSeekMoE architecture, trained on 14.8T tokens, with a **128K context window** per the [official model card](https://huggingface.co/deepseek-ai/DeepSeek-V3). Its claim to fame at release was efficiency: the report puts full training at 2.788M H800 GPU hours, a number that reset industry assumptions about frontier training budgets. The line evolved in three verified steps, and the checkpoint you evaluate matters: 1. **V3-0324** (March 2025, [model card](https://huggingface.co/deepseek-ai/DeepSeek-V3-0324)): the same architecture post-trained harder. DeepSeek's own numbers: MMLU-Pro 75.9 to 81.2, AIME 39.6 to 59.4, plus "increased accuracy in Function Calling, fixing issues from previous V3 versions." That last line is the one agent builders should read twice. 2. **V3.2** (2025, [model card](https://huggingface.co/deepseek-ai/DeepSeek-V3.2)): 685B parameters, introduces **DeepSeek Sparse Attention (DSA)**, which the card describes as substantially reducing computational complexity while preserving performance, a revised chat template with tool-aware thinking, and self-reported performance in GPT-5's neighborhood. 3. **V4** (2026): the successor generation, also open-weights on Hugging Face. The V3 line is no longer the newest DeepSeek, which is precisely why it is now cheap, well-understood infrastructure. V3 is also the base that [DeepSeek R1](/ai/deepseek-r1) was trained from, which is why the two route so cleanly together: same lineage, same tooling, different thinking budgets. ## The license, and what commercial use really permits The ladder is simple and worth getting exactly right. The **original V3** shipped with MIT code but weights under a DeepSeek "Model License" that explicitly supports commercial use. From **V3-0324 onward, the weights themselves are MIT**: the model card states "this repository and the model weights are licensed under the MIT License," and [V3.2](https://huggingface.co/deepseek-ai/DeepSeek-V3.2) carries the same MIT license. No user thresholds, no attribution badges, no naming rules on derivatives, distillation expressly on the table. Practical read: standardize on V3-0324 or later and the license conversation with counsel takes five minutes. The China question gets the same answer we gave in the R1 brief: self-hosted open weights are inspectable files in your VPC with no telemetry and no vendor in the request path, which is a stronger data posture than any hosted API, DeepSeek's or anyone else's. ## Real cost: API reality and self-hosting **The first-party API story changed in 2026, and older writeups will mislead you.** As of July 2026, [DeepSeek's API platform](https://api-docs.deepseek.com/quick_start/pricing) serves the V4 generation: `deepseek-v4-flash` (1M context, cache-miss input $0.14, output $0.28 per million tokens) and `deepseek-v4-pro` ($0.435 in / $0.87 out). The platform states plainly that the legacy names `deepseek-chat` and `deepseek-reasoner` are **deprecated on July 24, 2026**, with the old names mapping to V4-flash modes in the interim. There is no first-party per-token V3 endpoint to build on anymore. So the V3 line in 2026 is what the MIT license always made it: a model you run yourself or rent from a GPU cloud. Self-hosting footprints, derived from published parameter counts (weights only, before KV cache), mirror R1's, since it is the same chassis: - **Full V3 at native FP8**: roughly 670GB of weights, an 8-GPU datacenter node in the 96GB-141GB-per-card class. - **At ~4-bit quantization**: roughly 340GB, still multiple datacenter GPUs, with quality-versus-quant evals mandatory before committing. - **Per-token compute is the good news**: 37B active parameters means throughput behaves like a mid-size dense model once the memory is paid for. Memory is the entry fee; serving economics are the reward. If that footprint is out of budget, the honest alternatives are a hosted V3/V3.2 endpoint from a GPU cloud, or a smaller open model entirely, such as [Qwen 3](/ai/qwen-3) mid-sizes. ## The V3-versus-R1 routing decision This is the section this brief exists for. The industry default of "buy the smartest model and send everything to it" is exactly backwards with a reasoning model in the stack, because reasoning is priced in output tokens and time. DeepSeek's own R1-0528 notes put average reasoning depth around **23K tokens per AIME-class question**; a V3-class model answers a routine request in a few hundred tokens, seconds sooner. The decision rule we deploy: - **Default lane: V3.** Extraction, summarization, drafting, classification, translation, straightforward tool calls, RAG answer synthesis. This is 80-95% of real traffic in the systems we run, and on this traffic a reasoning model is strictly worse: slower, costlier, no measurable quality gain. - **Escalation lane: R1.** Multi-step planning, math-adjacent logic, code synthesis with intricate constraints, anything where your evals show chain-of-thought actually moves accuracy. - **The router is a classifier, not a vibe.** Route on measurable signals (task type, schema complexity, retry history), audit the escalation rate, and treat "we route by difficulty" as a claim your logs must prove. V3.2's tool-aware thinking modes blur this line inside a single checkpoint, which is convenient, but the budget discipline is identical: thinking tokens are spend, and the router (or the mode flag) is where you control it. ## When to use it, and when not **Use DeepSeek V3 when:** - You need a frontier-adjacent generalist inside your own infrastructure boundary under a clean MIT license. - You are building the two-lane DeepSeek stack: V3 for volume, R1 for the hard 5-20%, one toolchain, one deployment story. - Function calling and structured outputs on-prem matter (use V3-0324 or later; the card's own fix notes tell you why). **Do not use it when:** - You have no GPU story and wanted a first-party API: the V3-era endpoints are gone as of July 2026, and pretending otherwise is technical debt with a countdown timer. - Your workload fits a 32B-class open model; a [smaller model](/ai/qwen-3) at a fraction of the memory footprint wins the total-cost math. - You need a contractual counterparty behind the model. MIT weights ship with no warranty and no SLA; that is the trade. ## How we would architect it for a client The pattern is the [sovereign-cloud](/services/sovereign-cloud) two-lane stack we described in the R1 brief, with V3 promoted to its rightful place as the default lane: 1. **Routing gateway in front**, classifying every request; V3 (or a hosted V3.2 endpoint) takes the volume lane, R1 or V3.2-thinking takes the escalation lane, and the escalation rate is a monitored SLO, not a hope. 2. **One serving stack**: same MoE chassis, same quantization and serving toolchain for both lanes, which halves the operational surface compared to mixing model families. 3. **Deterministic verification after generation**: schema validation, code-checked claims, bounded re-prompts. Non-reasoning models fail faster and cheaper; they still fail, and the harness, not the model, owns correctness. 4. **An eval track against V4**: the successor weights are open too, and the V3-to-V4 promotion should happen when your task evals justify it, on your schedule rather than a vendor's deprecation calendar. That optionality is what owning weights buys, and it is the quiet, compounding argument for the whole [open-weights approach](/compare/open-source-vs-closed-source-llm). #### Frequently Asked Questions **Q: Is DeepSeek V3 free for commercial use?** Yes. From the V3-0324 checkpoint onward, the model card states the repository and the model weights are licensed under MIT, and V3.2 carries the same license: no user thresholds, no attribution requirements, no naming rules, distillation permitted. The original December 2024 V3 release used a separate DeepSeek Model License for the weights that also explicitly supports commercial use, but the clean recommendation is to standardize on V3-0324 or later and keep the stack fully MIT. **Q: When should we use DeepSeek V3 instead of R1?** For most traffic. Extraction, summarization, drafting, classification, RAG synthesis, and routine tool calls, which make up 80-95% of requests in typical production systems, get no measurable benefit from a reasoning model, just higher latency and output-token cost (DeepSeek's own notes show R1-0528 averaging around 23K reasoning tokens on hard questions). Route V3 as the default lane and escalate to R1 only where your evals show chain-of-thought moves accuracy, and audit that escalation rate in your logs. **Q: Can we still use DeepSeek V3 through DeepSeek's own API?** Effectively no, as of July 2026. The first-party platform now serves the V4 generation (deepseek-v4-flash and deepseek-v4-pro), and DeepSeek's pricing page states the legacy deepseek-chat and deepseek-reasoner names are deprecated on July 24, 2026, mapping to V4 modes in the interim. The durable ways to run the V3 line are the MIT weights in your own infrastructure or a hosted endpoint from a third-party GPU cloud. **Q: What does it cost to self-host DeepSeek V3?** The published architecture is 671B total parameters with 37B active per token, so weights alone are roughly 670GB at native FP8: an 8-GPU datacenter node before you account for KV cache. Around 4-bit quantization that drops to roughly 340GB, still multi-GPU, and quantized reasoning quality needs evaluation before you commit. The compensating economics: only 37B parameters are active per token, so once memory is provisioned, throughput behaves like a mid-size dense model. **Q: Which V3 checkpoint should we deploy?** V3-0324 at minimum, and evaluate V3.2. V3-0324 is where the weights went MIT and where DeepSeek's card reports major post-training gains (MMLU-Pro 81.2, AIME 59.4) plus explicit function-calling fixes, which matter for agent work. V3.2 (685B, MIT) adds DeepSeek Sparse Attention for cheaper long-context compute and a tool-aware thinking mode. The original December 2024 checkpoint is now primarily of historical and research interest. **Q: How does DeepSeek V3 relate to R1 and V4?** V3 is the base model R1's reasoning training was built on, which is why the two pair so naturally in a routed architecture: same lineage and toolchain, different thinking budgets. V4 is the 2026 successor generation, and notably its weights are also published on Hugging Face. Owning V3 in production therefore comes with a built-in upgrade path: run V4 in an offline eval track and promote it when your own task metrics justify it, on your schedule. **Q: Does BearPlex deploy DeepSeek V3 in client work?** We evaluate it as the default volume lane in DeepSeek-based architectures: a routing gateway in front, V3-class serving for the majority of traffic, R1 or a thinking mode for the hard minority, deterministic verification behind both, all inside the client's cloud boundary. It is the standard candidate wherever data residency, license cleanliness, and cost control at scale outweigh the convenience of a hosted frontier API, and the client's own evals make the final call. --- ### Mistral Large 3: Open-Weights LLM *Publisher: Mistral AI · Paper date: 2025.12.02 · Brief date: 2026.07.03 · 10 min read* *Parameters: 675B MoE (41B active) incl. 2.5B vision encoder · License: Apache 2.0* URL: https://www.bearplex.com/ai/mistral-large-3 arXiv: https://mistral.ai/news/mistral-3 Model card: https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512 **Excerpt**: BearPlex's engineering brief on Mistral Large 3: Apache 2.0 terms verified, the EU-sovereignty case, hosted API vs self-host economics, and when to choose it. Mistral Large 3 is the model you evaluate when the first requirement in the RFP is not a benchmark, it is a jurisdiction. Plenty of models are capable; very few combine frontier-class scale, a genuinely permissive license, and a European vendor. As of July 2026 this is the only place where all three meet, and that intersection, not the leaderboard, is where Large 3 wins engagements. This brief covers what shipped, what Apache 2.0 actually settles, the sovereignty argument stated precisely, and the deployment math. ## What it actually is Mistral Large 3 shipped on [December 2, 2025](https://mistral.ai/news/mistral-3) as the flagship of the Mistral 3 family, and Mistral's own description is the headline: "a state-of-the-art, open-weight, general-purpose multimodal model." The verified specs from the [model card](https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512): a granular Mixture-of-Experts design totaling **675B parameters with 41B active**, composed of a 673B-parameter language model (39B active) plus a 2.5B vision encoder, with a **256K context window**. The family context matters for architecture: the same Apache 2.0 release included **Ministral 3** at 14B, 8B, and 3B, giving the Large 3 stack a same-vendor, same-license distill ladder for volume traffic. Above it in Mistral's catalog (verified at [docs.mistral.ai](https://docs.mistral.ai/getting-started/models/models_overview/), July 2026) sit newer specialist releases, with **Mistral Medium 3.5** (v26.04) positioned as the frontier-class agentic and coding model; the flagship *open-weight* large model remains Large 3 (v25.12). The MoE math is the same lesson as every model in this class: 41B active parameters means per-token compute like a mid-size dense model, while 675B total means memory sized to the full parameter count. Cheap to run per token, expensive to hold. ## The license, and what commercial use really permits **Apache 2.0, full stop.** The release post states "all models are released under the Apache 2.0 license," and the model card confirms it. In the taxonomy of open-weight licensing this is the clean end of the spectrum, alongside [Qwen 3](/ai/qwen-3) and MIT-licensed [DeepSeek](/ai/deepseek-v3), and it is a meaningful break from Mistral's own earlier era of research-only licenses on flagship weights: no MAU gates, no attribution badges, no derivative naming rules (the [Llama 4 contrast](/ai/llama-4)), fine-tuning and redistribution permitted, plus Apache 2.0's explicit patent grant, which some counsel prefer even over MIT. One caution from our verification pass: Mistral's own pricing-page FAQ still carries stale generic language about commercial deployments requiring a separate Mistral license. For the Mistral 3 family, the release post and the Hugging Face model card are the controlling documents, and both say Apache 2.0. Have counsel cite those, not the FAQ. ## The EU-sovereignty angle, stated precisely "Sovereign AI" is mostly a marketing word; here is the engineering substance. A sovereignty-constrained deployment (EU public sector, healthcare, financial services under EU data-protection regimes) typically needs three properties: data never leaves a controlled boundary, the vendor relationship survives geopolitical friction, and the stack can be audited. Mistral Large 3 is the strongest combined answer in the July 2026 market because each property has a concrete mechanism: - **Data boundary**: Apache 2.0 weights run in your EU datacenter or EU-region VPC. No inference traffic to any vendor, no third-country data transfer to analyze in the DPIA. This is a property of self-hosted open weights generally; what Large 3 adds is frontier scale under that posture. - **Vendor jurisdiction**: Mistral AI is a French company. For procurement teams whose legal review stalls on US or Chinese vendor jurisdiction, an EU-headquartered publisher removes the hardest questions, and does so even in the hosted-API scenario. - **Auditability**: weights are inspectable artifacts; your red team evaluates the actual deployed model, not a vendor attestation. The honest caveat: if you self-host on Azure, AWS, or GCP EU regions, your *infrastructure* provider is still a US hyperscaler. Full sovereignty arguments end at EU-owned infrastructure, and that is an infrastructure decision Large 3 enables but does not make for you. We work through exactly this layering in [sovereign cloud](/services/sovereign-cloud) engagements. ## Real cost: hosted API versus self-host **Hosted path.** Large 3 serves as `mistral-large-latest` on Mistral's own platform, and the [pricing page](https://mistral.ai/pricing) lists Mistral Large at **$2 per million input tokens and $6 per million output tokens** (as of July 2026). That is an aggressive flagship price: below [Sonnet-class](/ai/claude-sonnet-4-5) $3/$15 and far below [gpt-5.5's](/ai/gpt-5) $5/$30 on list. Availability is unusually broad for day one: Amazon Bedrock, Azure Foundry, Hugging Face, IBM WatsonX, and the major GPU clouds, per the release post. **Self-host path.** Derived from the published parameter counts (weights only, before KV cache): - **BF16**: roughly 1.35TB of weights, a multi-node deployment. Rarely the right call. - **FP8**: roughly 675GB, an 8-GPU node in the 96GB-141GB-per-card class, the same footprint class as DeepSeek V3. - **~4-bit**: roughly 340GB, several datacenter GPUs, with mandatory quality evals at that quantization. - **The Ministral ladder** (3B/8B/14B, Apache 2.0) covers workstation-to-single-GPU territory for the traffic that does not need the flagship. The pattern that survives contact with a budget: hosted API or Ministral-class self-host first, and the 675B flagship on your own metal only once volume, sovereignty requirements, or unit economics prove out the cluster. ## When to use it, and when not **Use Mistral Large 3 when:** - Sovereignty or data-residency constraints are load-bearing and you want frontier scale without a US or Chinese vendor in the loop. - You want one license (Apache 2.0) and one vendor family from 3B to 675B, with the same toolchain from pilot to flagship. - Multimodal input (the built-in vision encoder) matters inside a self-hosted boundary, where open-weight options are scarce. **Do not use it when:** - You would only ever use the hosted API and have no sovereignty constraint; then it competes purely on price and your task evals against the frontier APIs, and the verdict is workload-specific. - Your workload is agentic coding and Mistral's own catalog points you at Medium 3.5 instead; run that comparison rather than assuming the flagship wins. - You need a warranty. Apache 2.0 weights carry none; contractual accountability comes from the hosted platforms, not the license. ## How we would architect it for a client For an EU-regulated deployment, the reference shape we use: 1. **Two-tier serving inside the boundary**: Ministral 3 14B on a single GPU for volume traffic, Large 3 at FP8 on an 8-GPU node for the requests that need the flagship, one Apache 2.0 license file covering both. 2. **A routing gateway** enforcing the tier split and logging escalation rates, the same discipline as our [DeepSeek two-lane pattern](/ai/deepseek-v3). 3. **Hosted-API burst lane, jurisdiction permitting**: `mistral-large-latest` at $2/$6 absorbs load spikes so the on-prem cluster is sized for the median, not the peak. Where the DPIA forbids it, the gateway simply queues instead. 4. **Eval-gated everything**: quantization level, Ministral-versus-Large routing thresholds, and any migration to newer Mistral releases all move only when task-level evals say so, the standard we apply in every [model engineering](/services/model-engineering) engagement. The strategic read: Large 3 made "European frontier model" a real procurement category instead of a wish. If your constraints point there, it is not one option among many; as of July 2026 it is essentially the only complete answer. #### Frequently Asked Questions **Q: Is Mistral Large 3 really Apache 2.0, including commercial use?** Yes. The December 2, 2025 release post states all Mistral 3 family models are released under Apache 2.0, and the Hugging Face model card for Mistral-Large-3-675B-Instruct-2512 confirms it. That means commercial use, modification, fine-tuning, and redistribution with no MAU gates, attribution badges, or naming rules, plus Apache 2.0's patent grant. Note that Mistral's pricing-page FAQ still contains stale generic language about commercial licensing; the release post and model card are the controlling documents. **Q: What hardware does Mistral Large 3 need to self-host?** It is a 675B-parameter Mixture-of-Experts model (41B active), so weights alone are roughly 1.35TB at BF16, about 675GB at FP8, and around 340GB at 4-bit quantization: an 8-GPU datacenter node at FP8, before KV cache. Per-token compute behaves like a 41B model once memory is provisioned. Most deployments should start on the hosted API or the Apache 2.0 Ministral 3 models (3B/8B/14B) and promote to self-hosted Large 3 only when volume or sovereignty requirements justify the cluster. **Q: What does the Mistral Large 3 API cost?** Mistral's pricing page lists Mistral Large at $2 per million input tokens and $6 per million output tokens as of July 2026, served as mistral-large-latest on Mistral's platform. It is also available through Amazon Bedrock, Azure Foundry, IBM WatsonX, and major GPU clouds per the release announcement. On list price that undercuts the US frontier APIs, but we always compare cost per completed task on the client's own workload rather than per-token rates. **Q: Why does Mistral Large 3 matter for EU sovereignty?** It is the only July 2026 combination of frontier scale, a permissive Apache 2.0 license, and an EU-headquartered publisher. Self-hosted, the weights run entirely inside your EU boundary with no vendor in the inference path and no third-country transfer to analyze; and the vendor-jurisdiction questions that stall EU procurement on US or Chinese providers largely disappear. The honest caveat: hosting on a US hyperscaler's EU region still leaves a US infrastructure provider in the stack, so full sovereignty is an infrastructure decision the model enables but cannot make for you. **Q: How does Mistral Large 3 compare to DeepSeek V3 as an open-weights flagship?** They occupy the same hardware class (roughly 670GB at FP8, 8-GPU nodes) with similar MoE economics, and both carry clean permissive licenses: MIT for DeepSeek from V3-0324 onward, Apache 2.0 for Large 3. The differentiators are jurisdiction and modality: Mistral is a French publisher, which matters enormously in EU procurement, and Large 3 ships with a built-in vision encoder. Where sovereignty is not a constraint, we let task-level evals decide; where it is, the jurisdiction usually decides first. **Q: Should we use Mistral Large 3 or a smaller Mistral model?** Route by difficulty, not by brand tier. The Apache 2.0 Ministral 3 line (3B, 8B, 14B) handles classification, extraction, and routine generation on single-GPU or smaller footprints, and in our architectures that lane carries most traffic. Reserve Large 3 for the requests that measurably need frontier capability, and note Mistral's own catalog positions Medium 3.5 as its agentic coding specialist. The right split is an eval result on your workload, then enforced by a routing gateway with logged escalation rates. **Q: Does BearPlex deploy Mistral models in client work?** We evaluate the Mistral 3 family as the primary candidate wherever EU sovereignty or data-residency constraints are load-bearing: a two-tier Apache 2.0 stack (Ministral for volume, Large 3 at FP8 for the hard cases) inside the client's boundary, optionally with a hosted-API burst lane where the data-protection assessment permits it. As with every model we deploy, quantization levels, routing thresholds, and upgrade decisions are gated on the client's own task evals, not on vendor announcements. --- ### Grok 4.3: Frontier LLM *Publisher: xAI · Paper date: 2026.05.05 · Brief date: 2026.07.03 · 9 min read* *Parameters: Undisclosed · License: Proprietary (API)* URL: https://www.bearplex.com/ai/grok-4-3 arXiv: https://docs.x.ai/docs/models **Excerpt**: BearPlex's engineering brief on Grok 4.3: verified $1.25/$2.50 pricing and 1M context, the xAI API access reality, model churn risk, and where it fits. Grok 4.3 is the cheapest flagship-tier API of mid-2026 by a wide margin, and that single fact is why it lands on evaluation shortlists that would not otherwise include xAI. The engineering job is to work out what the price does and does not buy you. This brief covers the verified numbers, the state of the xAI API as an integration target, and a candid account of where Grok fits in a production portfolio and where it does not. ## What it actually is Grok 4.3 went live on the xAI API in the spring of 2026 (xAI's own announcement that it was live on the API is dated May 5, 2026), and xAI's [models documentation](https://docs.x.ai/docs/models) is unambiguous about its positioning: "It is the most intelligent and fastest model we've built," recommended for chat and coding. The verified spec: a **1M-token context window** at **$1.25 per million input tokens and $2.50 per million output tokens** (as of July 2026). The current lineup around it, verified on the same page, tells you how xAI thinks about model surface: - **grok-4.20** (the March 2026 generation, snapshot-dated `-0309`) ships as three explicit variants: **reasoning**, **non-reasoning**, and **multi-agent**, all at 1M context and the same $1.25/$2.50 price. - **grok-build-0.1**: a 256K-context builder-focused model at $1.00/$2.00. - The docs recommend the aliased names (`` or `-latest`) so you pick up point releases automatically; in production we pin snapshots instead, on every vendor, for the same reason we pin dependencies. Note what is *not* on that page as of July 2026: the original Grok 4 and Grok 3. Barely a year after Grok 4's mid-2025 launch, the pre-4.2 lineup has been retired from the current model table entirely. Hold that thought for the risk section. ## Commercial terms Hosted API under xAI's terms; no weights, no self-host path. Pricing is published and simple, which is genuinely to xAI's credit: one flagship price, no long-context surcharge published for the 1M window, no priority-tier multiplication table. Two commercial observations from our verification pass: 1. **The price is the strategy.** $1.25/$2.50 against [gpt-5.5's](/ai/gpt-5) $5/$30 and [Sonnet-class](/ai/claude-sonnet-4-5) $3/$15 is a 4-12x list-price gap on output tokens. xAI is buying market share, and buyers should enjoy it while pricing it as promotional rather than structural. 2. **The enterprise story is thinner than the price sheet.** The public model docs we verified say little about data residency, compliance attestations, or capacity reservations compared to the documentation depth of Azure OpenAI, Bedrock, or Vertex paths. That is not an accusation, it is a procurement work item: if your compliance regime needs specific guarantees, get them in writing from xAI before committing an architecture, because the docs alone will not answer your security questionnaire. ## Real API cost Per the [official models page](https://docs.x.ai/docs/models), as of July 2026, per million tokens: | Model | Context | Input | Output | |---|---|---|---| | grok-4.3 | 1M | $1.25 | $2.50 | | grok-4.20 (reasoning / non-reasoning / multi-agent) | 1M | $1.25 | $2.50 | | grok-build-0.1 | 256K | $1.00 | $2.00 | The output price is the story. Agentic workloads are output-heavy (tool calls, drafts, retries, reasoning tokens), and at $2.50 per million output tokens, Grok 4.3's cost per completed agent task can undercut rivals even if it needs more attempts. That "even if" is measurable, and you should measure it: our standard is cost per *completed, verified* task, not cost per token, and models with higher retry rates lose more of their price advantage than teams expect. What we could not verify on the public page: cached-input pricing and per-call pricing for live search or web-search tooling. Budget conservatively until your own invoices tell you otherwise. ## API access reality Integration is deliberately low-friction, and the docs' recommended defaults reveal the intended audience: aliased model names, automatic feature pickup, one price. For an engineering team, the realities to plan around: - **Model churn is fast and real.** The original Grok 4 went from flagship launch to absent-from-the-docs in roughly a year, and the 4.20 generation shipped with date-stamped snapshots barely two months before 4.3 superseded it. That cadence is [OpenAI-speed](/ai/gpt-5), without the published minimum-notice deprecation policy we could verify on OpenAI's side. Pin snapshots, wrap the vendor behind your gateway, and keep your eval suite ready for forced migrations. - **The stale-cutoff trap.** The docs state a November 2024 knowledge cutoff for Grok 3 and Grok 4 and publish no cutoff for 4.3 on the models page we verified. Treat parametric knowledge as unreliable for anything recent and ground the model with retrieval or search tooling, which is good practice on every vendor and mandatory here. - **The variant surface is unusual.** Explicit reasoning versus non-reasoning versus multi-agent SKUs at identical prices means routing is by capability rather than by budget: pick the variant per lane the same way you would set `reasoning_effort` elsewhere. ## When to use it, and when not **Use Grok 4.3 when:** - Output-token economics dominate: high-volume agentic or generation workloads where $2.50 output changes the unit economics, validated by your own cost-per-completed-task evals. - You need a 1M-token window without a long-context surcharge and your inputs genuinely are that long. - It serves as the price-pressure lane in a multi-vendor gateway: even when Grok does not win a lane, its quote disciplines your negotiation with the vendors that do. **Do not use it when:** - Your compliance regime needs documented residency, attestations, or contractual capacity that you have not yet obtained from xAI in writing. Price does not answer a security questionnaire. - Brand-sensitivity review is part of your deployment gate and you have not run it here. Grok's consumer persona is distinctive by design; your evaluation should include tone and refusal-behavior testing against your own content policies, exactly as we recommend for every vendor, and with extra care here. - You need weights, a self-host path, or multi-year model stability. None of the three is on offer; for the first two see the [open-weights conversation](/compare/open-source-vs-closed-source-llm). ## How we would architect it for a client Grok enters our architectures the same way every frontier API does: as a lane behind a gateway, never as a hard dependency. 1. **Gateway-wrapped, snapshot-pinned.** Aliased names are convenient and we do not use them in production. Model IDs live in config; xAI's churn record makes this non-negotiable. 2. **Cost-per-verified-task bake-off** against the incumbent lane on the client's own workload, with retry and failure rates in the denominator. If Grok's list-price advantage survives that math, it earns the volume lane it is priced for. 3. **Grounding by default**: retrieval or search tooling in front of any knowledge-dependent workload, given the unverifiable cutoff situation. 4. **Procurement in parallel with engineering**: the compliance questions (residency, attestations, capacity terms) go to xAI in writing during the pilot, not after it, so a technical win cannot be stranded by an unanswerable questionnaire. This is standard [model engineering](/services/model-engineering) discipline; Grok just makes it visibly necessary. The one-line verdict: Grok 4.3 is a serious, aggressively priced option for cost-dominated agentic workloads, and a portfolio instrument even where it does not win. It is not yet the model you build a compliance-constrained system around on documentation alone. #### Frequently Asked Questions **Q: What does Grok 4.3 cost via the xAI API?** Per xAI's models documentation as of July 2026: $1.25 per million input tokens and $2.50 per million output tokens, with a 1M-token context window and no published long-context surcharge. The grok-4.20 variants (reasoning, non-reasoning, multi-agent) share the same price, and grok-build-0.1 runs $1.00/$2.00 at 256K context. Cached-input and search-tool pricing were not published on the page we verified, so budget those conservatively. **Q: Is Grok 4.3 really cheaper than GPT-5.5 and Claude?** On list price, dramatically: $2.50 per million output tokens versus $30.00 for gpt-5.5 and $15.00 for Sonnet-class Claude models as of July 2026, a 6-12x gap on the token type that dominates agentic workloads. Whether that survives contact with your workload depends on quality: a model that needs more retries to produce a verified result gives back part of its price advantage. We compare cost per completed, verified task rather than per-token rates, and recommend you do the same before migrating anything. **Q: What happened to the original Grok 4?** It no longer appears on xAI's current models page as of July 2026, roughly a year after its mid-2025 launch; the current lineup is Grok 4.3, the grok-4.20 variant family, and grok-build. The takeaway for buyers is cadence: xAI ships and retires models quickly, and we could not verify a published minimum-notice deprecation policy comparable to OpenAI's six-month commitment. Pin dated snapshots, route through a gateway, and keep migration evals ready. **Q: What is Grok 4.3's knowledge cutoff?** xAI's docs state a November 2024 cutoff for Grok 3 and Grok 4, and publish no cutoff for Grok 4.3 on the models page we verified in July 2026. Engineering-wise the answer is to make the question irrelevant: ground any knowledge-dependent workload with retrieval or search tooling and treat parametric knowledge as a convenience, not a source of record. That is our standard practice on every vendor; the unverifiable cutoff here just makes it mandatory. **Q: Is the xAI API enterprise-ready?** The integration surface is genuinely easy: published pricing, simple model list, aliased names for automatic updates. What we could not verify on the public docs is the enterprise packaging around it: data-residency options, compliance attestations, and capacity reservation terms are documented far more thinly than the Azure OpenAI, Bedrock, or Vertex paths. That is a procurement work item rather than a disqualifier: if your regime needs specific guarantees, obtain them from xAI in writing during the pilot phase, not after the architecture is committed. **Q: Where does Grok 4.3 fit in a multi-model architecture?** Two places. First, as the volume lane for output-heavy agentic workloads where its $2.50 output price changes unit economics, provided it wins a cost-per-verified-task bake-off on your workload. Second, as portfolio pressure: even where Grok does not win a lane, having a credible 4-12x cheaper quote in your gateway disciplines pricing conversations with incumbent vendors. In both roles it sits behind a model-agnostic gateway with pinned snapshots, never as a hard dependency. **Q: Does BearPlex deploy Grok in client work?** We evaluate it wherever cost-dominated, output-heavy workloads meet a client risk profile that can absorb a fast-moving vendor: gateway-wrapped, snapshot-pinned, grounded with retrieval, and gated on a cost-per-verified-task bake-off plus tone and refusal-behavior testing against the client's content policies. For compliance-constrained deployments we require written answers from xAI on residency and attestations before it enters the architecture. Whether it ships is decided by those evals and that paperwork, not by the price sheet. --- ### SaulLM-7B: Legal LLM *Publisher: Equall.ai · Paper date: 2024.03.07 · Brief date: 2026.05.08 · 9 min read* *Parameters: 7B · License: MIT · Base: Mistral 7B* URL: https://www.bearplex.com/ai/saullm-7b arXiv: https://arxiv.org/abs/2403.03883 Model card: https://huggingface.co/Equall/Saul-7B-Instruct-v1 GitHub: https://huggingface.co/Equall **Excerpt**: BearPlex's engineering brief on SaulLM-7B, Equall.ai's open-source legal LLM: what it does, where it fits, and how we'd architect privilege-aware deployment. Until SaulLM-7B shipped in March 2024, the only realistic path to a legal LLM in production was either a privilege-leakage-prone GPT-4 deployment, a brittle handful of fine-tuned classifiers stitched together, or a multi-million-dollar bespoke training run. None of those answers fit the constraints we hit when building [Letti AI](/case-studies/letti-ai), and they weren't fitting the law-firm and legal-tech engagements we were scoping either. SaulLM-7B changed the math. ## What it actually is SaulLM-7B is a 7-billion-parameter language model from [Equall.ai](https://huggingface.co/Equall) ([Colombo et al., March 2024](https://arxiv.org/abs/2403.03883)): the first open-source LLM purpose-built for legal text. It takes Mistral 7B as the foundation and applies two transformations: 1. **Continued pretraining** on 30 billion tokens of curated legal text. 2. **Legal instruction fine-tuning** using LegalBench-Instruct, the team's own synthesized dataset of legal-task instructions. The result is released as two checkpoints under the [MIT License](https://huggingface.co/Equall): `Saul-7B-Base` (the continued-pretrain output) and `Saul-7B-Instruct-v1` (the instruction-tuned variant). MIT means we can ship it inside client products without per-token API fees, vendor lock-in, or third-party data handling. ## What's in the training data The pretraining corpus is the part most engineers underestimate. Equall.ai pulled from: - **FreeLaw** (subset of The Pile): 15B tokens - **EDGAR** (SEC corporate filings): 5B tokens - **English MultiLegal Pile** (commercially-licensed subset): 50B tokens - **EuroParl** (parallel proceedings): 6B tokens - **GovInfo Statutes, Opinions & Codes**: 11B tokens - **Law Stack Exchange**: 19M tokens - **EU & UK Legislation**: 505M tokens combined - **Court Transcripts** (CourtListener via Whisper): 350M tokens - **USPTO**: 4.7B tokens - **Commercial Open Australian Legal Corpus**: 0.5B tokens Raw total: 94B tokens. After aggressive filtering, deduplication, and KenLM perplexity-based junk removal: 30B tokens of high-quality legal text. The data spans US, UK, EU, and Australian jurisdictions: important detail if you're deploying outside US-only contexts. ## Where SaulLM-7B fits in a production architecture The honest answer is that you don't ship SaulLM-7B as the sole component of anything. You ship it as one specialized capability inside a broader system. Three patterns from BearPlex engagements: ### Pattern 1: SaulLM-7B as the privilege-tagged understanding layer For e-discovery and contract-review workflows where the source documents include privileged communications, SaulLM-7B runs **inside the firm's VPC** with no egress. It handles document classification, issue spotting, and structured extraction. A retrieval system (typically a hybrid BM25 + vector store with role-based access controls) sits in front of it. Final synthesis routes through a larger model (often Claude or GPT-4) only for non-privileged content surfaces. This pattern lets you keep privileged data inside the firm's infrastructure boundary while still benefiting from frontier-model reasoning where appropriate. ### Pattern 2: Citation-disciplined drafting Bar sanctions for fabricated citations are real: the 2023 *Mata v. Avianca* matter, where attorneys filed a brief with hallucinated case law generated by ChatGPT, was the first widely-publicized incident, but it wasn't the last. We treat citation accuracy as an architectural concern, not a prompting concern. SaulLM-7B in this pattern handles the legal-language generation; a deterministic post-processor extracts every citation; each citation is verified against a structured citation graph (Westlaw or [CourtListener](https://www.courtlistener.com/) APIs) before any draft surfaces to a human. Hallucinated citations are caught and re-prompted in a finite loop. The model never delivers an unverified citation to the user. ### Pattern 3: Multilingual EU work The MultiLegal Pile inclusion gives SaulLM-7B usable proficiency on EU legal text. For clients with cross-border practices, this matters more than the raw English benchmark numbers suggest. ## The benchmarks worth caring about The paper introduces **LegalBench-Instruct**, a refinement of [LegalBench](https://hazyresearch.stanford.edu/legalbench/) that strips distracting few-shot examples and forces the model to generate proper tags rather than verbose explanations. This is a non-trivial methodological contribution: the original LegalBench's verbose-evaluation protocol penalized open models that hadn't been instruction-tuned for tight output. On LegalBench-Instruct and the legal subset of MMLU (international law, professional law, jurisprudence), SaulLM-7B-Instruct outperforms both the base Mistral-7B and Llama-2-7B-chat across nearly every legal task. It does **not** outperform GPT-4 on most tasks: that's not the point. The point is that you can ship it on commodity hardware inside a client's network for a marginal cost approaching zero per query. ## License and deployment MIT is the right license for production deployment. Most "open" legal models in 2024 had restrictive licenses that excluded commercial use; SaulLM-7B is genuinely usable. Hardware footprint: - **`Saul-7B-Instruct-v1`** at FP16: ~14GB GPU memory. Fits on a single A10G or T4-XL. - At INT8 quantization (via `bitsandbytes` or `AWQ`): ~7GB. Comfortably runs on a single A10G. - At INT4 (`AWQ` or `GPTQ`): ~4GB. Fits on a workstation-class GPU. Quality drop is measurable but tolerable for most extraction tasks. For a typical mid-market law-firm deployment, a single A10G inferring at INT8 handles 200-400 concurrent requests at sub-second latency. Cost: roughly $0.02 per 1k tokens at typical AWS spot pricing: **two orders of magnitude cheaper** than an equivalent GPT-4 workload. ## Where we'd actually use it We've evaluated SaulLM-7B in three engagement contexts so far: - **In-house legal AI for a SaaS company**: used as the structured-extraction layer over inbound contracts. Saved roughly 60% of paralegal review time on non-novel agreements. - **A legal-tech product feature**: used as the issue-spotting model in a contract-review SaaS. Replaced a hand-built classifier ensemble. Improved F1 by 8-12% on customer test sets and reduced operational overhead. - **A litigation-support engagement**: used as the document-clustering and brief-extraction layer over a 200,000-document e-discovery production. Privilege tagging stayed inside the firm's VPC. In all three, SaulLM-7B replaced or augmented something: it didn't ship alone. The retrieval system, the citation verifier, the access-control layer, and the human-review loop matter as much as the model. ## What we'd watch in production If you're shipping this, here's what we'd flag from experience: - **Output verbosity**: The instruction-tuned version still skews toward verbose outputs even when you ask for tags. Build structural output enforcement (function calling or constrained decoding) into the inference layer. Don't rely on prompt-only formatting. - **Citation quality**: SaulLM-7B will fabricate citations at a non-zero rate. The architectural fix is verification against an authoritative citation database, not retraining, not better prompts. - **Privilege boundary leakage**: The model itself cannot enforce privilege; it's a stateless function. The architecture around it must enforce it. Audit your retrieval layer and your logging layer with privilege as a first-class concern. - **Updating with new jurisprudence**: SaulLM-7B's training cutoff predates whatever's happened in legal AI in the last 18 months. For practice-area work where recent precedent matters, route those queries through retrieval, not pure generation. ## What's next The Equall.ai team has continued the SaulLM line: `Saul-7B-Instruct-v1` is the publicly-released checkpoint, but you'll see continued iteration on Hugging Face. The methodology generalizes; expect equivalent open models to land for medical, financial, and regulatory verticals over the next 12-18 months. For BearPlex, SaulLM-7B is one of the components we now reach for first when scoping a legal AI engagement. Whether it stays the right answer depends on the workload, but it expanded the design space materially. #### Frequently Asked Questions **Q: Can SaulLM-7B replace GPT-4 for legal tasks?** Not as a drop-in. SaulLM-7B doesn't match GPT-4 on raw reasoning quality, but on specialized legal extraction, classification, and language tasks it's competitive at a fraction of the cost. Most production architectures use both: SaulLM-7B for high-volume extraction inside the firm's network, GPT-4 / Claude for non-privileged synthesis where reasoning quality matters most. **Q: Is SaulLM-7B HIPAA / SOC 2 compliant out of the box?** The model itself isn't a compliance artifact: compliance comes from the deployment architecture around it. Because SaulLM-7B is MIT-licensed and runs on commodity GPUs, you can deploy it inside a HIPAA-compliant or SOC 2-audited environment with no third-party data egress. That's typically what makes it preferable to API-based alternatives for legal work. **Q: How does SaulLM-7B handle attorney-client privilege?** It doesn't, and no model does. Privilege is enforced by the architecture surrounding the model: privilege-tagged retrieval, role-based access control on the document store, audit logging on every inference call, and inference happening inside the firm's network boundary. SaulLM-7B is well-suited to this pattern because it doesn't require external API calls. **Q: What's the inference cost for a typical legal workload?** On AWS spot pricing with a single A10G at INT8 quantization, we measure roughly $0.02 per 1k tokens for typical contract-review workloads. That's roughly 100× cheaper than equivalent GPT-4 token costs, before considering the operational overhead of API calls. **Q: Will SaulLM-7B fabricate citations?** Yes: every LLM does at a non-zero rate, and SaulLM-7B is no exception. The architectural fix is to extract every citation from generated output and verify each against an authoritative source (Westlaw, LexisNexis, CourtListener API) before any draft surfaces to a human. Fabricated citations are caught and re-prompted in a verification loop. This is the same pattern that the post-Mata v. Avianca generation of legal AI tools converged on. **Q: How does SaulLM-7B compare to Llama or Mistral fine-tuned on legal text?** On LegalBench-Instruct and the legal subset of MMLU, SaulLM-7B outperforms both Mistral-7B and Llama-2-7B variants. The continued-pretraining on 30B legal tokens is doing real work: it's not just instruction tuning. If you're considering whether to fine-tune your own model versus use SaulLM-7B, the cost comparison usually favors SaulLM-7B unless you have a very specific corpus the open model wasn't exposed to. **Q: Does BearPlex deploy SaulLM-7B in client work?** Yes: we've evaluated it in three engagements covering in-house legal AI, legal-tech product engineering, and litigation-support work. It's now one of the first components we reach for when scoping legal AI work. We always pair it with retrieval, citation verification, and privilege-aware access control as part of the broader architecture. --- ### Apollo: Multilingual Medical LLM *Publisher: FreedomIntelligence · Paper date: 2024.03.07 · Brief date: 2026.05.08 · 8 min read* *Parameters: 0.5B / 1.8B / 2B / 6B / 7B · License: Apache 2.0* URL: https://www.bearplex.com/ai/apollo arXiv: https://arxiv.org/abs/2403.03640 Model card: https://huggingface.co/FreedomIntelligence GitHub: https://github.com/FreedomIntelligence/Apollo **Excerpt**: BearPlex's engineering brief on Apollo, FreedomIntelligence's multilingual medical LLM: ApolloCorpora, XMedBench, and medical AI beyond English-only. Almost every medical LLM that mattered before March 2024 was English-only. [Med-PaLM 2](https://www.nature.com/articles/s41586-023-06291-2), [PMC-LLaMA](https://arxiv.org/abs/2304.14454), [Clinical-Camel](https://arxiv.org/abs/2305.12031), [BioMedLM](https://arxiv.org/abs/2403.18421): all heavily English-leaning, all underperforming meaningfully on non-English clinical text. For a US-based health system serving an English-speaking patient population, this was tolerable. For everyone else (and quietly, for any US health system serving non-English-speaking patients in their own communities), it was a ceiling. Apollo broke the ceiling. ## What it actually is Apollo is a family of multilingual medical LLMs from [FreedomIntelligence](https://github.com/FreedomIntelligence/Apollo) ([Wang et al., March 2024](https://arxiv.org/abs/2403.03640)). It comes in five sizes: 0.5B, 1.8B, 2B, 6B, and 7B parameters, and supports six languages: English, Chinese, Hindi, Spanish, French, and Arabic. Combined, those six languages cover roughly 6.1 billion native speakers across 132 countries. The model is released under [Apache 2.0](https://github.com/FreedomIntelligence/Apollo) with full open weights, training corpus, and evaluation benchmark. A live demo runs at [apollo.llmzoo.com](https://apollo.llmzoo.com/). ## What's in ApolloCorpora The training corpus (ApolloCorpora) is the most interesting engineering artifact in the paper. It's a 2.5B-token multilingual medical dataset assembled from: - **Medical books** in all 6 languages - **Clinical guidelines** (regional, where available) - **Wikipedia medical articles** in 6 languages - **Medical exams** (MCQA datasets across regions) - **Doctor-patient dialogues** (synthesized + curated) - **Medical research papers** - **Online medical forums and Q&A** The cross-language balance matters. ApolloCorpora isn't English-with-translations: it's natively multilingual, with corpus weights tuned per-language to reflect both speaker population and content availability. That's why Apollo's per-language performance gap is much narrower than equivalent translate-then-fine-tune approaches. ## The benchmarks: XMedBench The paper introduces **XMedBench**: a multilingual medical benchmark created by translating relevant slices of [MMLU](https://github.com/hendrycks/test) into Chinese, Hindi, Spanish, French, and Arabic, plus including native-language medical multiple-choice tasks where they exist. The headline result: **Apollo-7B is the state-of-the-art multilingual medical LLM up to 70B parameters.** Even Apollo-1.8B outperforms much larger general-purpose models on non-English medical tasks. That's the part most engineering teams underestimate when scoping multilingual deployments: the size-vs-domain-specialization tradeoff favors specialization more than the popular "just use a bigger general model" narrative suggests. ## The architectural innovation: proxy tuning The paper's secondary contribution is a proxy-tuning recipe. You can apply Apollo's multilingual medical capabilities to a larger general LLM **without fine-tuning the larger model itself.** The proxy-tuning math is: > output = larger_general_model + (Apollo_tuned - Apollo_base) That's a meaningful shift in the deployment economics. It means a hospital system that's already running, say, a Llama-3-70B internal deployment for general clinical workflows can layer Apollo-7B's multilingual medical capabilities on top without retraining the 70B model. The compute and risk economics are very different. ## Where Apollo fits in a production architecture Three patterns from the multilingual healthcare engagements we've scoped: ### Pattern 1: Multilingual triage with sovereign deployment For health systems serving multilingual patient populations, an Apollo deployment inside the hospital's GPU cluster (PHI never leaves) handles initial triage, intake summarization, and translation-aware clinical NLP. A retrieval system over the patient record sits on top. Apollo's multilingual native capability means the same model serves all language groups without per-language model variants. ### Pattern 2: International health-system rollout For organizations operating across regions (large NGOs, pharmaceutical companies running multi-country trials, telehealth providers expanding into new markets), Apollo provides a unified multilingual baseline. A US team can ship a system that works equivalently for English, Spanish, French, and Arabic patient populations without maintaining six different models. ### Pattern 3: US health system with multilingual staff Less obvious but increasingly important. Many US health systems have clinical staff whose first language is Spanish, Hindi, Tagalog (not directly supported), or Mandarin. Apollo-deployed clinical-note generation and summarization tools that respect the clinical staff's native language (and translate accurately into English for the medical record) measurably improve documentation quality. ## License and deployment Apache 2.0: usable in commercial products, including inside HIPAA-bound deployments. The full Apollo line is on Hugging Face under the [FreedomIntelligence org](https://huggingface.co/FreedomIntelligence). Hardware footprint scales: - **Apollo-0.5B** at FP16: ~1GB. Runs on a workstation CPU at acceptable latency for batch workloads. - **Apollo-1.8B** at FP16: ~3.6GB. Single consumer GPU. - **Apollo-7B** at FP16: ~14GB. Single A10G or T4-XL. INT8 fits on workstation-class GPUs. For a typical hospital deployment, Apollo-7B at INT8 on a single A10G handles 200-300 concurrent inference requests at sub-second latency. Per-query cost approaches zero relative to API-based alternatives. ## Where we'd actually use it Three patterns we've scoped or evaluated in the past 12 months: - **Multilingual ambient scribe** for a hospital system serving primarily Spanish-speaking patient population in the US Southwest. Apollo handles native-Spanish clinical-note generation; downstream English-language EMR integration is straightforward. - **Clinical-guideline translation pipeline** for a global pharmaceutical client running multi-country trials. Apollo handles the medical-terminology-faithful translation that off-the-shelf translation services consistently failed at. - **Cross-language medical Q&A** for a digital-health startup with Hindi and Arabic patient populations alongside English. Apollo's per-language benchmark performance allowed a single model deployment instead of three. ## What we'd watch in production If you're shipping this, here's what we'd flag: - **US clinical terminology** (UMLS, ICD-10, RxNorm) is well-represented but the model is not a HIPAA compliance artifact. Compliance comes from the deployment architecture: sovereign cluster, audit logging, RBAC, etc. - **Hallucination rate is non-zero** and meaningfully higher than in English-only specialist models for non-English medical tasks. Always pair with a retrieval-grounded architecture for clinical-decision-support paths. - **Updating cadence**: medical knowledge moves fast. The Apollo training cutoff predates whatever's happened recently in the field. Build retrieval over current literature for any clinical-decision-support workflow; don't rely on the model's parametric knowledge alone. - **Per-language quality gaps** still exist: French and Spanish tend to be strongest, Hindi and Arabic stronger than expected but somewhat below English. Stress-test on your specific clinical-language workload before committing. ## Why this matters even for English-only health systems The cleanest argument: the methodology Apollo introduced (multilingual continued pretraining + ApolloCorpora composition recipe + proxy tuning) is broadly applicable. Equivalent open multilingual specialist models will land for radiology reporting, mental health, oncology decision support, and other sub-specialties over the next 12-18 months. Apollo is the proof that the recipe works at this size budget: it expanded the design space for medical AI deployment economics meaningfully. For BearPlex's [healthcare model-engineering practice](/services/model-engineering), Apollo is now part of the standard evaluation set when scoping non-English-required deployments. Whether it ships depends on the workload, but it changed our default assumption from "we'll need a bigger model" to "we should evaluate the specialist first." #### Frequently Asked Questions **Q: Is Apollo HIPAA-compliant?** The model itself isn't a compliance artifact: compliance comes from the deployment architecture. Because Apollo is Apache 2.0 with open weights, you can deploy it inside a HIPAA-compliant infrastructure boundary (sovereign GPU cluster, audit logging, RBAC, no PHI egress). That's typically what makes it preferable to API-based alternatives for clinical work. **Q: Can Apollo handle US clinical terminology like UMLS or ICD-10?** Yes: UMLS, ICD-10, and RxNorm coverage is well-represented in the English portion of ApolloCorpora. Performance on US-specific clinical extraction tasks is competitive with English-only specialists. For non-English medical terminology (e.g., mainland Chinese clinical coding standards), per-language benchmarks become the relevant signal. **Q: What's the inference cost on a typical clinical workload?** On AWS spot pricing with a single A10G at INT8 quantization, Apollo-7B costs roughly $0.015 to 0.025 per 1k tokens for clinical-note generation and summarization workloads. That's two orders of magnitude cheaper than API-based alternatives, before factoring in the data-residency benefits. **Q: Which Apollo size should I start with?** Apollo-7B if you have GPU budget and a quality bar: it's the SOTA multilingual medical model in its class. Apollo-1.8B for resource-constrained edge deployments (e.g., on-device clinical tools). Apollo-0.5B is genuinely useful as a draft model for speculative decoding paired with a larger model. Skip the 0.5B and 1.8B for primary-inference clinical workloads where quality matters. **Q: Does Apollo work for languages outside the supported six?** Not natively. The supported languages are English, Chinese, Hindi, Spanish, French, and Arabic. For other languages, you have two paths: (1) translate to one of the supported languages, run Apollo, translate back; quality drops measurably; (2) fine-tune Apollo on your target-language medical corpus using the published methodology. We've done the latter for two engagements; it's tractable but engineering-intensive. **Q: How does Apollo compare to Med-PaLM 2 or Med-Gemini?** Different design points. Med-PaLM 2 and Med-Gemini are closed-weight, English-primary, accessed via API. Apollo is open-weight, multilingual-native, deployable inside a sovereign cluster. On English-only US-clinical benchmarks, the closed models are still ahead on raw quality. On non-English clinical work or any workload requiring full data residency, Apollo is in a class by itself. **Q: Does BearPlex deploy Apollo in client work?** Yes: we've scoped or evaluated Apollo for three engagements covering multilingual ambient-scribe deployment, multi-country clinical-guideline translation, and cross-language patient Q&A. It's part of our standard evaluation set when scoping multilingual healthcare AI work. --- ### GameNGen: Neural Game Engine *Publisher: Google DeepMind · Paper date: 2024.08.27 · Brief date: 2026.05.08 · 7 min read* *Parameters: Stable Diffusion 1.4 base + custom fine-tune · License: Research-only (Google DeepMind) · Base: Stable Diffusion 1.4* URL: https://www.bearplex.com/ai/gamengen-ai-game-development arXiv: https://arxiv.org/abs/2408.14837 Model card: https://gamengen.github.io/ **Excerpt**: BearPlex's engineering brief on GameNGen: Google DeepMind's neural game engine, what it does, why it's hard, and where it transfers beyond gaming. For thirty years, video-game engines have been the canonical example of complex software that AI couldn't replace. State to manage, physics to simulate, rendering to compose, input to react to: all at 60 frames per second, deterministically, with no margin for the kind of generation drift that LLMs are notorious for. The rule was: AI builds tools for game engines, not game engines themselves. That rule has now broken. ## What GameNGen actually does GameNGen ([Valevski et al., Google DeepMind, August 2024](https://arxiv.org/abs/2408.14837)) is a diffusion model fine-tuned to predict the next frame of a video game given the most recent frames and the player's input. Specifically: a fine-tuned [Stable Diffusion 1.4](https://huggingface.co/CompVis/stable-diffusion-v1-4) running at **20 frames per second** producing **playable DOOM**. Not pre-rendered footage. Live, interactive, action-conditioned generation. A human can sit down at the controls, move forward, fire weapons, take damage, navigate levels, and the entire game world is being rendered by a neural network on the fly, with no underlying game engine at all. ## Why this is hard A long-running joke about game engines is that the hard part isn't the rendering: it's *consistency*. If you turn around 360 degrees, the thing in front of you should be exactly what was there before. If you fire a weapon, the bullet's trajectory should reflect physics. If you step into water, you should be wet next frame. The state is enormous and the consistency requirements are unforgiving. GameNGen handles this through three architectural decisions: 1. **Action-conditioned context**: The model receives both recent frames AND the player's input action sequence. The input embedding shapes the noise prediction so the next frame reflects the action. 2. **History buffer**: The last 64 frames feed into the conditioning. This gives the model a working memory of recent state without requiring an explicit state representation. 3. **RL agent for training data**: The training data isn't human gameplay; it's 900 million frames generated by a reinforcement-learning agent playing DOOM. The agent's behavior is engineered to cover the state space evenly, not to play "well" in the human sense. The RL agent is the unsung hero. Generating 900M frames of training data from human gameplay would be infeasible; an RL agent that systematically covers the level geometry, enemy types, weapon usage, and state transitions in DOOM produces uniformly-covered training data at machine speed. ## The benchmarks worth knowing The paper reports three results that matter: - **20 fps on a single TPU**. Below standard 60fps, but well into "playable" territory. - **PSNR of 29.4 dB** on next-frame prediction across a held-out trajectory: comparable to standard lossy video compression. - **Human rater accuracy of 58%** when asked whether a 1.6-second clip is real DOOM or generated. Not perfect, but close enough that humans struggle to tell. The 58% rater accuracy is the headline. We're inside the noise floor of human perception of game-engine output, with a model that has no underlying game-engine code at all. ## Where the architecture transfers The honest BearPlex perspective: we don't currently ship neural game engines. We're tracking GameNGen because the architecture transfers to several adjacent domains where we do ship. ### Action-conditioned simulation The hardest part of any [autonomous-agent](/services/autonomous-agents) deployment is reliable simulation of the agent's environment for evaluation and training. Real environments are slow and expensive; static simulators miss the long tail of edge cases. An action-conditioned diffusion model trained on real environment trajectories can serve as a high-fidelity simulator that matches the real environment's distribution. ### Digital twins for industrial systems Industrial-twin systems (manufacturing, energy, logistics) currently rely on physics-engine-based simulation that's expensive to author and brittle in the face of novel inputs. An action-conditioned generative model trained on telemetry from the real system can serve as a learned twin: useful for what-if analysis, training scenarios, and operator practice. ### Training environments for RL agents The chicken-and-egg problem of RL: you need a high-fidelity simulator to train an RL agent, and historically you needed physics simulation to build a high-fidelity simulator. GameNGen breaks the chicken-and-egg by training the simulator on real-world trajectories from a less-capable agent or human operator. ## Limitations to internalize Three honest caveats: 1. **Memory horizon is short**. The 64-frame context window is enough for moment-to-moment gameplay but doesn't preserve longer-term state (which keys you've collected, which doors you've unlocked, deep level navigation). For applications requiring persistent state, you bolt on an explicit state representation. 2. **Trained on one game, transfers poorly**. The model is fine-tuned for DOOM. Re-purposing it for another game or domain requires substantial retraining with domain-specific RL agent rollouts. There's no zero-shot transfer. 3. **Compute footprint**. 20fps on a TPU, not a consumer GPU. The economics for production deployment of GameNGen-derived systems are still tilted toward enterprise infrastructure, not user-device inference. ## The license question GameNGen is research-only. Google DeepMind has not released model weights or training code as of this writing. For BearPlex client work, this means GameNGen is interesting as an architectural blueprint, not a deployable artifact. Any production deployment of action-conditioned diffusion in 2026 would replicate the architecture from open primitives (a finetuned Stable Diffusion variant, a custom RL data pipeline) rather than use Google's weights directly. ## Why we're tracking this The clearest signal: this paper expanded the design space for what "AI-engineered software" can replace. Five years ago, the answer to "can AI replace a video game engine?" was confidently "no, the consistency requirements are too high." Today, the answer is "yes, at 20fps, with caveats." Three years from now, those caveats will be smaller. For BearPlex, this changes how we scope simulation, twin, and training-environment work. When a client asks whether a learned simulator could substitute for hand-coded physics, the answer was "no" and is now "let's evaluate." We're not building neural game engines for clients today. But the architecture has crossed the line from research curiosity into production-relevant pattern, and we expect to see derivative work over the next 18 months that's genuinely deployable. #### Frequently Asked Questions **Q: Is GameNGen open source?** No. GameNGen is research-only. Google DeepMind has not released model weights or training code. The paper describes the architecture in enough detail that the approach can be replicated with open primitives (a fine-tuned Stable Diffusion variant + a custom RL training data pipeline), but you can't download GameNGen and run it yourself. **Q: Could this replace Unreal or Unity?** Not soon. GameNGen is a single-game neural model that runs at 20fps on enterprise hardware. Modern game engines render diverse content at 60-120fps on consumer GPUs and support persistent state across long play sessions. The interesting question is whether components of game engines (visual fidelity, physics, NPC behavior) get incrementally replaced by neural models inside otherwise-conventional engines. That's already happening in non-trivial ways. **Q: What's the GPU cost for a GameNGen-style deployment?** GameNGen runs at 20fps on a single TPUv5. For non-game applications (digital twins, training simulators), you're looking at one A100 80GB or H100 per concurrent user at acceptable latency. Production economics are still tilted toward enterprise infrastructure, not user-device inference. **Q: How does GameNGen compare to other AI game-generation work?** Earlier work (Genie, World Models) built playable environments but at lower fidelity, lower frame rates, or smaller state spaces. GameNGen is the first to combine commercial-game fidelity, real-time interactive frame rates, and human-grade plausibility on a long-played-out, well-known game. The 58% human-rater confusion rate is the headline metric. **Q: Where does this architecture transfer beyond gaming?** Three high-value patterns: (1) action-conditioned simulation for RL agent training, (2) learned digital twins for industrial systems where physics-engine modeling is expensive, (3) training environments for autonomous-agent evaluation that match the real environment's distribution. None of these are deployable in 2026 without substantial bespoke work, but the architecture is now proven. **Q: Does BearPlex deploy this in client work?** Not currently: we don't ship neural game engines. We're tracking GameNGen because its architectural pattern (action-conditioned diffusion as a learned environment model) transfers to simulation, digital-twin, and training-environment work where we do ship. We expect derivative work to become deployable over the next 18 months and we're scoping accordingly. --- ## Glossary (50 AI engineering terms) Each term has a full explainer page with use cases, production examples, and FAQs at the linked URL. The one-sentence definitions below are the canonical BearPlex definitions. ### RAG RAG (Retrieval Augmented Generation) is an AI architecture that retrieves relevant documents from a knowledge base and injects them into a large language model's context window before generating an answer: grounding responses in source material instead of relying purely on the model's parametric memory. URL: https://www.bearplex.com/glossary/rag ### RLHF RLHF (Reinforcement Learning from Human Feedback) is a post-training technique where human annotators rank model outputs by quality, those rankings train a reward model, and the reward model fine-tunes the base language model via reinforcement learning: aligning the model's behavior with human preferences for helpfulness, honesty, and safety. URL: https://www.bearplex.com/glossary/rlhf ### LoRA LoRA (Low-Rank Adaptation) is a parameter-efficient fine-tuning technique that freezes the pretrained model weights and trains small low-rank matrices injected into the attention layers: achieving 90%+ of full fine-tuning quality at roughly 1% of the GPU memory and training cost, while producing tiny adapter files that can be swapped in and out at inference time. URL: https://www.bearplex.com/glossary/lora ### Fine-tuning Fine-tuning is the process of continuing a pretrained language model's training on a smaller, task-specific dataset to adapt its behavior, style, or capabilities to a particular domain: modifying the model's weights so the adapted version performs better on the target task than the generic base model. URL: https://www.bearplex.com/glossary/fine-tuning ### Embedding An embedding is a numerical vector representation of text, image, or other data (typically a dense array of 384 to 4,096 floating-point numbers) that encodes semantic meaning in a high-dimensional space, where similar concepts are mathematically close to each other and dissimilar concepts are far apart. URL: https://www.bearplex.com/glossary/embedding ### Vector Database A vector database is a specialized database optimized for storing and querying high-dimensional embedding vectors: supporting fast nearest-neighbor search across millions or billions of vectors using approximate nearest-neighbor (ANN) algorithms like HNSW, IVF, or DiskANN, typically with sub-100ms query latency at scale. URL: https://www.bearplex.com/glossary/vector-database ### Agent An agent is an AI system that perceives its environment, makes decisions, and takes actions to achieve goals, typically by using tools, executing multi-step plans, and adapting based on feedback. In modern LLM context, an agent is a language model with the ability to call tools, maintain state across steps, and iterate until a goal is reached, rather than producing a single one-shot response. URL: https://www.bearplex.com/glossary/agent ### AI Agent An AI agent is a language model-powered system that autonomously perceives its context, plans multi-step actions, calls tools or APIs, and iterates toward a goal: distinguished from a chatbot by its ability to take actions in the world (not just respond) and from traditional software by its capacity to reason about novel situations using LLM intelligence. URL: https://www.bearplex.com/glossary/ai-agent ### Multi-Agent System A multi-agent system (MAS) is an AI architecture where multiple specialized agents (each with its own role, tools, and prompt) coordinate to accomplish tasks too complex for a single agent. Common patterns include hierarchical (orchestrator agent delegates to worker agents), conversational (agents debate or negotiate), and pipeline (agents pass work through stages like an assembly line). URL: https://www.bearplex.com/glossary/multi-agent-system ### MCP MCP (Model Context Protocol) is an open standard introduced by Anthropic in late 2024 that defines how AI assistants connect to external data sources and tools. It standardizes the way LLMs request resources, invoke tools, and access prompts from external servers: making AI integrations interoperable across LLM providers, like USB-C for AI tooling. URL: https://www.bearplex.com/glossary/mcp ### Tool Use Tool use (also called function calling) is the capability of an LLM to invoke external functions, APIs, or other systems during inference: allowing the model to retrieve information it doesn't know, perform actions in the real world, or compute results that would be unreliable to generate from training data alone (math, code execution, database queries). URL: https://www.bearplex.com/glossary/tool-use ### Chain-of-Thought Chain-of-Thought (CoT) is a prompting technique that encourages an LLM to articulate its reasoning step-by-step before producing a final answer: significantly improving performance on multi-step problems by making the model 'think out loud' rather than jumping directly to a conclusion. Modern frontier models (Claude Sonnet 4.5, GPT-5, Gemini 2.5) often do this automatically; explicit CoT prompting is most impactful with smaller or older models. URL: https://www.bearplex.com/glossary/chain-of-thought ### Prompt Engineering Prompt engineering is the discipline of designing the inputs (instructions, context, examples, formatting cues) given to a language model to elicit reliable, high-quality outputs. It combines linguistic precision, system understanding, and iterative refinement to turn LLM capabilities into production-grade behavior: the foundational skill underlying every other AI engineering technique. URL: https://www.bearplex.com/glossary/prompt-engineering ### Hallucination Hallucination is the failure mode where a language model generates content that is fluent, confident-sounding, and incorrect: fabricating facts, citations, code, or details that don't exist in reality. Hallucinations occur because LLMs predict statistically likely text, not factually verified text: making confident-but-wrong outputs structurally inevitable without specific defenses like RAG with citation tracking, output validation, or human review. URL: https://www.bearplex.com/glossary/hallucination ### Context Window A context window is the maximum number of tokens an LLM can process in a single inference call (including both the input prompt (system instructions, conversation history, retrieved documents) and the model's generated output), measured in tokens, where roughly 1 token ≈ 0.75 English words. URL: https://www.bearplex.com/glossary/context-window ### Function Calling Function calling is the LLM capability to generate structured JSON output that invokes external functions or APIs, enabling the model to read databases, call services, run computations, and take real-world actions instead of producing only natural-language text. URL: https://www.bearplex.com/glossary/function-calling ### System Prompt A system prompt is a special message sent to an LLM at the start of a conversation that defines the model's persona, capabilities, constraints, and instructions: separate from user messages and treated by the model as higher-priority guidance throughout the conversation. URL: https://www.bearplex.com/glossary/system-prompt ### Temperature Temperature is a numerical parameter (typically 0.0 to 2.0) that controls randomness in LLM output sampling: lower values make the model more deterministic and conservative, higher values make it more creative and varied, with 0 producing nearly identical outputs for identical inputs and 1+ producing genuinely diverse responses. URL: https://www.bearplex.com/glossary/temperature ### Token A token is the fundamental unit of text that LLMs process (typically a word, subword, or punctuation mark) produced by a tokenizer that splits raw text into integer IDs the model can compute on, where roughly 1 token equals 0.75 English words on average. URL: https://www.bearplex.com/glossary/token ### Inference Inference is the process of running a trained machine learning model on new input to generate output: for LLMs specifically, this means taking a prompt, processing it through the model's billions of parameters, and producing the next tokens in the response, with cost and latency determined by model size, input length, and output length. URL: https://www.bearplex.com/glossary/inference ### Transformer The Transformer is a neural network architecture introduced in the 2017 paper 'Attention Is All You Need' that uses self-attention mechanisms instead of recurrence or convolution to process sequences: it's the foundation of every modern large language model including GPT, Claude, Gemini, Llama, and most other frontier AI systems. URL: https://www.bearplex.com/glossary/transformer ### Attention Mechanism The attention mechanism is a neural network component that lets each position in a sequence dynamically weight how much to attend to every other position (computed via query, key, and value matrices) enabling models to capture long-range dependencies and is the core mathematical operation that makes Transformer-based LLMs work. URL: https://www.bearplex.com/glossary/attention-mechanism ### Zero-Shot Learning Zero-shot learning is the ability of an LLM to perform a task without being given any examples (relying entirely on the task description in the prompt and the model's pre-trained knowledge), distinct from few-shot learning where a handful of examples are provided in the prompt. URL: https://www.bearplex.com/glossary/zero-shot-learning ### Few-Shot Learning Few-shot learning is the technique of providing an LLM with a small number of input-output examples (typically 1-10) in the prompt to teach the model the desired pattern, format, or behavior: distinct from zero-shot (no examples) and fine-tuning (training on hundreds of examples). URL: https://www.bearplex.com/glossary/few-shot-learning ### Evaluation Harness An evaluation harness is the automated test infrastructure that measures LLM system quality across a representative set of inputs (combining held-out test datasets, scoring rubrics (LLM-as-judge, programmatic checks, or human review), regression detection, and continuous integration), so prompt and model changes can be measured before they ship to production. URL: https://www.bearplex.com/glossary/evaluation-harness ### Guardrails LLM guardrails are programmatic safety and quality checks applied to model inputs and outputs (including content filters, structured output validation, hallucination detection, prompt injection defense, PII redaction, and topic restrictions) that constrain LLM behavior beyond what prompting alone can guarantee. URL: https://www.bearplex.com/glossary/guardrails ### AI Alignment AI alignment is the discipline of ensuring AI systems behave in ways consistent with human intentions and values: spanning training-time techniques (RLHF, Constitutional AI, DPO) that shape model behavior during creation and deployment-time techniques (prompts, guardrails, evaluation, monitoring) that constrain behavior in production. URL: https://www.bearplex.com/glossary/ai-alignment ### Prompt Injection Prompt injection is an attack technique where adversarial input causes an LLM to ignore its original instructions and follow new instructions embedded in user-provided content, including direct prompt injection (the user types the malicious instruction) and indirect prompt injection (the malicious instruction comes from data the LLM retrieves, like email content, web pages, or document text). URL: https://www.bearplex.com/glossary/prompt-injection ### Semantic Search Semantic search is a retrieval technique that finds documents based on meaning rather than exact keyword match (converting both queries and documents to high-dimensional vector embeddings and ranking documents by vector similarity to the query), enabling search to surface conceptually-related results even when no shared terms exist. URL: https://www.bearplex.com/glossary/semantic-search ### Reranking Reranking is the second-stage scoring of candidate documents in a retrieval pipeline: using a more accurate but slower model (typically a cross-encoder) to reorder the top 50-200 candidates from initial retrieval into a precisely-ranked list of 5-10 documents that go to the LLM or user. URL: https://www.bearplex.com/glossary/reranking ### Chunking Chunking is the preprocessing step in RAG pipelines where source documents are split into smaller passages (typically 200-800 tokens each) before embedding and indexing: enabling fine-grained retrieval that returns just the relevant section of a long document rather than the entire document. URL: https://www.bearplex.com/glossary/chunking ### Knowledge Graph A knowledge graph is a structured representation of entities (people, products, places, concepts) and the relationships between them (typically stored in graph databases as nodes and edges with typed properties) enabling queries that traverse relationships, infer new connections, and provide structured context to AI systems for grounded reasoning. URL: https://www.bearplex.com/glossary/knowledge-graph ### Embedding Model An embedding model is a neural network trained to convert text (or images, audio, or other content) into fixed-size dense numerical vectors that capture semantic meaning: enabling similarity comparisons, retrieval, classification, and clustering by mathematical operations on the resulting vectors rather than the raw content. URL: https://www.bearplex.com/glossary/embedding-model ### Mixture of Experts Mixture of Experts (MoE) is a neural network architecture where each forward pass routes tokens through a small subset of specialized 'expert' subnetworks rather than the entire model, keeping total parameter count high (giving model capacity) while activating only a fraction of parameters per token (saving inference compute). It is used in production by Mixtral, DeepSeek-V2/V3, GPT-4 (rumored), and others. URL: https://www.bearplex.com/glossary/mixture-of-experts ### DPO Direct Preference Optimization (DPO) is a fine-tuning method that aligns language models to human preferences directly from a dataset of preferred vs rejected response pairs (without training a separate reward model or running reinforcement learning) making preference alignment dramatically simpler and cheaper than traditional RLHF. URL: https://www.bearplex.com/glossary/dpo ### Constitutional AI Constitutional AI (CAI) is Anthropic's alignment approach where a language model is trained to critique and revise its own responses according to a written set of principles (a 'constitution'): reducing the need for large amounts of human-labeled preference data and producing more transparent, steerable safety behavior than standard RLHF. URL: https://www.bearplex.com/glossary/constitutional-ai ### PEFT Parameter-Efficient Fine-Tuning (PEFT) is a family of fine-tuning techniques (including LoRA, QLoRA, prefix tuning, and prompt tuning) that update only a small subset of a model's parameters (typically 0.1-3%) while keeping the rest frozen, dramatically reducing memory requirements, training time, and storage cost compared to full fine-tuning. URL: https://www.bearplex.com/glossary/peft ### Quantization Quantization is the process of reducing the numerical precision of a neural network's weights and activations (typically from 16-bit floating point (FP16/BF16) to 8-bit integer (INT8) or 4-bit integer (INT4)) to reduce model size, memory bandwidth, and inference cost, with modest accuracy trade-offs that are usually acceptable for production deployment. URL: https://www.bearplex.com/glossary/quantization ### Distillation Knowledge distillation is a model compression technique where a smaller 'student' model is trained to mimic the outputs of a larger 'teacher' model: producing a compact model that approximates the larger model's capabilities while being much cheaper and faster to deploy. URL: https://www.bearplex.com/glossary/distillation ### ReAct Pattern ReAct is an LLM agent design pattern where the model alternates between Reasoning steps (thinking through what to do next) and Acting steps (calling tools or taking actions): producing an interpretable trace of the agent's decision-making process and dramatically improving task completion compared to direct action without reasoning. URL: https://www.bearplex.com/glossary/react-pattern ### Scaling Laws Scaling laws are empirical mathematical relationships discovered in deep learning research showing that model performance improves predictably as a power-law function of model size, dataset size, and compute: establishing the theoretical foundation for the modern frontier-model arms race and explaining why investing in larger models, more data, and more compute yields predictable capability gains. URL: https://www.bearplex.com/glossary/scaling-laws ### AI Safety AI safety is the multidisciplinary field focused on building AI systems that don't cause harm: spanning technical alignment research (making models do what we want), robustness (making models behave well on novel inputs), interpretability (understanding what models learn), governance (policies and norms for AI development), and existential safety (concerns about future AI systems whose capabilities exceed human oversight). URL: https://www.bearplex.com/glossary/ai-safety ### Structured Output Structured output is the LLM capability to generate output that conforms to a specific schema (typically JSON matching a defined structure with typed fields) enabling reliable parsing, validation, and downstream use without the brittleness of extracting structure from natural language. URL: https://www.bearplex.com/glossary/structured-output ### FlashAttention FlashAttention is an exact attention algorithm by Tri Dao that computes the same mathematical result as standard attention but with much better GPU memory bandwidth utilization (typically 2-4× faster on long sequences and dramatically lower memory peak) now standard in nearly every production LLM inference engine. URL: https://www.bearplex.com/glossary/flash-attention ### Tree of Thoughts Tree of Thoughts (ToT) is an LLM reasoning technique that explores multiple reasoning paths in parallel (branching at each reasoning step to consider alternatives, evaluating partial solutions, and selecting the best path), extending chain-of-thought from sequential reasoning to deliberate search over a reasoning tree. URL: https://www.bearplex.com/glossary/tree-of-thoughts ### KV Cache The KV Cache (key-value cache) is a memory structure used during LLM inference that stores the computed key and value matrices from previous tokens: enabling each new token to be generated by attending to all prior context without recomputing the attention for tokens already seen, dramatically reducing per-token compute cost during autoregressive generation. URL: https://www.bearplex.com/glossary/kv-cache ### Speculative Decoding Speculative decoding is an LLM inference optimization where a small fast 'draft model' proposes multiple tokens, the large target model verifies them in a single forward pass, and accepted tokens are returned, typically achieving 1.5-3× speedup on token generation with no quality loss compared to running the target model alone. URL: https://www.bearplex.com/glossary/speculative-decoding ### Self-Consistency Self-consistency is an LLM reasoning technique where multiple chain-of-thought reasoning paths are sampled at temperature > 0 for the same problem, then the most common final answer across paths is selected: improving accuracy on math, logic, and reasoning tasks by 5-15% over single chain-of-thought at the cost of generating multiple reasoning chains per problem. URL: https://www.bearplex.com/glossary/self-consistency ### Dataset Curation Dataset curation is the discipline of constructing, cleaning, labeling, and maintaining the datasets used to train, fine-tune, evaluate, and align AI systems (encompassing data collection, quality control, annotation, deduplication, bias mitigation, and ongoing maintenance) and is consistently the most important and most under-invested factor in production AI quality. URL: https://www.bearplex.com/glossary/dataset-curation ### AI Observability AI observability is the practice of instrumenting and monitoring production AI systems to understand their behavior, detect issues, and continuously improve them, including request / response tracing, evaluation against golden datasets, drift detection, cost tracking, latency monitoring, and the production AI ops infrastructure that distinguishes mature deployments from prototypes that escaped into production. URL: https://www.bearplex.com/glossary/ai-observability --- ## Technology Comparisons (43) Decision-framework pages comparing the technologies BearPlex builds with. Each page carries a full comparison table, scenario recommendations, and FAQs; the TL;DR verdicts below are the canonical summaries. ### RAG vs Fine-Tuning: Which to Choose in 2026 Choose RAG when your knowledge changes frequently, when you need source citations, or when you have role-based access controls, which describes the majority of enterprise AI use cases. Choose fine-tuning when you need to change the model's style or behavior consistently and you have 1,000+ high-quality training examples. In production, the strongest enterprise systems use BOTH: fine-tune for tone and format consistency, RAG for grounded facts. URL: https://www.bearplex.com/compare/rag-vs-fine-tuning --- ### Build vs Buy AI: Enterprise Decision Framework for 2026 Buy when the AI capability is commoditized and not strategic to your differentiation (general-purpose chatbots, off-the-shelf transcription, generic copilots). Build when AI is core to your competitive moat, when your data is unique enough to require custom modeling, or when sovereign data handling and deep workflow integration aren't negotiable. The hybrid path (buy the foundation, build on top) wins more often than either pure approach. Most enterprise AI failures come from over-buying (generic vendor tools that don't fit) or over-building (rebuilding commodity capabilities from scratch). URL: https://www.bearplex.com/compare/build-vs-buy-ai --- ### LangChain vs LangGraph: Which to Choose in 2026 Since the joint 1.0 releases on October 22, 2025, this stopped being a framework rivalry: LangChain's create_agent now executes on the LangGraph runtime, so you are choosing an abstraction level within one stack, not picking a side. Use LangChain (1.3.11 as of July 2026) when a standard tool-calling agent loop plus middleware (human approval, summarization, PII redaction) covers the job. Drop to LangGraph (1.2.8) directly when you need explicit control: custom state schemas, durable execution that survives restarts, human-in-the-loop interrupts at arbitrary points, or multi-agent topologies. Both are MIT-licensed with a stated commitment of no breaking changes until 2.0. Our production default: start with create_agent, and because it is LangGraph underneath, dropping down later is a refactor, not a rewrite. URL: https://www.bearplex.com/compare/langchain-vs-langgraph --- ### OpenAI vs Anthropic: Which to Choose in 2026 Both OpenAI (GPT-4o, GPT-5, o-series reasoning models) and Anthropic (Claude 3.5/4 Sonnet, Opus, Haiku) are frontier-class options viable for nearly any production AI workload. OpenAI leads on ecosystem maturity, lowest-cost endpoints, and image generation; Anthropic leads on long-context reliability, agentic workflows, code generation, prompt caching economics, and safety transparency. For most production engagements we recommend benchmarking BOTH on your specific task: quality differences are real but task-dependent. For agentic and long-context workloads we default to Claude; for image-heavy or cost-sensitive workloads we lean OpenAI. The strongest production architecture is provider-portable code that lets you switch (or A/B test) between them. URL: https://www.bearplex.com/compare/openai-vs-anthropic --- ### Pinecone vs Qdrant: Which Vector Database to Choose in 2026 Use Pinecone if you want a managed vector database with zero operational burden, accept vendor lock-in, and operate at small-to-medium scale (under 30M vectors). Use Qdrant if you need self-hosted deployment for sovereignty / data residency requirements, want lower cost at large scale (30M+ vectors), or want to avoid vendor lock-in. Both are production-quality choices; for most BearPlex engagements the choice comes down to deployment requirements, not technical merits: Qdrant wins for sovereign / on-prem; Pinecone wins for managed simplicity at small-to-medium scale. URL: https://www.bearplex.com/compare/pinecone-vs-qdrant --- ### LoRA vs Full Fine-Tuning: Which to Choose in 2026 Default to LoRA for production fine-tuning in 2026 and treat full fine-tuning as a deliberate exception, not a gold standard you are settling below. Thinking Machines' September 2025 research showed LoRA matches full fine-tuning on typical post-training datasets when adapters are applied to every layer (not just attention) with a learning rate about 10x the full fine-tuning optimum, and it matches even at rank 1 for reinforcement learning. Full fine-tuning still wins in three regimes: training data that exceeds adapter capacity (continued pretraining on billions of domain tokens), very large batch sizes, and repeated sequential fine-tuning of the same model, where LoRA's 'intruder dimensions' accumulate. The economics land hard on LoRA's side: at July 2026 on-demand rates (about $4 per H100 GPU-hour), an adapter experiment costs tens of dollars while a multi-day full fine-tune on an 8x H100 node runs into the thousands. URL: https://www.bearplex.com/compare/lora-vs-full-fine-tuning --- ### Self-Hosted vs Managed LLM: Which to Choose in 2026 Use managed LLMs (Anthropic API, OpenAI, AWS Bedrock, Vertex AI) for the first 6-18 months of any AI initiative: the operational simplicity is dramatic. Switch to self-hosted (open-source models on vLLM / TGI / SageMaker) only when (a) data residency / sovereignty requires it, (b) cost crosses a threshold where self-hosted economics dominate (typically 1M+ requests/month), or (c) you need customization that managed APIs don't support. The hybrid path (managed for some workloads, self-hosted for others) wins more often than either pure approach. URL: https://www.bearplex.com/compare/self-hosted-vs-managed-llm --- ### DPO vs RLHF: Which Alignment Method to Choose in 2026 Use DPO (or its variants ORPO, KTO, SimPO) for 90%+ of preference-tuning use cases: much simpler, much cheaper, comparable results on most tasks. Use full RLHF only when (a) you're at the frontier of capability where the last 1-3% quality matters, (b) you have abundant high-quality preference data and the infrastructure to run reward model + PPO at scale, or (c) you need specific properties RLHF provides that DPO doesn't. The default for production preference-tuning has shifted to DPO; RLHF is the exception, not the rule. URL: https://www.bearplex.com/compare/dpo-vs-rlhf --- ### LangGraph vs CrewAI vs AutoGen: Which Agent Framework to Choose Use LangGraph for production agent systems requiring explicit state management, human-in-the-loop checkpoints, and reliable debugging: our default for production work. Use CrewAI for quick multi-agent prototypes with role-based design where execution speed matters more than production maturity. Use AutoGen for research-heavy work where you're exploring novel multi-agent patterns. For Claude-committed production work, Claude Agent SDK is competitive with LangGraph. For most BearPlex client engagements requiring production reliability and operational maturity, LangGraph wins. URL: https://www.bearplex.com/compare/langgraph-vs-crewai-vs-autogen --- ### Snowflake vs Databricks: Which to Choose in 2026 Snowflake and Databricks spent 2025 and 2026 converging on each other's territory: each now sells a lakehouse, a managed Postgres (Databricks Lakebase went GA on February 3, 2026; Snowflake Postgres followed on February 24, 2026), Apache Iceberg support, and a full agent platform (Agent Bricks vs Cortex AI with CoWork and CoCo). So the decision is no longer a feature checklist; it is about workload center of gravity and team skills. Choose Snowflake when SQL analytics, BI concurrency, and governed data sharing dominate and you want the lowest operational overhead a small data team can run. Choose Databricks when data engineering, ML training, and agentic AI products dominate and your team is fluent in Python and Spark. Large enterprises increasingly run both, with Iceberg as the neutral storage layer between them. URL: https://www.bearplex.com/compare/snowflake-vs-databricks --- ### Fine-Tuning vs Prompt Engineering: Which to Choose in 2026 Start with prompt engineering for nearly every LLM use case in 2026, and treat fine-tuning as a deliberate second step, not a default. Two things changed the math this year: prompt caching now prices cached input at 10 percent of the normal rate on both OpenAI and Anthropic, which guts the old 'long prompts are expensive' argument, and the managed fine-tuning path has narrowed sharply, with OpenAI winding down its self-serve fine-tuning platform (closed to new customers since May 7, 2026, and closed to everyone for new training jobs on January 6, 2027). In practice, fine-tuning in 2026 means tuning open-weight models (Qwen, Llama, Gemma, Mistral) with LoRA or QLoRA on infrastructure you control: a bigger commitment, but the resulting model cannot be deprecated out from under you. Reach for it when a rigorously evaluated prompt still misses your quality bar, when unit economics at millions of requests per month favor a small specialized model, or when you need format and voice compliance that instructions in context cannot hold. Most production systems we ship are prompts plus RAG, with fine-tuning reserved for the narrow, high-volume slices that clearly justify it. URL: https://www.bearplex.com/compare/fine-tuning-vs-prompt-engineering --- ### Multi-Agent vs Single-Agent AI Systems: Which to Build in 2026 Default to a single agent. The 2024 reflex of reaching for agent crews has aged badly, and the two most-cited engineering write-ups on this question, Anthropic's multi-agent research system and Cognition's 'Don't Build Multi-Agents' (both June 2025), agree more than they disagree: multi-agent wins on breadth-first, parallelizable, high-value work (Anthropic measured a 90.2% lift over a single-agent Claude Opus 4 baseline on their internal research eval) and loses on tightly coupled work like coding, while burning about 15x chat-level tokens versus about 4x for a single agent. Meanwhile frontier models keep raising the single-agent ceiling: Claude Sonnet 5 (released June 30, 2026) plans, drives browsers and terminals, and runs autonomously at a $3 input / $15 output per million tokens standard rate ($2 / $10 introductory through August 31, 2026). Go multi-agent only when the task exceeds one context window, splits into genuinely independent sub-questions, and carries enough value to absorb roughly 4x the token spend. URL: https://www.bearplex.com/compare/multi-agent-vs-single-agent --- ### Azure OpenAI vs AWS Bedrock: Which Cloud AI Platform to Choose Use Azure OpenAI when you're committed to the Microsoft / Azure stack, want OpenAI models with enterprise BAA / compliance, and have predominantly Microsoft-stack engineering. Use AWS Bedrock when you're on AWS, want multi-vendor model access (Anthropic Claude, Meta Llama, Mistral, Cohere, Stability AI) through one API, or need flexibility across model providers. Both are competitive enterprise AI platforms; the choice usually comes down to your existing cloud commitments and whether you prefer Azure's OpenAI-only depth or AWS Bedrock's multi-vendor breadth. URL: https://www.bearplex.com/compare/azure-openai-vs-aws-bedrock --- ### Open-Source vs Closed-Source LLMs: Which to Use in 2026 Use closed-source frontier models (GPT-5, Claude Sonnet / Opus, Gemini 2.5) when you want best-in-class quality without operating infrastructure, accept vendor lock-in, and operate at scale where managed pricing is acceptable. Use open-source models (Llama 3.3, Qwen 2.5, DeepSeek-V3, Mistral) when you need sovereign deployment, want lower per-call cost at scale, need to fine-tune or customize, or want vendor independence. The hybrid path (closed-source for highest-quality use cases, open-source for cost-optimized workloads) wins more often than either pure approach. Open-source has caught up dramatically; for most production tasks, frontier open-source is competitive with frontier closed-source. URL: https://www.bearplex.com/compare/open-source-vs-closed-source-llm --- ### Promptfoo vs Braintrust vs LangSmith: Which LLM Eval Tool in 2026 The right answer changed in 2026. Promptfoo (MIT open source, acquisition by OpenAI announced March 9, 2026 with a public commitment to stay open source and model-agnostic) is the pick for pre-ship evals, CI regression gates, and automated red teaming; it is free and it still ships multiple releases a month. Braintrust (flat $249/month Pro, unlimited seats on every tier) is the pick when production tracing, online scoring, and dataset curation by mixed engineering and product teams is the job. LangSmith ($39 per seat/month Plus) is the pick when you are on LangChain or LangGraph, both of which hit stable 1.0 in October 2025: nothing traces those graphs as natively. Most production teams we work with run two of the three: Promptfoo in CI plus one observability platform. Running zero of them is the only wrong answer. URL: https://www.bearplex.com/compare/promptfoo-vs-braintrust-vs-langsmith --- ### LangChain vs LlamaIndex: Which RAG Framework to Choose in 2026 Use LlamaIndex for document-heavy RAG where ingestion / indexing / retrieval depth matters: our default for production RAG over diverse document types. Use LangChain for broader LLM application work, agent systems, and mixed RAG + non-RAG workloads. They're complementary, not competitive: many production engagements use both (LlamaIndex for retrieval, LangChain / LangGraph for orchestration). For pure RAG use cases, LlamaIndex's depth wins; for mixed use cases, LangChain's breadth wins. URL: https://www.bearplex.com/compare/langchain-vs-llamaindex --- ### AI Agents vs RPA: Which Automation Approach to Choose in 2026 Use RPA (UiPath, Automation Anywhere, Blue Prism) for high-volume rule-based automation of repetitive structured workflows where the process is well-defined and rarely changes. Use AI agents for workflows that require natural-language understanding, judgment under ambiguity, handling exceptions, or rapid iteration. The hybrid path (RPA for structured rules-based automation, AI agents for the parts requiring judgment) wins more often than either pure approach. Many enterprise automation initiatives in 2026 are migrating from pure RPA to RPA + AI agent hybrid. URL: https://www.bearplex.com/compare/ai-agents-vs-rpa --- ### MLflow vs Weights & Biases: Which MLOps Platform to Choose Use MLflow for production model registry, deployment, and lifecycle management: open-source, enterprise-friendly, integrates with Databricks and standard MLOps stacks. Use Weights & Biases (W&B) for experiment tracking, dataset versioning, and ML research workflows: polished UX, strong collaboration features, paid platform. Many production ML organizations use both: MLflow for model registry and deployment, W&B for experimentation and team collaboration. The right choice depends on whether your priority is production ops (MLflow) or research / experimentation (W&B). URL: https://www.bearplex.com/compare/mlflow-vs-weights-and-biases --- ### Semantic vs Hybrid Search: Which Retrieval Approach to Choose Use hybrid search (semantic + keyword) for almost every production RAG and search use case: combines the meaning understanding of semantic search with the exact-match precision of keyword search. Use pure semantic search only for use cases with minimal proper noun content where keyword precision doesn't matter. Use pure keyword search rarely: only when meaning understanding is genuinely irrelevant. Modern production retrieval is hybrid by default; the 30-40 lines of code it takes to add keyword search to a semantic pipeline is one of the highest-ROI improvements available. URL: https://www.bearplex.com/compare/semantic-search-vs-hybrid-search --- ### OpenAI vs Cohere vs Voyage: Which Embedding Model to Choose Use OpenAI text-embedding-3 (large or small) for general-purpose production retrieval: strong quality, well-supported, reasonable cost, the default choice for most BearPlex engagements. Use Cohere Embed v3 for multilingual workloads or when you want native reranking integration. Use Voyage AI for domain-specific work where their domain-tuned models (voyage-code, voyage-finance, voyage-law) outperform general-purpose models. Use open-source (BGE, E5) for self-hosted requirements. Quality differences between top embedding models on most production tasks are 1-5%: choose based on operational fit (cost, multilingual, sovereignty) rather than chasing benchmark differences. URL: https://www.bearplex.com/compare/embedding-models-comparison --- ### Toptal vs a Dedicated Agency Team: Which to Choose in 2026 Choose Toptal when you need one vetted senior specialist quickly, you already have engineering management in place, and the engagement is measured in weeks or a few months. Choose a dedicated agency team when you need an outcome delivered rather than a seat filled: a pod that brings its own delivery process, covers multiple roles (engineering, QA, design, project management), and stays accountable for a roadmap over quarters. The honest framing: Toptal sells excellent individuals, an agency sells a functioning team. Match the model to whoever will own delivery. If that person is you, Toptal works well. If you need the vendor to own it, use an agency pod. BearPlex is one instance of the dedicated-team model; Toptal is the strongest brand in the freelance-network model. URL: https://www.bearplex.com/compare/toptal-vs-dedicated-agency --- ### Turing vs a Dedicated Agency Team: Which to Choose in 2026 Choose Turing when you want individual full-time remote engineers at rates typically estimated below US onshore fully loaded cost, you have management capacity to direct them, and long-term individual seats are the shape of your need. Choose a dedicated agency team when you need a vendor to own delivery of a roadmap with a coordinated multi-role pod. One important 2026 context point: Turing has publicly repositioned around AI-lab work (training data, model evaluation, enterprise AI), with developer staffing as one line among several, so evaluate the staffing product on current service quality, not 2021-era marketing. As with all these comparisons, the models differ more than the brands: Turing sells vetted individuals, an agency sells an accountable team. URL: https://www.bearplex.com/compare/turing-vs-dedicated-agency --- ### Lemon.io vs a Dedicated Agency Team: Which to Choose in 2026 Choose Lemon.io when you are an early-stage startup that needs one or two affordable, vetted senior developers fast, month to month, and you can direct their work yourself. Its niche is speed and price: 24-hour average matching from a pool of 1,500+ vetted developers in Europe, Latin America, the US, and Canada. Choose a dedicated agency team when the job is a whole product or workstream that needs engineering, QA, design, and project management operating as one accountable unit over quarters. Lemon.io is arguably the best-fit freelance network for startup budgets; it is simply a different product than vendor-owned delivery. URL: https://www.bearplex.com/compare/lemon-io-vs-dedicated-agency --- ### Andela vs a Dedicated Agency Team: Which to Choose in 2026 Choose Andela when you are an enterprise embedding individual vetted engineers (increasingly AI-focused ones) into squads you already run, you can absorb reported 12-month minimum terms, and global time-zone distribution suits you. Choose a dedicated agency team when you need a vendor accountable for delivering a roadmap with a coordinated pod, transparent scoping, and smaller initial commitments. The two models overlap more than the other network comparisons here: Andela has moved up-market into blended teams and AI services, positioning itself in 2026 as 'the human layer powering production AI'. The practical differences are commitment shape, pricing transparency, and who owns delivery. URL: https://www.bearplex.com/compare/andela-vs-dedicated-agency --- ### Freelancers vs an Agency Team: Which to Choose for Software Development in 2026 Choose freelancers when the scope is small and well-bounded, the budget is tight, you can technically direct the work yourself, and continuity risk is acceptable. Choose an agency team when the work spans multiple disciplines, will run for quarters, must survive any individual leaving, or needs a vendor who is accountable for the outcome rather than the hours. The decision is really about three questions: who owns delivery, who absorbs turnover, and how many skills the work needs at once. One senior freelancer under a capable CTO is often the highest-value engineering money can buy; three freelancers coordinating themselves on a product build is usually where the model breaks. URL: https://www.bearplex.com/compare/freelancers-vs-agency --- ### In-House vs Outsourced Development: Which to Choose in 2026 Build in-house when the software is your core product and competitive moat, the horizon is measured in years, and you can win the hiring market for the skills you need. Outsource when speed matters more than headcount growth, when the skills are scarce or temporarily needed, or when the work is important but not the differentiator you must own culturally. The strongest engineering organizations in 2026 run a deliberate hybrid: a lean in-house core that owns architecture, product judgment, and vendor management, extended by external teams for capacity, specialties, and parallel workstreams. The failure modes are symmetric: all-in-house teams move at hiring speed, and all-outsourced companies wake up owning a product nobody inside understands. URL: https://www.bearplex.com/compare/in-house-vs-outsourced-development --- ### Offshore vs Nearshore vs Onshore Development: Which to Choose in 2026 Choose offshore when cost efficiency and access to deep global talent pools matter most and your delivery process is strong enough to work across large time-zone gaps (or your partner guarantees overlap hours). Choose nearshore when you want meaningful savings while keeping most of the working day shared. Choose onshore when regulatory constraints, on-site presence, or same-market context genuinely require it, and the budget supports it. The dirty secret of this debate in 2026: process quality and partner seniority predict outcomes far better than geography does. A disciplined offshore team with guaranteed overlap hours outperforms a mediocre onshore one at a fraction of the cost, and a bad partner fails in every time zone. URL: https://www.bearplex.com/compare/offshore-vs-nearshore-vs-onshore --- ### Staff Augmentation vs Dedicated Team: Which Model to Choose in 2026 Choose staff augmentation when you have strong engineering management and defined processes, and simply need more hands inside your existing structure: the augmented engineers report into your leads and work your backlog. Choose a dedicated team when you need a self-sufficient unit that brings its own delivery management and owns a workstream end to end, accountable for outcomes rather than hours. The models answer different questions: augmentation answers 'we need more capacity in our machine', a dedicated team answers 'we need another machine'. Most vendor disappointment traces to buying one when the situation called for the other. URL: https://www.bearplex.com/compare/staff-augmentation-vs-dedicated-team --- ### Hiring on Upwork vs an Agency: Which to Choose in 2026 Choose Upwork when the task is small, well-specified, and severable: a script, a fix, a bounded feature, a short specialist engagement, especially at budgets no agency can serve. The marketplace's scale is unmatched and its escrow and review mechanics de-risk small bets. Choose an agency when the work is a product or a roadmap: multi-skill, multi-month, and expensive to restart if a contractor disappears. The honest boundary is stakes times duration. Upwork is a spectacular market for buying tasks; it is a risky way to buy a product. Agencies are overkill for tasks and built for products. URL: https://www.bearplex.com/compare/upwork-vs-agency --- ### Accenture vs a Boutique Agency: Which to Choose in 2026 Choose a global consultancy like Accenture when the program is genuinely enormous: multi-year, multi-country, spanning strategy, operations, and technology, with board-level risk cover and procurement mandates that require a vendor of that scale. Choose a boutique agency when the work is a defined build or capability, when senior-practitioner attention matters more than organizational breadth, and when you want your budget buying engineering rather than coordination layers. These are different instruments, not competitors on most deals: Accenture (roughly 799,000 people and $69.7 billion in FY25 revenue, per its own fact sheet) is built for transformation programs; a boutique is built for shipping. The expensive mistake in both directions is buying scale you do not need, or boutique intimacy for a program that genuinely needs an army. URL: https://www.bearplex.com/compare/accenture-vs-boutique-agency --- ### Fixed Price vs Time and Materials: Which Contract Model to Choose in 2026 Choose fixed price when scope is genuinely known, stable, and specifiable in advance: migrations with defined endpoints, well-understood builds, compliance deliverables. You buy budget certainty and transfer estimation risk to the vendor, and you pay a risk premium for it. Choose time and materials when the work involves discovery, iteration, or changing requirements, which describes most product development and nearly all AI work in 2026. The mature answer on real projects is usually a hybrid: a fixed-price discovery phase that buys certainty about the unknowns, followed by capped T&M delivery with strong governance. The contract model does not remove uncertainty; it only decides who carries it and at what markup. URL: https://www.bearplex.com/compare/fixed-price-vs-time-and-materials --- ### AI Development Agency vs Generalist Agency: Which to Choose in 2026 Choose an AI development agency when the AI system IS the deliverable: RAG over enterprise knowledge, agent workflows, model fine-tuning, or anything where accuracy, evaluation, and cost engineering determine success. The specialist disciplines (golden datasets, evaluation harnesses, retrieval engineering, guardrails) are learned in production and rarely exist at firms that added 'AI' to their services page in 2023. Choose a generalist agency when the product is fundamentally a web or mobile application with a modest AI feature inside it, or when an existing trusted partner's context outweighs specialist depth for a small AI surface. The test that cuts through marketing: ask any agency how they evaluate AI system quality before shipping. Specialists answer with specifics; generalists answer with adjectives. URL: https://www.bearplex.com/compare/ai-development-agency-vs-generalist-agency --- ### Off-the-shelf SaaS vs Building Your Own: Which to Choose in 2026 For most teams, buying the SaaS is the right call. If your workflow is standard, your headcount is modest, and no unusual compliance constraint applies, an off-the-shelf product is cheaper, faster, and better maintained than anything you could commission, and you should go buy it. Building wins in four specific situations: the seat-price math crosses over (per-seat fees times headcount times 36 months clearly exceeds the cost of building and maintaining your own), the workflow is the product (the process you run is your differentiation, and renting the generic version caps it), compliance or data sovereignty rules keep your data out of a vendor cloud, or an integration duct-tape audit shows you already pay for five tools plus Zapier plus spreadsheets to fake one system. Run the 3-year math before deciding either way; the numbers in the table below are computed from vendor pricing pages checked in July 2026. URL: https://www.bearplex.com/compare/build-vs-buy-software --- ### Salesforce vs Building Your Own: Which to Choose in 2026 For most teams, buy Salesforce (or a cheaper CRM) and move on: if your pipeline looks like leads, opportunities, and quotes, a mature platform your ops person can run beats a custom build on speed, risk, and often on cost. The build case is narrow and specific: your workflow IS the product (intake, compliance, scheduling, billing flows that never mapped to Salesforce's object model), your seat count makes per-user math absurd (50 Enterprise seats is $315,000 in licenses alone over 3 years at the current $175 per user per month), or you are already paying more in admin and Apex customization than in licenses just to bend the platform. A custom CRM or ops system typically runs $25,000-$70,000 to build plus a monthly care plan, and its cost does not grow when your headcount does. Decide on workflow fit first and seat math second; sticker price alone is the wrong tiebreaker. URL: https://www.bearplex.com/compare/salesforce-vs-custom-crm --- ### HubSpot vs Building Your Own: Which to Choose in 2026 Most teams should just use HubSpot. If you run a standard pipeline with ten or fewer sales seats and a marketing list in the low thousands, nothing you can build will beat a CRM that is free for up to 2 users and working the same day, with Starter seats at a $20 per month list price (promoted at $7 per seat for new customers as of July 2026). The build case appears at three specific triggers: your marketing contact list is large enough that tier pricing compounds (Marketing Hub Professional includes only 2,000 marketing contacts at $800 per month billed annually, and additional contacts start at $250 per month per 5,000), your business runs on entities and relationships HubSpot's object model cannot express (custom objects are Enterprise-only and capped at 10), or you are paying an Enterprise-stack bill of roughly $248,000 over three years for a 20-seat sales plus marketing setup at list prices while using a fraction of the features. A custom CRM typically costs $25,000 to $70,000 to build plus a monthly care plan, flat with respect to seats and contacts. Buy first; build when the bill or the data model stops fitting. URL: https://www.bearplex.com/compare/hubspot-vs-custom-crm --- ### Asana (and Monday.com) vs Building Your Own: Which to Choose in 2026 If your team needs a tool to track its own tasks and projects, buy Asana or monday.com and move on. At verified July 2026 pricing, a 15-person team on Asana Advanced spends roughly $13,500 over three years, and the same team on monday.com Pro spends about $10,260. No serious custom build beats that, and the vendors have spent a decade polishing views, mobile apps, and integrations you would have to rebuild. Building only makes sense when the workflow IS the business: a branded client portal your customers log into, a production pipeline wired to inventory and stage gates, field operations with scheduling and compliance evidence, or per-seat licensing that has crept past a hundred users and unlimited external collaborators. In those cases a custom ops system (from $15,000, typically $25,000-$70,000) is flat with respect to headcount and fits the workflow exactly instead of bending the workflow to the tool. URL: https://www.bearplex.com/compare/asana-vs-custom-project-management --- ### Retool vs Building Your Own: Which to Choose in 2026 For most teams, Retool is the right call and you should just use it: if your internal tool is CRUD screens, admin panels, and approval flows for a team of 5 to 30 people, Retool's editor gets you there in days for $5 to $50 per user per month, and up to 5 users it is free. Build custom when one of four triggers fires: seat counts push per-user fees past custom-build economics (roughly 75-100+ seats on the Business tier), the tool needs to face customers or partners at volume, your requirements have outgrown the component model and every change is another JavaScript workaround, or compliance requires owning the stack outright. The honest middle path most teams should take: start on Retool, keep your business logic in your database and APIs rather than in Retool queries, and graduate the one or two tools that outgrow it. BearPlex builds the graduation path: custom internal tools typically run $25,000-$70,000, which is a real number to weigh against three years of seat fees, not a reflex answer. URL: https://www.bearplex.com/compare/retool-vs-custom-internal-tools --- ### Zendesk vs Building Your Own: Which to Choose in 2026 If what you need is a help desk (agents answering tickets across email, chat, and a help center), buy Zendesk. It is mature software, it deploys in days, and at typical team sizes it is cheaper than any custom build: 5 agents on Suite Team run about $9,900 over three years at live pricing. The build case is narrow and it is not about replacing ticketing. Build when the portal is part of your product: customers need to see orders, invoices, entitlements, contracts, or job status alongside support, when the experience must live inside your app rather than on a themed help center, or when per-agent fees plus add-on stacking (Copilot $50, Workforce Engagement $50, Contact Center $83, each per agent per month on top of the base seat) outgrow the cost of software you would own. The strongest pattern for many B2B teams is hybrid: keep Zendesk as the ticketing engine and build a custom portal on its APIs. URL: https://www.bearplex.com/compare/zendesk-vs-custom-support-portal --- ### Shopify vs Building Your Own: Which to Choose in 2026 For most merchants, Shopify is the right call, full stop. At $19-$299 per month on annual billing (live shopify.com pricing, July 2026) you get hosting, PCI compliance, a proven checkout, and an app ecosystem that would cost multiples to rebuild, and no custom build will out-run that math for a standard store. Building custom makes sense in five specific situations: your all-in Shopify bill is approaching Plus territory (from $2,300/month on a 3-year term, $2,500 on 1-year), you need checkout control that Shopify gates to Plus, you want a headless storefront with full frontend ownership, your operations are ERP-heavy (NetSuite, SAP, custom pricing and inventory logic), or you are running a marketplace model with multiple sellers. If none of those describe you, use Shopify and spend the difference on marketing. If two or more do, the custom math starts winning, and a scoped storefront build starts at $4,500 with typical projects landing between $8,000 and $25,000. URL: https://www.bearplex.com/compare/shopify-vs-custom-ecommerce --- ### BambooHR vs Building Your Own: Which to Choose in 2026 For most companies, the honest answer is: buy BambooHR. At $10 to $25 per employee per month (verified on bamboohr.com, July 2026), a 50-person US company pays about $18,000 over three years for a mature HRIS. You cannot build and maintain anything comparable for that. Build custom only when one of three narrow triggers applies: your HR or recruitment workflow IS your product or a revenue line, your compliance requirements live in a niche no HRIS models (NDIS care providers, for example), or the HR layer must be embedded inside a larger custom platform. We say this as a team that built its own HR and recruitment SaaS, PeoplePlus. It was worth building because it is our product, not our back office. If BambooHR were only going to run our internal HR, we would have bought it. URL: https://www.bearplex.com/compare/bamboohr-vs-custom-hr-software --- ### Airtable vs Building Your Own: Which to Choose in 2026 For most teams, Airtable is the right call and you should not build anything: at $20 per editor per month on the Team plan (verified July 2026), no custom build competes for internal trackers, light workflows, and databases under about 50,000 records. Build custom when you hit specific walls: record volume approaching the 50,000 (Team) or 125,000 (Business) per-base caps, automations pausing at monthly run limits, a customer-facing product that needs real authentication and UI, per-seat spend crossing build economics at roughly 25+ Business seats, or compliance requirements that demand owning the data. The typical arc is healthy, not a failure: prototype in Airtable, prove the workflow, then migrate the proven system to a custom build once the limits start costing real money. When that moment comes, a custom internal tool or ops system starts at $15,000 and typically runs $25,000-$70,000, and your Airtable base becomes the best requirements document you will ever hand a development team. URL: https://www.bearplex.com/compare/airtable-vs-custom-database-app --- ### SharePoint vs Building Your Own: Which to Choose in 2026 If your company already runs on Microsoft 365, SharePoint is usually the right intranet call: the licensing is already paid, document collaboration is genuinely best in class, and no custom build beats an incremental cost of zero for internal news, policies, and file sharing. Build custom in four specific situations: the portal is client-facing (external guest access is SharePoint's most fragile surface), the tool is workflow-first rather than content-first (operations, approvals, custom data, dashboards), your company runs on Google Workspace or another non-Microsoft stack, or the portal is part of the product you sell and needs product-grade UX. For everyone else, the honest answer is stay on SharePoint and fix governance before paying anyone to rebuild what you already own. URL: https://www.bearplex.com/compare/sharepoint-vs-custom-intranet --- ### Slack vs Building Your Own: Which to Choose in 2026 Buy Slack. For almost every team asking this question, that is the honest answer: Slack Pro costs $7.25 per user per month on annual billing (verified at slack.com/pricing, July 2026), so a 50-person company pays about $4,350 a year for a product category Slack has spent over a decade perfecting. No custom build competes with that on chat, and a from-scratch clone will always be a worse Slack. The real question hiding inside 'should we build our own Slack' is usually different: tool sprawl, client communication leaking across email and DMs, and operational data living in screenshots and scrollback. You fix that by keeping Slack for conversation and building the thing Slack cannot be: an ops hub, a client portal, or a notification layer on top of your actual systems. That build starts at $15,000, typically runs $25,000-$70,000, and unlike a chat clone it pays for itself. URL: https://www.bearplex.com/compare/slack-vs-custom-internal-comms --- ## Hire AI Talent (25 roles) BearPlex staffs vetted senior engineers by role. Each page details the skills matrix, vetting process, and engagement model. ### LLM Engineers An LLM engineer at BearPlex owns the full lifecycle of a production language model system. That means designing the prompt and retrieval architecture, building the evaluation harness BEFORE writing the agent loop, integrating with your existing data sources and IAM, hardening for production with proper observability (LangSmith, Arize, OpenTelemetry), and operating the system after launch. Our LLM engineers ship to production within the first sprint of an engagement: they don't write demos that get thrown away. They've worked with the full stack: GPT-5, Claude Sonnet 4.5, Llama 3.3, fine-tuning with LoRA and DPO, RAG with Pinecone/Qdrant/Weaviate, agent frameworks like LangGraph and the Claude Agent SDK, and the operational tooling that distinguishes a prototype from a system you can run at scale. They also know what NOT to build: they'll push back on architecture decisions that feel sophisticated but won't survive production. URL: https://www.bearplex.com/hire/llm-engineers --- ### ML Engineers An ML engineer at BearPlex owns the complete production lifecycle of machine learning systems: distinct from LLM engineers (who specialize in language models) and data scientists (who lean toward exploratory analysis). That includes feature engineering pipelines (Airflow, Dagster, dbt), model training and experiment tracking (MLflow, Weights & Biases), online serving infrastructure (BentoML, Seldon, SageMaker), feature stores when warranted (Feast, Tecton), monitoring for data drift and model decay (Evidently, Arize, WhyLabs), and CI/CD for models (testing, validation, blue-green deployment). They work across the full ML stack: classical ML for tabular problems where it still beats LLMs (XGBoost, LightGBM, Random Forests), deep learning for vision and time-series (PyTorch, TensorFlow), and increasingly hybrid systems where classical ML and LLMs work together. Our ML engineers have shipped recommendation systems, fraud detection pipelines, demand forecasting models, computer vision systems, and time-series forecasting at production scale across logistics, financial services, healthcare, and retail. URL: https://www.bearplex.com/hire/ml-engineers --- ### AI Engineers An AI engineer at BearPlex is the hybrid generalist that production AI roadmaps actually need: equally comfortable building an LLM agent system on Monday, training an XGBoost fraud model on Tuesday, and shipping the platform infrastructure to operate both on Wednesday. The role exists because most enterprise AI projects don't fit neatly into either 'pure LLM' or 'pure ML' boxes. A real customer support deployment needs RAG (LLM expertise), intent classification (classical ML), feature pipelines for analytics (data engineering), and serving infrastructure (platform). One engineer who can own all four ships dramatically faster than a team of specialists who have to coordinate across handoffs. AI engineers at BearPlex have backgrounds spanning LLM systems, classical ML, and platform engineering: they're the engineers we deploy when the project shape requires breadth, when team size is constrained, or when the client team needs a single technical owner across multiple AI initiatives. URL: https://www.bearplex.com/hire/ai-engineers --- ### RAG Engineers A RAG engineer at BearPlex specializes in production retrieval systems: the kind that handle real enterprise document corpora (10M+ documents), enforce role-based access control at retrieval time, integrate with existing IAM systems, track citations back to source paragraphs, and survive the daily reality of regulated industries. Generic RAG tutorials stop working at scale; production-grade RAG is its own discipline. Our RAG engineers know that chunking strategy matters more than people realize, that hybrid search (BM25 + vectors + reranking) consistently beats vector-only, that filter-first retrieval is how you enforce permissions, that RAGAS evaluation is non-negotiable. They've shipped systems handling millions of legal documents (with privilege preservation), enterprise knowledge bases (with org-chart-aware permissions), customer support corpora (with citation tracking), and regulated healthcare retrieval (with HIPAA boundary enforcement). They specialize in the production hardening that turns 'RAG demo' into 'RAG you trust with business-critical workflows.' URL: https://www.bearplex.com/hire/rag-engineers --- ### AI Agent Developers An AI agent developer at BearPlex specializes in the production architecture of autonomous AI workflows: the systems where an LLM doesn't just answer questions but takes actions, runs multi-step processes, and operates with appropriate autonomy under human oversight. This is its own engineering discipline, distinct from chatbot development. Our agent developers know that LangGraph's explicit state management beats LangChain's AgentExecutor for production work, that tool design discipline (clear descriptions, argument validation, structured error handling) is the #1 driver of agent reliability, that evaluation harnesses must be built BEFORE the agent loop, that human checkpoints on consequential actions are non-negotiable, and that step/cost limits prevent runaway agents from racking up thousands of dollars in API calls. They've shipped autonomous workflows for customer support (multi-step issue resolution), document processing (classify-route-extract-summarize pipelines), DevOps (incident triage and runbook execution), and complex domain-specific agents (legal contract analysis, healthcare prior authorization, financial fraud explanation). They build for production reliability, not demo magic. URL: https://www.bearplex.com/hire/ai-agent-developers --- ### MLOps Engineers An MLOps engineer at BearPlex builds the production infrastructure that ML and LLM systems need to operate reliably: distinct from data science (modeling) and ML engineering (model development). Their work is the platform layer: data pipelines (Airflow, Dagster, Prefect), model registries and versioning (MLflow, Weights & Biases), CI/CD for ML (testing, validation, blue-green deployment), serving infrastructure (BentoML, Seldon, SageMaker, vLLM for LLMs), monitoring for data drift and model decay (Evidently, Arize, WhyLabs), feature stores when warranted (Feast, Tecton), and the operational discipline that distinguishes a working model from a system you can run for years. They handle both classical ML systems (where model decay over months is the concern) and LLM systems (where prompt changes, model version updates, and cost optimization are the concerns). Our MLOps engineers have shipped MLOps platforms for fraud detection running on millions of daily transactions, recommendation systems with hourly retraining cycles, LLM applications serving thousands of concurrent users, and the meta-platforms (CI/CD for ML, feature stores) that ML teams build their work on top of. URL: https://www.bearplex.com/hire/mlops-engineers --- ### Fine-Tuning Engineers A fine-tuning engineer at BearPlex owns the full model adaptation lifecycle: dataset construction (the hardest and most under-appreciated part), evaluation harness design, training run management, hyperparameter sweeps, model evaluation against held-out tasks, and production deployment of the resulting model. They work across the full method stack: LoRA and QLoRA for parameter-efficient adaptation, full fine-tuning when scale justifies it, DPO and ORPO for preference alignment, supervised fine-tuning (SFT) for instruction-following, and continued pre-training for domain shifts. They've shipped fine-tuned models on Hugging Face TGI, vLLM, Together AI, AWS Bedrock custom models, and Azure ML. They know when fine-tuning is the right answer (rigid format compliance, per-call cost optimization at scale, specialized domain language) and when it's the wrong one (fast-changing knowledge that should be in RAG instead). Most importantly, they build evaluation harnesses BEFORE training, because a fine-tuned model with no evaluation is just an expensive version of the base model. URL: https://www.bearplex.com/hire/fine-tuning-engineers --- ### Prompt Engineers A prompt engineer at BearPlex is part engineer, part product designer, part QA lead. They own the system prompts, agent prompts, function-call schemas, evaluation rubrics, and prompt-versioning infrastructure for production AI systems. The role goes far beyond writing prompts: they instrument prompts with structured logging so you can analyze production behavior, build evaluation harnesses that catch regressions before they ship, design prompt A/B test infrastructure, and translate ambiguous business requirements into testable model specifications. They work across providers (Claude, GPT, Gemini, Llama) and know each model's quirks (Claude's preference for XML structure, GPT's JSON mode reliability, Gemini's long-context behavior). They've shipped systems that depend on prompts at scale: customer support copilots handling thousands of tickets/day, autonomous research agents, internal knowledge assistants. They also know when to escalate from prompting to fine-tuning, RAG, or different model selection: a great prompt engineer doesn't try to solve every problem with cleverer wording. URL: https://www.bearplex.com/hire/prompt-engineers --- ### Data Engineers A data engineer at BearPlex owns the full data pipeline lifecycle: source ingestion (Fivetran, Airbyte, custom CDC, Kafka), warehouse modeling (dbt, SQL, dimensional design), transformation pipelines (batch and streaming), data quality engineering (tests, observability, lineage), and operational ownership of the resulting systems. They work across the modern data stack: Snowflake, BigQuery, Databricks, ClickHouse for warehousing; Kafka, Kinesis, RabbitMQ for streaming; Airflow, Dagster, Prefect for orchestration; dbt for transformation; Hightouch and Census for reverse ETL, and know which tools fit which problems. They've shipped pipelines that handle 50B+ events per month, built customer 360 models that reconcile identity across 8+ source systems, and stood up AI-ready feature stores for production ML. Importantly, they push back on the 'Modern Data Stack' diagram when it doesn't fit the actual problem, sometimes the right answer is Postgres + cron, not Snowflake + Airflow + dbt + Hightouch. URL: https://www.bearplex.com/hire/data-engineers --- ### NLP Engineers An NLP engineer at BearPlex covers the full natural language processing stack: classical methods (regex, CRF, spaCy, transformers for classification), modern LLM-based methods (prompting, RAG, fine-tuning, agents), and the engineering work that turns NLP research into production systems. They've shipped: named entity recognition pipelines extracting structured data from unstructured text, document classification systems serving millions of inferences per day, RAG systems over millions of documents, multilingual NLP for global products, and conversational systems that combine intent classification with LLM generation. They know when to reach for a 7B-parameter LLM (most modern needs) vs when classical NLP is still the right answer (high-volume classification with strict latency budgets, multilingual entity extraction at scale). They build evaluation harnesses appropriate to NLP, not just LLM eval, but precision/recall on classification, span F1 on entity extraction, BLEU/ROUGE/BERTScore where appropriate. URL: https://www.bearplex.com/hire/nlp-engineers --- ### Computer Vision Engineers A computer vision engineer at BearPlex covers the full CV stack: classical computer vision pipelines (OpenCV preprocessing, classical detection algorithms), deep learning CV models (CNNs, Vision Transformers, segmentation, object detection), modern vision-language models (CLIP, GPT-4V, Claude vision, Gemini), and the production engineering required to deploy CV at scale. They've shipped: real-time object detection on edge devices, document understanding pipelines for insurance and legal, defect detection systems for manufacturing QA, video analytics for retail and security, multimodal RAG systems combining images and text. They know when to use a fine-tuned YOLO model (low latency, high volume, well-defined object types) vs when to use GPT-4V or Claude vision (zero-shot understanding, novel object categories, complex visual reasoning). They handle the operational challenges that distinguish production CV from research demos: model serving on GPUs and edge devices, dataset annotation strategies, evaluation across diverse imaging conditions, and the domain shift problems that kill most CV deployments. URL: https://www.bearplex.com/hire/computer-vision-engineers --- ### Generative AI Engineers A generative AI engineer at BearPlex specializes in production systems whose primary job is generating content. The role spans text generation (chatbots, content tools, code generation), image generation (DALL-E, Stable Diffusion, FLUX, Midjourney API), video generation (Runway, Sora, Veo, Kling), audio generation (ElevenLabs, OpenAI TTS, music generation), and structured generation (JSON, XML, code, database schemas). They work with frontier APIs (GPT-5, Claude 4 Opus, Gemini 2.5, DALL-E 3, Midjourney, Runway) and self-hosted open-source generative models (Stable Diffusion XL, FLUX.1, Llama 3.3 for text, Mixtral, AudioGen). They've shipped: marketing content generation tools that produce on-brand copy at scale, code generation systems that build whole features from natural language specs, image generation pipelines for ecommerce product photography, video generation workflows for short-form content marketing, and synthetic data generation pipelines for ML training. They know the production realities of generative work: prompt engineering at the level of measured eval rather than vibes, brand and quality controls, content safety filtering, IP and copyright considerations, and the cost economics of generation (which is dramatically more expensive per call than classification or retrieval). URL: https://www.bearplex.com/hire/generative-ai-engineers --- ### ChatGPT Developers A ChatGPT developer at BearPlex builds production applications on the OpenAI platform end-to-end. They know the OpenAI API surface deeply: chat completions, Assistants API for stateful applications, function calling and structured outputs for reliable tool integration, the o-series reasoning models for hard reasoning tasks, fine-tuning for narrow tasks at scale, embeddings for retrieval, DALL-E for image generation, Whisper for speech-to-text, voice models for STT/TTS. They've shipped: customer support copilots, internal knowledge assistants, autonomous workflow agents, content generation tools, code generation features, multimodal applications combining text and image, and high-volume classification pipelines. They know the platform's idioms: when GPT-4o-mini is the right cost-quality answer vs when GPT-5 is required; when to use Assistants API vs raw chat completions; when function calling is enough vs when structured outputs are required; how to use prompt caching for cost optimization. Equally important: they know the platform's limitations and when to reach for Anthropic / open-source / custom solutions instead. URL: https://www.bearplex.com/hire/chatgpt-developers --- ### AI Chatbot Developers An AI chatbot developer at BearPlex builds production conversational systems end-to-end. The role spans: conversational design (turn structure, persona, fallback patterns), retrieval design (RAG over the customer's knowledge), tool integration (chatbot taking actions in your stack), evaluation harness construction (catching regressions before they ship), front-end integration (web widgets, mobile, in-product, voice), and operational ownership (monitoring, incident response, continuous improvement). They know the chatbot stack: foundation models (Claude, GPT, Gemini), agent frameworks (LangGraph, Claude Agent SDK), retrieval infrastructure (Pinecone, Qdrant, pgvector), front-ends (Vercel AI SDK, custom React, Intercom Custom Channels, Slack, Teams), and the operational layer (LangSmith, Helicone, custom analytics). They've shipped chatbots that actually work in production: deflecting 60-75% of tier-1 support tickets, generating qualified leads, handling complex multi-step workflows, and operating at scale across thousands of customer tenants. They also know what makes chatbots fail (the patterns that produce demo-quality chatbots but break in production) and design around them. URL: https://www.bearplex.com/hire/ai-chatbot-developers --- ### AI Consultants An AI consultant at BearPlex bridges executive strategy and engineering execution. The role spans: AI strategy and roadmap development (where to invest, what to build vs buy, how to sequence), AI vendor evaluation (model providers, vector databases, agent platforms, MLOps tooling), technical due diligence on AI products (for investors, acquirers, partners), AI governance framework design (NIST AI RMF, EU AI Act, ISO 42001), build-vs-buy decision support, prompt engineering and architecture review, and the executive-facing communication that translates technical decisions into business outcomes. Our consultants are senior engineers who can also produce the kind of analysis that survives board-level scrutiny: they've shipped production AI themselves and bring those scars to advisory work. Standard engagements include Discovery Sprints (1-2 weeks, intensive scoping for new initiatives), Technical Due Diligence (focused investor / acquirer audits), AI Strategy Engagements (4-8 weeks producing roadmap + governance + execution plan), and Fractional CTO arrangements for clients without senior in-house AI leadership. URL: https://www.bearplex.com/hire/ai-consultants --- ### AI Architects An AI architect at BearPlex owns the architecture of production AI systems: making the design decisions that shape how systems work, scale, and evolve over months and years. The role spans: agent architecture (single-agent vs multi-agent, state management, tool design, HITL patterns), RAG architecture (retrieval pipelines, chunking strategy, multi-tenancy, evaluation infrastructure), data architecture for AI (feature stores, embedding pipelines, training data infrastructure), governance architecture (model registry, audit logging, compliance integration), and the platform architecture that supports all the above. Our architects are senior engineers (typically 10+ years) who've shipped production AI systems and now design them at scale. They produce: architecture documents that survive engineering team scrutiny, design reviews that catch problems before they're built, and the patterns that get implemented across many AI projects within an organization. They also know when NOT to architect: when a simple solution beats a sophisticated one, when the right answer is to use someone else's product instead of building, when premature abstraction will slow the team down. URL: https://www.bearplex.com/hire/ai-architects --- ### AI Solutions Engineers An AI solutions engineer at BearPlex is the technical bridge between your sales team and prospective customers. The role spans: discovery (technical scoping calls with customer engineering teams), proof-of-concept development (build a working integration with the customer's data and stack in 1-3 weeks), technical demos (live coding for executive audiences), RFP responses (the technical sections), customer-specific implementation work (the first 60-90 days post-contract), and the feedback loop back to product on what real customers need. They've shipped production AI for enterprise customers across Fortune 500 financial services, healthcare, and B2B SaaS: meaning they know what enterprise procurement actually evaluates and how to address it. They're equally comfortable in customer architecture review, on a sales call, and in their IDE shipping code. This is a rare skill set; most engineers can't talk to customers comfortably and most sales engineers can't actually ship code. URL: https://www.bearplex.com/hire/ai-solutions-engineers --- ### AI Platform Engineers An AI platform engineer at BearPlex builds the infrastructure that powers all AI initiatives across an organization. The role spans: shared model serving infrastructure (frontier models, fine-tuned models, self-hosted open-source), centralized retrieval infrastructure (vector indexes, reranking, embedding pipelines), model governance and registry (version tracking, validation, MRM integration), evaluation infrastructure (golden datasets, LLM-as-judge pipelines, regression detection), developer experience (internal SDKs, documentation, templates), cost monitoring and optimization, and the operational layer that makes the platform reliable. They work across the modern AI infrastructure stack: vLLM and Triton for serving, Pinecone and Qdrant for retrieval, MLflow and custom registries for governance, Promptfoo and Braintrust for evaluation, LangSmith and Helicone for observability. They've built platforms supporting 5-50+ AI initiatives across organizations and know what scales vs what doesn't. URL: https://www.bearplex.com/hire/ai-platform-engineers --- ### Claude Developers A Claude developer at BearPlex builds production applications on the Anthropic platform end-to-end. They know the Claude API surface deeply: chat completions, tool use with parallel call support, prompt caching with 90% discount on cached prefixes, citations API for RAG-grounded outputs, extended thinking for hard reasoning, computer use for desktop application automation, and the Claude Agent SDK for production agent systems. They've shipped: production agent systems, code-generation features (Claude's strongest task category), customer support copilots, document analysis pipelines (long-context Claude shines here), and multi-step reasoning agents. They know when Claude is the right answer (code, long context, agents, safety-conscious enterprise) and when it isn't (image generation, cost-sensitive bulk classification, vendor-portable code). Equally important: they know Claude's idioms (XML structure for prompts, tool use ergonomics, prompt caching configuration) that produce significantly better results than generic LLM patterns. And they work in Claude Code daily, as does every engineer at BearPlex: agentic coding is our default delivery workflow, and we bring the same setup (custom skills, hooks, MCP connectors, CI integration) to client teams adopting Claude Code. URL: https://www.bearplex.com/hire/claude-developers --- ### AI Developers An AI developer at BearPlex ships production AI features across the full stack: from the React component that renders the AI feature, through the backend API that handles the LLM integration, through the prompt engineering and evaluation that determines quality, to the deployment and operational infrastructure that keeps it running. They work in TypeScript / Python with modern frontend frameworks (Next.js, React, Vue), backend APIs (FastAPI, Express, NestJS), LLM integrations (Vercel AI SDK, Anthropic SDK, OpenAI SDK), and the production AI tooling stack (LangSmith, Helicone, Promptfoo). They're generalists who can take an AI feature from spec to production deployment without needing handoffs between frontend, backend, and ML specialists. This makes them ideal for fast-moving teams that need to ship AI features quickly without coordinating across multiple engineering disciplines. URL: https://www.bearplex.com/hire/ai-developers --- ### Machine Learning Consultants A machine learning consultant at BearPlex brings senior ML expertise to organizational and architectural decisions. The role spans: ML strategy and roadmap (where to invest, build vs buy, how to sequence), ML architecture review (audit existing systems, identify risks and opportunities), ML vendor evaluation (frameworks, platforms, tooling), model risk management for regulated industries, ML governance framework design, technical due diligence on ML products (for investors and acquirers), and the executive communication that translates ML decisions into business outcomes. Our consultants are senior ML practitioners (typically 10+ years) who've shipped production ML themselves and bring those scars to advisory work. Standard engagements include ML Discovery Sprints (1-2 weeks of intensive scoping), ML Architecture Reviews (3-4 weeks of audit + recommendations), Strategy Engagements (4-8 weeks producing roadmap + governance + execution plan), and Fractional Head of ML arrangements. URL: https://www.bearplex.com/hire/machine-learning-consultants --- ### AI Product Managers An AI product manager at BearPlex owns the product side of AI features: discovery (which AI features should we build?), prioritization (which first?), spec writing (what exactly does this AI feature do?), evaluation design (how do we know it's good?), launch coordination, and iteration based on production behavior. The role is harder than traditional PM because AI features have probabilistic behavior, evaluation is non-trivial, and the design space is constantly changing as new model capabilities emerge. Our AI PMs combine product judgment with deep AI literacy: they understand what current models can and can't do, design eval rubrics that catch real problems, write specs that engineers can implement, and own the iteration cycle that makes AI features improve over time. They've shipped production AI features for B2B SaaS, fintech, healthcare, and consumer products. URL: https://www.bearplex.com/hire/ai-product-managers --- ### LangChain Developers A LangChain developer at BearPlex builds production applications on the LangChain ecosystem end-to-end. They know LangChain (the framework) for integrations and prototyping, LangGraph (the production agent orchestration library) for stateful agents, LangSmith (the observability product) for production monitoring, and the broader LangChain integration ecosystem. They've shipped: production agent systems on LangGraph, RAG pipelines using LangChain primitives, multi-agent workflows, and the migration work that takes legacy LangChain AgentExecutor systems to modern LangGraph. They know the framework's strengths and weaknesses honestly: when LangChain accelerates work and when bypassing the abstractions to direct API calls is the right move. They're equally comfortable in Python (the primary LangChain ecosystem) and TypeScript (the LangChainJS port), though Python is preferred for production work where feature parity matters. URL: https://www.bearplex.com/hire/langchain-developers --- ### AI Research Engineers An AI research engineer at BearPlex sits between academic research and production deployment. The role spans: implementing recent papers (turning arXiv preprints into working production systems), designing novel architectures for problems that don't have off-the-shelf solutions, evaluating frontier model capabilities for emerging use cases, leading research-heavy phases of complex AI engagements, and translating research insights into production systems. They've worked across the full spectrum: fine-tuning research, RLHF / DPO / Constitutional AI, agent system design, novel retrieval architectures, multi-modal AI, reasoning systems. They read papers continuously, implement the most-promising recent work, and have the engineering skill to take research from notebook to production. This is a rare profile; most engineers don't read papers continuously, and most researchers don't ship production systems. URL: https://www.bearplex.com/hire/ai-research-engineers --- ### Deep Learning Engineers A deep learning engineer at BearPlex specializes in production neural network engineering: architecture design (CNNs, Transformers, hybrid architectures), distributed training (multi-GPU, multi-node), optimization (quantization, distillation, pruning), and deployment (high-throughput inference, edge deployment, model serving infrastructure). They work across the full stack: PyTorch (primary), JAX, Hugging Face Transformers, ONNX Runtime, TensorRT, custom CUDA when needed. They've shipped: production deep learning models for computer vision (object detection, segmentation, OCR), NLP (classification, NER, custom Transformer architectures), audio (speech recognition, audio classification), multimodal (CLIP variants, vision-language models), and increasingly fine-tuned LLMs for specific domains. They're the right hire when frontier APIs can't solve the problem: when you need custom architecture, specialized models for your domain, or production infrastructure for self-hosted deep learning at scale. URL: https://www.bearplex.com/hire/deep-learning-engineers --- ## Industries (8) ### B2B SaaS & Software Audience: VPs of Engineering, CTOs, Heads of Product at growth-stage and enterprise SaaS companies Key industry statistics: - $232B Global SaaS market 2025 (source: Gartner 2025) - 78% of SaaS companies actively building AI features (source: Bessemer Cloud Benchmark 2025) - 47% average reduction in support ticket volume after deploying AI agents (source: Gainsight 2025 PX Benchmark) - $0.40 median cost-per-resolution after agentic deployment vs $4.20 human-only (source: Intercom Customer Service Trends 2025) URL: https://www.bearplex.com/industries/saas --- ### Financial Services (FinTech, Banking, Insurance) Audience: Chief Technology Officers, Heads of AI, Compliance Officers at banks, insurers, FinTechs, and asset managers Key industry statistics: - $25B FinTech AI market 2025 (source: Boston Consulting Group 2025) - 92% of large banks running AI pilots in 2025 (source: McKinsey Global Banking Annual Review 2025) - $1.2T global financial services AI spend forecast for 2030 (source: Statista 2025) - 73% of insurers report AI as critical to fraud detection roadmap (source: Coalition Against Insurance Fraud 2025) URL: https://www.bearplex.com/industries/financial-services --- ### Healthcare (Providers, Pharma, Medical Devices) Audience: Chief Medical Information Officers, VPs of Clinical Informatics, Heads of AI at health systems, payors, pharma, and medical device companies Key industry statistics: - $187B Healthcare AI market by 2030 (source: Grand View Research 2025) - 67% of US health systems piloting LLM agents in 2025 (source: American Hospital Association 2025) - 65.3% AI Overview coverage on healthcare queries (highest of any vertical we tracked) (source: Backlinko Healthcare AI Search Study 2025) - 2.7 hours average daily clinician burden on EHR documentation eliminated by AI ambient scribes (source: Mayo Clinic AI Initiative 2025) URL: https://www.bearplex.com/industries/healthcare --- ### Legal (LegalTech, Law Firms, In-House Counsel) Audience: Chief Innovation Officers, Heads of Legal Operations, Practice Group Leaders at AmLaw firms, corporate legal departments, and LegalTech vendors Key industry statistics: - $1.45B LegalTech AI market 2025 (source: Thomson Reuters Institute 2025) - 77.7% AI Overview coverage on legal queries (highest of any vertical we tracked) (source: Backlinko Legal AI Search Study 2025) - 85% of AmLaw 100 firms have at least one production GenAI deployment (source: Wolters Kluwer Future Ready Lawyer 2025) - 11× speedup on first-pass contract review with AI clause extraction (source: Stanford CodeX Legal Informatics 2025) URL: https://www.bearplex.com/industries/legal --- ### E-commerce & Retail Audience: Heads of Engineering, VPs of Product, Chief Digital Officers at DTC brands, marketplaces, and large retailers Key industry statistics: - $24B E-commerce AI market 2025 (source: Statista 2025) - 67% of online shoppers expect AI-personalized experiences (source: Salesforce Connected Customer 2025) - 21% average lift in conversion rate from AI-powered product discovery (source: Algolia AI Search Benchmark 2025) - $338B global retail revenue from AI personalization by 2027 (source: McKinsey Retail AI Report 2025) URL: https://www.bearplex.com/industries/ecommerce --- ### Manufacturing & Industrial Audience: Chief Information Officers, VPs of Operations, Plant Managers at industrial manufacturers and OEMs Key industry statistics: - $28B Manufacturing AI market 2025 (source: Deloitte Manufacturing Industry Outlook 2025) - 40% of manufacturers report AI-driven productivity gains above 15% (source: World Economic Forum Industrial AI 2025) - $1.4T potential global manufacturing value from generative AI by 2030 (source: McKinsey Generative AI Report 2025) - 73% of manufacturing AI projects stall before production due to OT/IT integration (source: Gartner Industrial AI Survey 2025) URL: https://www.bearplex.com/industries/manufacturing --- ### Logistics, Supply Chain & 3PL Audience: Chief Supply Chain Officers, VPs of Operations, Heads of Logistics at carriers, 3PLs, and shippers Key industry statistics: - $23B Logistics AI market 2025 (source: Allied Market Research 2025) - $1.6T global logistics market 2025 (source: Statista 2025) - 47 AI agents BearPlex deployed in 90 days for one Fortune 100 logistics client (source: BearPlex case study, December 2025) - $14M annualized cost savings from that single deployment (source: BearPlex case study, December 2025) URL: https://www.bearplex.com/industries/logistics --- ### Government & Public Sector Audience: Chief AI Officers, Chief Data Officers, Agency CTOs at federal civilian agencies, defense, intelligence, state, and municipal government Key industry statistics: - $3.3B US federal AI contract spend FY2024 (source: Bloomberg Government 2025) - 1,757 AI use cases inventoried across 41 federal agencies (source: AI.gov use case inventory 2025) - M-24-10 OMB memo on agency AI governance: sets baseline requirements for all federal AI (source: Office of Management and Budget 2024) URL: https://www.bearplex.com/industries/government --- ## Products ### AdForge: AI-Powered Ad Intelligence & Creative Engine BearPlex's productized ad-intelligence and creative-generation engine. URL: https://www.bearplex.com/products/adforge --- ## Contact - Email: hello@bearplex.com - Website: https://www.bearplex.com - Contact form: https://www.bearplex.com/contact - Pricing: https://www.bearplex.com/pricing - LinkedIn: https://www.linkedin.com/company/bearplex - X: https://x.com/bearplexdigital - Instagram: https://www.instagram.com/bearplexdigital/ - Facebook: https://www.facebook.com/BearPlex/ - GitHub: https://github.com/bearplex - RSS feed: https://www.bearplex.com/feed.xml --- *This file is auto-generated from https://www.bearplex.com/llms-full.txt and updated at every build.*