00:00:14 --> 00:00:25
Gemma 4 is a new open-weight AI from Google that can run entirely on a contractor’s own hardware, keeping controlled data inside the security boundary. How does that work, analyst?
00:00:25 --> 00:00:43
Gemma 4 was released on March 31, 2026 in several sizes: E2B, E4B, 26B A4B, 31B Dense, and a 12B Unified added on June 3, 2026. It ships under an Apache 2.0 license.
00:00:43 --> 00:00:48
So the license is standard open source, but what makes it practical for defense contractors?
00:00:49 --> 00:01:00
Apache 2.0 is a commercially permissive license that eliminates a vendor-specific terms-of-use document. Legal teams already review it for other software, so the process is straightforward.
00:01:01 --> 00:01:04
And the hardware? What do we need to run the biggest model?
00:01:04 --> 00:01:17
The unquantized 31B Dense model fits a single 80GB NVIDIA H100 GPU. Quantized versions run on consumer GPUs, even on phones or Raspberry Pi.
00:01:18 --> 00:01:21
That’s a wide range. How do the smaller models perform?
00:01:22 --> 00:01:34
E2B and E4B use per-layer embeddings and run offline with near-zero latency on a phone or a Jetson Orin Nano. They also support 30-second audio clips natively.
00:01:34 --> 00:01:38
What about context windows? How much text can each model handle?
00:01:39 --> 00:01:56
E2B and E4B support 128K tokens. The 12B Unified, 26B A4B, and 31B Dense models support 256K tokens, giving a full technical document in one request.
00:01:57 --> 00:02:01
That’s impressive. Why is local inference so critical for CUI?
00:02:02 --> 00:02:16
When a contractor pastes CUI into a hosted AI, the data leaves the system security plan and requires additional controls. Local inference keeps prompts, documents, and outputs inside the existing boundary.
00:02:16 --> 00:02:24
So it satisfies DFARS 252-7012 and NIST SP 800-171?
00:02:24 --> 00:02:35
Exactly. The inference host sits on a segmented VLAN or air-gap, so the same access, audit, and encryption controls that protect other CUI apply to the AI stack.
00:02:36 --> 00:02:38
How does that affect the System Security Plan?
00:02:38 --> 00:02:53
Adding a model expands the SSP boundary to include the runtime, checkpoint storage, and inference cache. Those components must be mapped to the 110 controls of NIST SP 800-171.
00:02:53 --> 00:02:54
And if a requirement isn’t met?
00:02:55 --> 00:03:07
It goes into a Plan of Actions and Milestones, which is allowed under 32 CFR 170.21. The model file itself isn’t compliant; the deployment is what gets assessed.
00:03:08 --> 00:03:10
Can you walk us through Petronella’s seven-stage method?
00:03:11 --> 00:03:20
Sure. First, define the data boundary and applicable frameworks. Second, run a paid scoping engagement to build a prototype on the client’s data.
00:03:20 --> 00:03:22
What does the prototype involve?
00:03:22 --> 00:03:29
It measures context lengths, concurrency, and accuracy requirements, then records those metrics for hardware sizing.
00:03:30 --> 00:03:32
So hardware sizing follows the prototype results?
00:03:33 --> 00:03:47
Yes. The GPU cluster is sized based on the prototype’s performance. A 31B Dense model under concurrency requires an H100, while an E4B field utility might fit on a single RTX 5090.
00:03:47 --> 00:03:49
After sizing, what’s next?
00:03:49 --> 00:03:57
The cluster is isolated-either a VLAN or a full air-gap-so only approved traffic reaches the inference host.
00:03:57 --> 00:03:59
How are access and audit controls layered?
00:04:00 --> 00:04:17
Role-based access controls map to NIST 3.1, encryption at rest and in transit uses FIPS-validated cryptography, and audit logging satisfies NIST 3.3. All logs are preserved for the 90-day retention required by DFARS.
00:04:17 --> 00:04:18
And the final stage?
00:04:19 --> 00:04:28
Validate the deployment against the mapped controls, document the results in the SSP, and then hand off to a C3PAO for certification.
00:04:28 --> 00:04:32
What about Section 1532 of the FY2026 NDAA?
00:04:32 --> 00:04:45
Section 1532 defines covered artificial intelligence as AI developed by DeepSeek or by High Flyer and its affiliates. Gemma, developed by Google DeepMind, is not mentioned.
00:04:45 --> 00:04:48
So Gemma isn’t restricted by that provision?
00:04:48 --> 00:04:55
Correct. The law prohibits only the named developers; it does not reference the model family or its licensing.
00:04:55 --> 00:04:58
Are there other Chinese-origin models that raise concerns?
00:04:59 --> 00:05:11
GLM-5.3, for example, is owned by Zhipu AI, which is on the BIS Entity List. Exporting its weights to a hosted service would require a license, but self-hosting is permissible.
00:05:11 --> 00:05:13
What about future legislation?
00:05:13 --> 00:05:25
The FY2027 Senate bill S. 4784 would add Zhipu AI to the list of covered AI. That could force a mid-contract change if the model is used.
00:05:25 --> 00:05:30
So risk management is key. How does Gemma compare to GPT-OSS?
00:05:30 --> 00:05:50
Both are Apache 2.0, but GPT-OSS 20B runs on a 16GB GPU and GPT-OSS 120B on an 80GB GPU with MXFP4 quantization. They are text-only and have smaller context windows of 128K tokens.
00:05:50 --> 00:05:53
Gemma offers more modalities and larger context windows, right?
00:05:54 --> 00:06:10
Exactly. Every Gemma 4 size supports images, and E2B, E4B, and 12B Unified also support audio. The larger models provide 256K tokens, which is useful for long technical documents.
00:06:10 --> 00:06:15
If a contractor needs a small, low-cost deployment, what’s the best choice?
00:06:15 --> 00:06:27
The E2B model in the QAT mobile format fits in just 1GB of memory and runs offline on a phone or Raspberry Pi. That covers field laptops or edge devices.
00:06:27 --> 00:06:29
And for a mid-size contractor with a data center?
00:06:30 --> 00:06:41
A single 80GB H100 can run the 31B Dense model, which is the foundation for fine-tuning. That fits within a typical CUI enclave.
00:06:41 --> 00:06:43
What about the legal review process for the model?
00:06:44 --> 00:06:53
Because the license is Apache 2.0, the review is a standard open-source assessment. No separate policy document or click-through gate is required.
00:06:53 --> 00:06:56
How does Petronella’s own deployment illustrate this approach?
00:06:56 --> 00:07:10
Petronella runs a private AI cluster in a hybrid SOC that never sends CUI to a public cloud. The AI never closes tickets autonomously, and all actions are logged for CMMC and HIPAA audit.
00:07:10 --> 00:07:13
What does a typical engagement look like in terms of duration?
00:07:13 --> 00:07:24
A CMMC Level 2 readiness engagement spans 12 to 14 weeks: discovery, gap analysis, remediation sprint, and C3PAO handoff.
00:07:24 --> 00:07:29
During remediation, the inference host is built and the model deployed, correct?
00:07:29 --> 00:07:39
Yes, the chosen checkpoint is installed, access controls are applied, audit logging is connected, and the system is validated against the SSP.
00:07:39 --> 00:07:42
Who is responsible for issuing the certificate?
00:07:42 --> 00:07:50
Only a C3PAO can issue a Level 2 certificate. An RPO prepares the contractor but cannot conduct the assessment.
00:07:50 --> 00:07:54
So the contractor’s C3PAO partner must be independent from Petronella?
00:07:55 --> 00:08:02
Exactly. Independence rules prohibit a provider from performing both preparation and assessment on the same client.
00:08:02 --> 00:08:05
What about change discipline when updating the model?
00:08:05 --> 00:08:14
Each new checkpoint is recorded with the developer name, license, and version. The model is checked against Section 1532 before deployment.
00:08:15 --> 00:08:17
And if a new version is released, how is it handled?
00:08:17 --> 00:08:25
The change is logged, the SSP is updated, and the new model undergoes the same validation process before going live.
00:08:25 --> 00:08:30
So the entire stack is mapped to NIST controls, and the boundary remains within the SSP.
00:08:30 --> 00:08:39
Yes, the inference host inherits the same protection as the rest of the enclave-segmented access, audit, encryption, and incident reporting.
00:08:40 --> 00:08:46
Now that we’ve covered the technical and compliance aspects, let’s talk about what organizations should do next.
00:08:46 --> 00:08:52
First, organizations should map the Gemma 4 checkpoint to the existing System Security Plan boundary.
00:08:53 --> 00:08:58
That means identifying which CUI assets will interact with the model and labeling them in the SSP.
00:08:59 --> 00:09:05
Once the boundary is defined, the next step is a paid scoping engagement to prototype on real data.
00:09:05 --> 00:09:11
During that prototype we measure context lengths, concurrency, and the accuracy required for the business use case.
00:09:12 --> 00:09:17
Those measurements feed directly into the GPU sizing calculation for the inference cluster.
00:09:18 --> 00:09:25
We saw that a 31B Dense model serving 256K tokens at moderate concurrency needs a single 80GB H100.
00:09:26 --> 00:09:32
In contrast, an E2B model can run on a phone or a Raspberry Pi with a 1GB mobile checkpoint.
00:09:32 --> 00:09:37
That flexibility lets a contractor span from field laptops to a hardened enclave server.
00:09:38 --> 00:09:47
After sizing, the cluster is isolated, typically on a segmented VLAN or an air-gap, to enforce the same controls as the rest of the system.
00:09:47 --> 00:09:56
Isolation ensures that only authorized traffic can reach the inference host, satisfying 3.13 System and Communications Protection.
00:09:56 --> 00:10:02
Once isolated, the open-weight model is downloaded under the Apache 2.0 license, with no click-through gate.
00:10:03 --> 00:10:07
That simplifies the legal review because the license is standard across the software stack.
00:10:08 --> 00:10:14
After download, the weights are installed on the inference stack, and role-based access control is defined.
00:10:14 --> 00:10:21
Role-based access maps to NIST 3.1 Access Control, ensuring only certain users can submit prompts.
00:10:22 --> 00:10:27
Audit logging is then wired to the same collection point used for other CMMC controls.
00:10:28 --> 00:10:34
This satisfies 3.3 Audit and Accountability, capturing who requested what and when.
00:10:34 --> 00:10:43
Encryption at rest and in transit uses FIPS-validated cryptography, as required by 800-171.
00:10:43 --> 00:10:49
The inference host also logs incident data for the 90-day preservation period mandated by DFARS.
00:10:49 --> 00:10:56
Because the model runs locally, every prompt, document, and output stays inside the enclave boundary.
00:10:56 --> 00:11:01
That eliminates the risk of data crossing into a third-party cloud, which would require additional agreements.
00:11:02 --> 00:11:08
The next concrete step is to document the model checkpoint in the SSP, listing its version and license.
00:11:09 --> 00:11:14
That documentation becomes part of the evidence repository during the C3PAO assessment.
00:11:14 --> 00:11:21
If a new checkpoint is released, the organization must log the change, update the SSP, and re-validate.
00:11:21 --> 00:11:27
Re-validation ensures the new model still meets the same NIST control mappings and performance metrics.
00:11:27 --> 00:11:34
Common mistakes include adding the model to the SSP without detailing the inference host’s segmentation.
00:11:34 --> 00:11:38
Another mistake is overlooking the cache memory requirement for long context windows.
00:11:39 --> 00:11:46
The KV cache grows with context length, so a 256K token request can double the VRAM usage.
00:11:47 --> 00:11:50
Failing to account for that can lead to out-of-memory errors during production.
00:11:51 --> 00:11:56
Another pitfall is assuming the Apache 2.0 license covers all downstream use cases.
00:11:56 --> 00:12:01
While it is permissive, you still need to verify that downstream partners can accept an open-weight model.
00:12:02 --> 00:12:09
Organizations should also review the model’s provenance, ensuring the download source is a trusted channel like Hugging Face.
00:12:09 --> 00:12:15
That way you can trace the weights back to the developer and confirm compliance with Section 1532.
00:12:15 --> 00:12:23
Because Gemma 4 is not covered by Section 1532, it poses no statutory restriction for DoD contracts.
00:12:23 --> 00:12:32
In contrast, the pending FY2027 bill could add Zhipu AI to that list, making GLM-5.3 risky.
00:12:32 --> 00:12:37
Export controls also apply to Zhipu, since its entities are on the BIS Entity List.
00:12:37 --> 00:12:43
That means any hosted-API usage would require an export license, but self-hosting avoids that.
00:12:43 --> 00:12:52
But the security record of GLM-5.3 raises additional concerns, as advisory reports note its agentic capabilities.
00:12:52 --> 00:12:58
Petronella’s approach is to keep the model local, so it stays under the same controls as the rest of the system.
00:12:58 --> 00:13:05
That also simplifies incident response because the logs are already integrated into the existing SOC workflow.
00:13:05 --> 00:13:09
Now, let’s discuss the hardware options for different business sizes.
00:13:10 --> 00:13:17
A small contractor might deploy the E2B model on a single 80GB H100 if they need high throughput.
00:13:17 --> 00:13:23
But if the workload is light, the 1GB mobile checkpoint can run on a Raspberry Pi without any GPU.
00:13:23 --> 00:13:35
For a mid-size contractor, the 12B Unified model offers 256K context and runs on a single RTX 5090 or H200 class card.
00:13:36 --> 00:13:45
The 31B Dense model is best suited for a data-center enclave with an 80GB H100 and 128GB unified memory per node.
00:13:45 --> 00:13:58
Our benchmarks showed 63.5 tokens per second on an RTX PRO 6000 Blackwell with 192GB DDR5 when running the 31B Dense model.
00:13:58 --> 00:14:06
When scaled to 32-way concurrency, the aggregate throughput reaches 2 tokens per second across the cluster.
00:14:06 --> 00:14:12
Those numbers illustrate the performance gap between a single GPU and a clustered environment.
00:14:12 --> 00:14:18
When choosing between Gemma 4 and GPT-OSS, consider modality and context requirements.
00:14:18 --> 00:14:24
Gemma 4 supports text, image, and native audio, which is useful for scanned documents and voice notes.
00:14:25 --> 00:14:30
GPT-OSS only handles text, so it may be sufficient for code analysis or structured drafting.
00:14:31 --> 00:14:43
However, GPT-OSS 120B offers 128K tokens, which may limit document size compared to Gemma’s 256K window.
00:14:44 --> 00:14:48
So the choice depends on the volume of text and the need for multimodal input.
00:14:48 --> 00:14:53
For defense contractors, the security posture of the model is as important as its performance.
00:14:54 --> 00:15:01
Because Gemma 4 is not covered by Section 1532, it avoids the statutory prohibition that applies to DeepSeek.
00:15:01 --> 00:15:08
If you adopt GPT-OSS, you still stay within the same compliance boundary because the model runs locally.
00:15:08 --> 00:15:16
But you must verify that the model’s license is Apache 2.0 and that no hidden clauses restrict use of CUI.
00:15:16 --> 00:15:24
The seven-stage method also includes a final validation step where the model’s outputs are compared against ground truth.
00:15:24 --> 00:15:29
That helps catch hallucinations and ensures the AI behaves as intended for mission-critical tasks.
00:15:30 --> 00:15:37
A common oversight is to skip that validation, assuming the model’s performance metrics alone guarantee reliability.
00:15:37 --> 00:15:41
That can lead to erroneous decisions in threat analysis or compliance reporting.
00:15:42 --> 00:15:47
During the remediation sprint, teams should also plan for model versioning and rollback procedures.
00:15:47 --> 00:15:52
Versioning ensures you can revert to a known good checkpoint if a new release introduces bugs.
00:15:53 --> 00:15:58
Rollback procedures should be documented in the SSP and tested in a staging environment.
00:15:59 --> 00:16:04
The Petronella team recommends using a separate staging cluster that mirrors the production hardware.
00:16:04 --> 00:16:09
That way you can validate the new model without impacting live CUI processing.
00:16:09 --> 00:16:15
After validation, the final step is to submit the evidence to the C3PAO for certification.
00:16:15 --> 00:16:23
The C3PAO will review the SSP, the audit logs, and the model documentation before issuing the certificate.
00:16:23 --> 00:16:30
Once certified, the contractor can confidently use Gemma 4 for CUI processing within the DoD ecosystem.
00:16:30 --> 00:16:37
If a new version is released, you repeat the validation and re-documentation steps to maintain compliance.
00:16:37 --> 00:16:42
That ensures the SSP remains current and the C3PAO can audit the system at any time.
00:16:42 --> 00:16:47
Another question we hear is whether the model can be fine-tuned on proprietary data.
00:16:48 --> 00:16:53
Fine-tuning is possible, but you need to keep the training pipeline inside the enclave and audit the process.
00:16:54 --> 00:17:00
The training data must also be classified as CUI and handled under the same NIST controls.
00:17:00 --> 00:17:05
If you use a third-party training service, that would require a separate data-processing agreement.
00:17:05 --> 00:17:11
Petronella’s approach is to build the fine-tuning pipeline on the same hardware used for inference.
00:17:11 --> 00:17:16
That keeps the entire lifecycle under your control and simplifies audit requirements.
00:17:16 --> 00:17:24
Finally, we recommend establishing a change management process that tracks model updates, version numbers, and licensing status.
00:17:24 --> 00:17:32
Recording those details in the SSP and evidence repository helps future assessors see the evolution of the AI stack.
00:17:32 --> 00:17:38
That completes the practical roadmap for deploying Gemma 4 under CMMC Level 2 compliance.
00:17:38 --> 00:17:42
Thanks for the deep dive into the compliance and deployment details.