GPT-OSS 20B and 120B run on your own hardware under Apache 2.0. Petronella Technology Group, Inc. deploys open-weight models inside compliance boundaries.
https://petronellatech.com/blog/using-gpt-oss-20b-and-120b-for-regulated-industries/
Chapters:
00:00 Introduction
00:14 Using GPT-OSS 20B and 120B for Regulated Industries
08:52 What organizations should do
16:36 How to reach us
A conversation about "Using GPT-OSS 20B and 120B for Regulated Industries" from the Petronella Technology Group, Inc. blog.
Subscribe to Encrypted Ambition and hear every episode: https://petronellatech.com/podcasts/
Questions about AI, cybersecurity, or compliance for your business? Call Petronella Technology Group, Inc. at 919-348-4912.
00:00:14 --> 00:00:20
Today we’re looking at OpenAI’s new GPT-OSS models and how they fit into regulated environments.
00:00:20 --> 00:00:30
The release on August 5, 2025 brought two open-weight models: gpt-oss-20b and gpt-oss-120b.
00:00:30 --> 00:00:35
They’re the first OpenAI models with downloadable weights since GPT-2 in 2019.
00:00:35 --> 00:00:41
Because the weights are on your premises, you can inspect, pin versions, and avoid external calls.
00:00:41 --> 00:00:47
But that alone doesn’t guarantee compliance, especially for PHI or Controlled Unclassified Information.
00:00:47 --> 00:00:51
Compliance attaches to the system around the model, not the model itself.
00:00:51 --> 00:00:56
So let’s dive into the technical details that drive those compliance decisions.
00:00:56 --> 00:01:04
gpt-oss-20b runs within 16 GB of memory, which means a single workstation GPU can handle it.
00:01:04 --> 00:01:08
That makes it a practical choice for team-level deployments like document summarization.
00:01:09 --> 00:01:20
In contrast, gpt-oss-120b requires a single 80 GB GPU, such as an NVIDIA H100, with MXFP4 quantization.
00:01:20 --> 00:01:28
If you skip quantization and run BF16 inference, you need roughly 240 GB, meaning two to four GPUs.
00:01:28 --> 00:01:33
So the hardware decision is the second key factor after choosing the model.
00:01:33 --> 00:01:35
Now, what about the legal side of things?
00:01:35 --> 00:01:46
Section 1532 of the FY2026 NDAA restricts use of only DeepSeek and High Flyer models during DoD contract performance.
00:01:46 --> 00:01:51
It does not name GPT-OSS, so self-hosting those weights is not prohibited by that statute.
00:01:52 --> 00:01:58
The law also says nothing about cloud versus on-premises or API versus downloaded weights.
00:01:58 --> 00:02:02
That clears up a common misconception that downloading a model makes it legal.
00:02:03 --> 00:02:08
But you still need to verify that the system around the model meets the applicable frameworks.
00:02:08 --> 00:02:13
For example, if you’re handling PHI, you’re under HIPAA’s Security Rule.
00:02:13 --> 00:02:21
HIPAA doesn’t certify software; it requires a risk assessment and controls that protect electronic protected health information.
00:02:22 --> 00:02:27
So you must document encryption, access controls, and audit logging around the inference stack.
00:02:27 --> 00:02:37
Similarly, CUI workloads for defense contractors fall under CMMC and DFARS 252-7012.
00:02:37 --> 00:02:42
Those frameworks demand FIPS-validated cryptography for data at rest and in transit.
00:02:42 --> 00:02:50
If you run gpt-oss-120b on a single 80 GB GPU, you’re already meeting the memory requirement.
00:02:50 --> 00:02:56
But that’s only the start; you still need a segmented VLAN or air-gap for the inference cluster.
00:02:56 --> 00:03:07
The article calls this the ‘private AI cluster pattern,’ using NVIDIA GB10 Grace Blackwell nodes with 128 GB unified memory each.
00:03:07 --> 00:03:14
Those nodes can be linked over a 400 G interconnect, giving you 256 GB for larger models.
00:03:14 --> 00:03:21
Those nodes can be linked over a 400 G interconnect, giving you 256 GB for larger models.
00:03:22 --> 00:03:26
However, you must still match the cluster size to the workload measured in a prototype.
00:03:26 --> 00:03:31
That’s why the article recommends a paid scoping engagement before buying any GPUs.
00:03:32 --> 00:03:38
The prototype should ingest a realistic set of documents, perhaps a few terabytes, to gauge concurrency needs.
00:03:38 --> 00:03:44
If you’re processing 40 TB, you’ll likely need a different build than a smaller company.
00:03:44 --> 00:03:49
Once you know the concurrency profile, you can size GPUs and add headroom for peak usage.
00:03:49 --> 00:03:56
Now let’s talk benchmarks and what they actually prove. This is especially true when dealing with structured outputs.
00:03:57 --> 00:04:05
On the OpenAI model card, gpt-oss-120b scores 92.5 on AIME 2025 competition math at high reasoning effort.
00:04:05 --> 00:04:12
The 20b version scores 91.7, showing similar reasoning capability in the same setting.
00:04:12 --> 00:04:18
Other benchmarks like MMLU and SWE-Bench show respectable scores, but they’re not the only measure you need.
00:04:18 --> 00:04:22
The real concern for regulated work is hallucination rates.
00:04:22 --> 00:04:31
On SimpleQA, gpt-oss-120b hallucinated on 78.2 percent of answers, with 16.8 percent accuracy.
00:04:31 --> 00:04:39
The 20b model’s hallucination rate jumps to 91.4 percent, with only 6.7 percent accuracy.
00:04:40 --> 00:04:44
That means you can’t rely on open-domain answers for clinical or legal questions.
00:04:44 --> 00:04:48
Retrieval grounding and human review become mandatory, not optional.
00:04:49 --> 00:04:54
Structured outputs help by forcing the model to fit responses into a JSON schema you define.
00:04:54 --> 00:04:59
The chain-of-thought visibility also lets reviewers see how the model arrived at its answer.
00:04:59 --> 00:05:05
But even with those guardrails, the compliance team must verify the entire stack against the controls.
00:05:05 --> 00:05:16
The article outlines a seven-stage deployment method that starts with defining the data boundary. This ensures that any policy or regulatory requirement is explicitly mapped.
00:05:17 --> 00:05:22
You need to map whether data is CUI, PHI, or financial, and which frameworks apply.
00:05:23 --> 00:05:27
Then you run a paid scoping engagement that builds an MVP on your own data.
00:05:28 --> 00:05:32
The prototype’s performance informs the GPU sizing and the isolation requirements.
00:05:33 --> 00:05:38
Next step is to isolate the inference cluster on a segmented VLAN or full air-gap.
00:05:39 --> 00:05:43
That keeps prompts and outputs inside the SSP boundary you already control.
00:05:43 --> 00:05:47
After isolation, you deploy the open-weight models on an inference stack you own.
00:05:48 --> 00:05:55
Layer role-based access, encryption at rest and in transit, and audit logging to map to framework controls.
00:05:55 --> 00:05:59
Then you validate the stack against those controls before production.
00:05:59 --> 00:06:04
The article emphasizes that the seven-stage method is a checklist for compliance readiness.
00:06:04 --> 00:06:09
If you skip any of those steps, you’ll likely fail a CMMC audit or a HIPAA assessment.
00:06:10 --> 00:06:20
For example, if encryption at rest isn’t FIPS-validated, DFARS 252-7012 would consider that a violation.
00:06:20 --> 00:06:31
Similarly, a model that logs prompts to an uncontrolled share would break the audit trail requirement. This can expose the organization to fines or audit findings.
00:06:31 --> 00:06:42
The article also notes that no model is certified for HIPAA or CMMC, so the system’s controls are what matter. A documented risk assessment also aids in incident response planning.
00:06:42 --> 00:06:51
The article cites early adopters like AI Sweden, Orange Business, and Snowflake who hosted on premises for data-security reasons.
00:06:51 --> 00:06:59
Those examples show that private deployment is feasible for large enterprises. It demonstrates to auditors that you have intentional controls in place.
00:07:00 --> 00:07:04
But they also illustrate the need for a clear hardware sizing plan before you buy.
00:07:05 --> 00:07:10
If you underestimate concurrency, you’ll end up with a single GPU that can’t keep up with real-world usage.
00:07:11 --> 00:07:20
Conversely, over-provisioning wastes capital and increases maintenance overhead. Proper sizing also reduces risk of model downtime during peak demand.
00:07:21 --> 00:07:25
The article also discusses the importance of a data boundary that aligns with the control frameworks.
00:07:26 --> 00:07:31
For CUI, that boundary is often an enclave that only authorized users can access.
00:07:31 --> 00:07:37
For PHI, the boundary is the HIPAA-compliant EHR and any attached data lakes.
00:07:38 --> 00:07:44
For financial data, the boundary may be a dedicated data warehouse with strict role-based access.
00:07:44 --> 00:07:52
Once you have that boundary, you can map the defense-industry or healthcare frameworks onto it. This alignment is critical for maintaining compliance during scaling.
00:07:53 --> 00:08:01
That mapping informs which NIST SP 800-171 controls you must implement around the inference node.
00:08:01 --> 00:08:06
The article stresses that FIPS-validated cryptography is required for CUI at rest.
00:08:06 --> 00:08:11
Encryption that is strong but not validated does not satisfy the regulation as written.
00:08:12 --> 00:08:17
That’s why you need to document the cryptographic algorithms and their validation status in the SSP.
00:08:17 --> 00:08:24
Another key point is that even if the model is on-premises, the data still needs to stay within the enclave.
00:08:24 --> 00:08:29
If a prompt leaks outside, you violate the data boundary and potentially breach the regulation.
00:08:30 --> 00:08:35
The article also highlights that the model’s chain-of-thought visibility can be logged for audit purposes.
00:08:36 --> 00:08:40
Those logs provide evidence that an assessor can review how a decision was made.
00:08:40 --> 00:08:52
So the next step is to decide how your organization will validate all these controls before putting the model in production. Only after successful validation can you consider a phased rollout.
00:08:52 --> 00:08:57
Now that we’ve mapped the boundaries, we should look at the deeper implications of running these models.
00:08:57 --> 00:09:10
Because the hallucination rate is 78.2% for gpt-oss-120b and 91.4% for gpt-oss-20b, any production use must ground answers.
00:09:11 --> 00:09:15
Grounding means pulling in verified documents and feeding them back into the inference loop.
00:09:15 --> 00:09:22
That retrieval layer must be part of the system, not an afterthought, otherwise you risk unverified claims.
00:09:22 --> 00:09:26
And the chain-of-thought is a built-in feature that lets you see exactly how the model reasoned.
00:09:27 --> 00:09:34
Logging that chain and the prompt gives auditors a clear audit trail, which is essential for CMMC and HIPAA.
00:09:34 --> 00:09:40
But remember, the model itself isn’t certified; the whole stack must pass validation.
00:09:40 --> 00:09:47
That means you need to document the SSP, map controls, and run a scoping prototype before buying GPUs.
00:09:48 --> 00:09:53
The prototype should ingest a realistic volume, like 40 TB, to see how many GPUs you’ll need.
00:09:54 --> 00:10:05
For gpt-oss-120b, a single 80 GB GPU with MXFP4 can serve one steady stream, but concurrency pushes you to two or four.
00:10:05 --> 00:10:12
If you skip quantization and run BF16, you’ll need roughly 240 GB, meaning two to four GPUs.
00:10:12 --> 00:10:18
That hardware decision is the third key choice, after the model and the compliance boundary.
00:10:18 --> 00:10:26
Now let’s talk cost. The article says that at 500 tokens per day, private deployment breaks even in 6 to 12 months.
00:10:27 --> 00:10:35
At 5 million tokens daily, you can cut API spend 60 to 80 percent, which is a huge advantage for large enterprises.
00:10:35 --> 00:10:43
But the upfront GPU cost can be high, especially for gpt-oss-120b, but the long-term savings can justify it.
00:10:44 --> 00:10:51
You should also factor in maintenance, power, cooling, and the cost of a qualified data-scientist to tune the model.
00:10:51 --> 00:10:55
Another common mistake is assuming the benchmark numbers are enough proof of safety.
00:10:56 --> 00:11:02
Benchmarks show reasoning ability, but they don’t measure hallucination against your specific documents.
00:11:02 --> 00:11:06
You need to run your own SimpleQA-style test on your data to see the real error rate.
00:11:06 --> 00:11:13
If you find 70% or more hallucinations, you must add a human-in-the-loop or a stricter grounding layer.
00:11:14 --> 00:11:21
Regarding Section 1532, the only covered models are DeepSeek and High Flyer, so GPT-OSS is not prohibited.
00:11:21 --> 00:11:27
But the law limits use during DoD contract performance; it does not ban commercial use outside that scope.
00:11:28 --> 00:11:33
So if your company is only handling private data, you’re in the clear-just keep documentation up to date.
00:11:34 --> 00:11:41
However, if you ever sign a DoD contract, you must audit the model’s origin and keep it out of that performance window.
00:11:41 --> 00:11:49
Let’s talk about FIPS validation. The article says you need FIPS-validated cryptography for CUI at rest and in transit.
00:11:49 --> 00:11:55
Using a non-validated algorithm, even if it’s strong, would still violate the regulation as written.
00:11:55 --> 00:12:03
That means you should use AES-256-GCM or similar FIPS-validated ciphers for encryption of weights and logs.
00:12:03 --> 00:12:13
Also remember that encrypted CUI remains CUI, so it still requires the same FedRAMP or DFARS controls if you move it to the cloud.
00:12:13 --> 00:12:19
For HIPAA, the security rule demands a risk analysis and mitigations, not a list of certified software.
00:12:19 --> 00:12:31
So you’ll need to show that your inference node, access controls, and audit logs satisfy the 164(a)(1)(ii)(A) requirements.
00:12:31 --> 00:12:36
Now, what about the practical steps you can take today to start the validation process?
00:12:36 --> 00:12:43
First, assemble a cross-functional team-security, compliance, data science, and operations-to own the SSP.
00:12:44 --> 00:12:50
Second, document the data boundary for each regulatory framework and map the required NIST controls.
00:12:51 --> 00:12:58
Third, run a prototype that ingests a realistic subset of your documents and measures latency and GPU usage.
00:12:58 --> 00:13:03
During that run, capture the hallucination rate and the accuracy against a gold-standard set.
00:13:04 --> 00:13:10
If the error exceeds your tolerance, iterate on retrieval, prompt design, or add a human review step.
00:13:10 --> 00:13:16
Fourth, size the hardware based on the prototype’s concurrency and document the GPU inventory in the SSP.
00:13:17 --> 00:13:25
Fifth, isolate the inference cluster on a segmented VLAN or an air-gap, and restrict egress to only necessary endpoints.
00:13:25 --> 00:13:33
Sixth, implement FIPS-validated encryption for the weights, logs, and any data at rest, and use TLS for in-transit.
00:13:33 --> 00:13:41
Seventh, set up logging of prompts, retrieved passages, outputs, and reviewer decisions, and feed that into your audit system.
00:13:41 --> 00:13:47
Eighth, run a formal security assessment or third-party audit to verify that every control is functioning.
00:13:48 --> 00:13:55
Once validated, you can move to a phased rollout, starting with non-critical documents and scaling up gradually.
00:13:55 --> 00:13:59
Now, let’s cover some questions listeners often ask about GPT-OSS.
00:13:59 --> 00:14:07
First, is GPT-OSS HIPAA compliant? No, the model itself isn’t; compliance comes from the system around it.
00:14:07 --> 00:14:12
Second, what hardware do I need for gpt-oss-120b on a single stream?
00:14:13 --> 00:14:23
You need one 80 GB GPU, like an H100, with MXFP4 quantization; BF16 would need roughly 240 GB.
00:14:24 --> 00:14:28
Third, does the model’s Apache 2.0 license affect compliance?
00:14:28 --> 00:14:36
The license allows you to download and run the weights on your own hardware; it does not grant any regulatory certification.
00:14:36 --> 00:14:40
Fourth, can I use GPT-OSS for clinical drafting workflows?
00:14:41 --> 00:14:49
HealthBench scores show gpt-oss-120b is capable, but you must ground outputs, review them, and log the review.
00:14:50 --> 00:14:55
Fifth, how does Section 1532 impact me if I’m not a DoD contractor?
00:14:55 --> 00:15:02
You’re not bound by the clause; it only applies to DoD contract performance, so the model’s origin is irrelevant to you.
00:15:03 --> 00:15:06
Sixth, what about cost savings versus API usage?
00:15:07 --> 00:15:16
If you hit 5 million tokens per day, you can reduce API spend by 60 to 80 percent, which outweighs the GPU cost over time.
00:15:16 --> 00:15:20
Seventh, what are the biggest pitfalls when deploying GPT-OSS?
00:15:20 --> 00:15:31
Common mistakes include ignoring hallucination rates, over-estimating GPU capacity, skipping FIPS validation, and neglecting to log prompts and reviews.
00:15:31 --> 00:15:38
Another pitfall is assuming the model’s chain-of-thought is enough; you still need human oversight for critical decisions.
00:15:38 --> 00:15:47
Also, don’t forget that the model’s API or inference node must stay within the SSP’s network boundaries to avoid data leakage.
00:15:47 --> 00:15:51
Finally, what is the recommended way to handle model updates or version pinning?
00:15:51 --> 00:16:01
Download the weights, record the SHA-256 hash, and lock the version in your SSP; treat any update as a new configuration item.
00:16:01 --> 00:16:05
Do you have any final tips for organizations starting with GPT-OSS?
00:16:05 --> 00:16:13
Treat the model as a tool, not a replacement; always validate, audit, and document every step, and keep the human in the loop.
00:16:13 --> 00:16:17
That wraps up our deep dive into GPT-OSS compliance and deployment.
00:16:18 --> 00:16:25
Remember, the key is to build a secure, auditable, and validated inference stack that meets your regulatory framework.
00:16:25 --> 00:16:30
Thank you for the thorough walk-through; your expertise really helps demystify the process.
00:16:30 --> 00:16:35
It was a pleasure to share the practical steps and clarify the compliance nuances.
00:16:35 --> 00:16:37
Thank you for sharing your insights.