When a factory robot governed by a generative AI model swings an arm toward a target, its path can respond to environmental conditions. It may not hit a person. The same constraint appears in other less dramatic applications that fail in the same way. An edited portrait may change the light or the expression; it may not become someone else. A simulated physical field may take an unexpected shape; it may not leave the envelope the equipment was designed for. In each case, the generator is free to propose. The application is not free to accept a proposal that violates the rule.[1]
That gap is the subject of a paper titled “HardFlow: Hard-Constrained Sampling for Flow-Matching Models via Trajectory Optimization” by researchers Zeyang Li, Kaveh Alim, and Navid Azizan. The paper asks a narrow question with broad consequences: can a pretrained generative model retain the freedom to act that makes it useful, while the sampler guarantees that the result lies inside a hard constraint set?
Hard constraints are mandatory conditions on the finished object returned by a model, called a sample. For example, a robot trajectory must reach its target without intersecting an obstacle, and an edited image of a face must remain the same person under a fixed identity test. A sample that misses either condition is unusable, regardless of how close it is to a valid sample.
The hard-constraint requirement is difficult for most AI models to meet because they do not produce the finished object in one step. A flow-matching model starts with noise and carries it, a little at a time, along a learned velocity until the noise has become an image, a trajectory, or a field; diffusion models do the same work under a different formulation. Only the last state is the sample the user will keep. The states before it is scaffolding, and they need not obey the rule that will later be applied to the product. Most constrained samplers ignore the distinction. They treat every point on the path as if it were already the thing that must be legal, which is a stricter demand than the application makes and a less efficient use of the model.
Projection methods try to keep the evolving sample inside permissible bounds as it is generated. Some do this after every small step, some only near the end, some on a tightening schedule. The impulse is to stay safe throughout. The mistake is that the early steps are still mostly noise, not a draft of the final robot path or image. Forcing those noisy steps to already obey the final rule discards paths that would have become permissible only when the sample is finished, which is the only moment the application will check. The methods also tend to quit as soon as they have a valid answer, even when a shorter path, a cheaper field, or a stronger edit was available.
Soft guidance makes the other mistake. Instead of requiring the finished sample to obey the rule, it offers extra credit for getting closer. The model can then produce a better-looking image or a closer match to the prompt and still break the constraint. That may be acceptable when the only aim is appearance. It is not acceptable when a broken rule makes the sample unusable. What’s left is a choice between a tightly steered model whose outputs get worse and an unsteered model whose outputs are illegal.
HardFlow leaves the pretrained model unchanged and intervenes only while drawing a sample. It does not force every intermediate step to be legal. As the sample evolves from noise toward data, the method adds a correction to the model’s ordinary step. It penalizes large corrections, so the result stays close to what the unconstrained model would have produced. The correction can also shorten a path or save energy. The requirement that cannot be relaxed is that the finished sample obey the constraint. The researchers do this separately for each random start, because computing one steering rule for every possible state is impossible at the size of these models.
HardFlow keeps the finished sample legal by shrinking the steering problem instead of relaxing the rule. It does not plan the rest of the path at once. At each step, it asks a computationally cheaper question: after this correction, where does the model think the sample will end up? It uses that guess to judge both legality and quality, so it doesn't simulate the entire future at every step.
Choosing the next noisy state so the network’s prediction of the end is legal is a messy calculation. HardFlow works the other way. It chooses a predicted finished sample that already obeys the constraint, then takes one backward step to find the next state. At the last step, the prediction and the sample are the same object, so the point returned to the user is legal by construction.
The half-finished sample is allowed to leave the legal region. The finished sample is not. The research is explicit about what that shortcut costs. Looking only one step ahead, the method misjudges where the sample will truly end and ignores corrections it has not made yet. The one-step backward map also means the secondary aim, a shorter path, lower energy, a stronger edit, is optimized only approximately. Those errors are bounded. The constraint on the sample delivered to the user is not.
The experiments test that balance. In a robot-arm study, the model had to generate short motions that reached a target, obeyed the arm’s dynamics, avoided obstacles seen in training, and avoided new obstacles shown only at test time. Left alone, the model was safe in 6 percent of trials. Soft guidance stayed in the teens. Projection reached 76 percent and, when it succeeded, took longer routes. HardFlow was safe in every trial, reached the target in the fewest steps, and did not need the runtime of the heaviest combined baselines.
A longer maze-navigation task produced the same ranking. HardFlow was the only method with no constraint violations and the highest task score. A third test asked the model to control a Burgers equation, a standard one-dimensional flow problem, subject to bounds that moved over time, a discretized form of the physics whose viscosity was known only inside an interval, and a preference for low control energy. Projection could also keep every sample inside the bounds. Among the methods that did, HardFlow used the least energy.
A fourth experiment asked whether the same method could protect identity in an image edit. A model trained on celebrity faces had to follow a text prompt while remaining perceptually close to the source photograph, as measured by LPIPS, a standard distance between images, held below a fixed ceiling. Guidance that chased only the prompt often changed the person. Methods that pushed the image much closer than the ceiling required often preserved identity by editing very little and sometimes produced weak pictures with tidy scores. HardFlow stayed under the bound on every test image, matched the most aggressive editors on prompt score, and used about a third of the computation. The ceiling limited how far the edit could drift. It did not slam the image onto the nearest legal copy of the original.
The practical finding is that the original model can keep doing what it was trained to do: propose new samples. You can add the rules of a particular job later, when a sample is drawn, and apply them to the finished object rather than to every noisy step along the way. One trained network can then be reused under different rules: collision-free robot motion in one setting, bounds on a physical field in another, identity in an image edit.
The HardFlow method comes with a computational cost. It does the extra work when a sample is drawn, rather than building the rule into the trained network. Early on, the model’s guess of the finished sample is less reliable, so the researchers often wait until later steps before running the extra optimization. If the pretrained model is weak, or if the legal region is empty or badly specified, no amount of steering will rescue the result. They suggest later versions that put some of the constraint work into training, that reuse part of the runtime computation so deployment is cheaper, and that try the same idea on larger generators and on robot tasks that involve contact and vision, where forcing the wrong quantities to be legal would be more damaging.
None of that will invent a model never learned. HardFlow changes only the conditions under which a usable model is allowed to stop. Generation keeps the freedom that lets these systems propose new samples. The returned sample must lie in the set. A nearby point that does not lie in the set remains unusable. That is what HardFlow enforces.
This article used AI to craft the contents and is shared at no charge for educational and informational purposes only.
Red Sky Alliance is a Cyber Threat Analysis and Intelligence Service organization. We provide indicators of compromise information (CTI) via a notification/Tier I analysis service (RedXray) or an analysis service (CTAC). For questions, comments, or assistance, please contact the office directly at 1-844-492-7225 or feedback@redskyalliance.com
- Reporting: https://www.redskyalliance.org/
- Website: https://www.redskyalliance.com/
- LinkedIn: https://www.linkedin.com/company/64265941
Weekly Cyber Intelligence Briefings:
REDSHORTS - Weekly Cyber Intelligence Briefings
https://attendee.gotowebinar.com/register/7855487668891299929
[1] https://six3ro.substack.com/p/when-the-last-step-is-the-only-one
Comments