When modelling a gRPC library in Go, I noticed that steps added to global taint tracking via the isAdditionalFlowStep predicate do not cover cases where the object passed to the step has tainted content, but is not considered tainted itself.
Example code for the issue below:
type Container = struct{ Value int }
func link_in(Container) {
// magic CodeQL does not understand
}
func link_out() Container {
// more magic
// returns the argument of link_in
panic(0)
}
func main() {
a := Container{Value: taint()}
link_in(a)
a_ := link_out()
sink(a_.Value)
}
The link_in and link_out functions represent two places that pass data between them, but which CodeQL does not connect out of the box. Adding an additional flow step to establish this connection will only work if we consider the entire Container struct to be tainted; the taint on Container.Value in the example above is ignored.
After experimenting a bit, this behaviour seems consistent across different languages, which leads me to believe it was an intentional decision. However, I do not fully understand the reasoning behind it. Many libraries wrap their data in some kind of container (for example, the code emitted by the Go gRPC tooling generates custom structs for the message types in the protobuf), and the mechanisms provided by flow summaries are not always sufficient to model this (The above example can be solved by introducing a write/read to a synthetic global variable unique to this function pair, however that only works if the functions with this special behaviour are known in advance).
Since flow summaries can handle tainted content, would it be possible to do the same for isAdditionalFlowStep? Dropping the taint to the level of the entire containing struct is not always a worthwhile option, so this mechanism would be very helpful when modelling more complex library behaviour.
If this is a deliberate design decision, I feel like it should be mentioned in the data-flow tutorial.
When modelling a gRPC library in Go, I noticed that steps added to global taint tracking via the
isAdditionalFlowSteppredicate do not cover cases where the object passed to the step has tainted content, but is not considered tainted itself.Example code for the issue below:
The
link_inandlink_outfunctions represent two places that pass data between them, but which CodeQL does not connect out of the box. Adding an additional flow step to establish this connection will only work if we consider the entireContainerstruct to be tainted; the taint onContainer.Valuein the example above is ignored.After experimenting a bit, this behaviour seems consistent across different languages, which leads me to believe it was an intentional decision. However, I do not fully understand the reasoning behind it. Many libraries wrap their data in some kind of container (for example, the code emitted by the Go gRPC tooling generates custom structs for the message types in the protobuf), and the mechanisms provided by flow summaries are not always sufficient to model this (The above example can be solved by introducing a write/read to a synthetic global variable unique to this function pair, however that only works if the functions with this special behaviour are known in advance).
Since flow summaries can handle tainted content, would it be possible to do the same for
isAdditionalFlowStep? Dropping the taint to the level of the entire containing struct is not always a worthwhile option, so this mechanism would be very helpful when modelling more complex library behaviour.If this is a deliberate design decision, I feel like it should be mentioned in the data-flow tutorial.