Why don’t they marry the visual categorisation of objects with a sparse predefined model to simplify?
Point clouds seem like a noisy layer to use and better as verification.