New training data teaches AI to understand the forestides valuable data for forestry

Forestry information systems

M I S T R A D I G I T A L F O R E S T 2 0 2 5 H I G H L I G H T S

At a time when the forestry industry is undergoing rapid digitalisation, access to high-quality data is crucial. A recently published PhD thesis shows that harvester data and artificially generated information can be used to train AI-based decision support systems to classify trees and detect quality deviations.

Raul De Paula Pires
Raul de Paula Pires. Photo: Private.

There is a heated debate about how AI can contribute to more resource-efficient and sustainable use of the forest. At the same time, there is not enough detailed and site-specific training data to make such a development possible. In November, Raul de Paula Pires defended his thesis at SLU. Supported by Mistra Digital Forest, his research explores two new potential data sources that could be used to train AI-based decision support systems for forest inventory.

In this context, the harvesting data already collected by forestry companies has proved to be a goldmine. By combining this data source with airborne laser scanning data, an AI model was built which was able to classify tree species with a high degree of accuracy.

- It is particularly exciting that the AI model maps the value of the forest using relatively little training data, and that the forestry companies already have access to this type of data. For them, it is both free and up to date, says Raul de Paula Pires.

Synthetic data – when reality isn’t enough

Certain phenomena in the forest are rare enough that it would be extremely time-consuming to collect sufficient training data by means of traditional field surveys. One such example is crooked trees – they account for just over two per cent of all trees but they have a significant impact on timber value. Raul de Paula Pires, together with forestry machinery manufacturer Komatsu Forest, demonstrated that synthetic data can be used to train AI to identify crooked trees. In this case, it involved training the AI on thousands of computer-generated trees.


– Synthetic data is useful for training AI in everything from classifying tree species under varying weather conditions, to detecting unusual quality defects such as crookedness and fungal growth. We have a long, positive history of collaboration with universities and research institutes, and it is not at all uncommon that these partnerships form the foundations for various types of internal development projects, says Johan Fransson, Team Leader Strategic Research at Komatsu Forest.

Now these research findings on new potential training data have been linked to another project where Komatsu Forest, SLU and Skogforsk are investigating how information on crookedness can optimise the cutting of tree trunks.

– By diversifying its toolbox and utilising more data sources, the forestry industry can leverage AI to make well-informed decisions regarding forest management and nature conservation, Raul de Paula Pires concludes.

Captions:

Hero: The Komatsu Forest 951XC, the machine model used in the project tests.