BetaEntity Annotation Prototype
← Back to institutions

Annotated abstract

8274428 Small Language model based NAICS 2022 classification of Industry from products and services: CLIPS

oemed · 2025-10-06 · canonical JSON source

2 visible annotations · policy: published · automated confidence ≥ 75.00%

Document resource

Objective Numerous tools are now available that use natural language processing to automatically code free-text job descriptions to standardized occupation classification systems. However, the development of tools for coding industry has received much less attention. Here we describe CLIPS, a tool we developed to identify plausible standardized industry codes based on free-text responses to the question ‘what product is made or services provided.’ This tool can be incorporated into online questionnaires to help participants self-code their industry.Material and Methods CLIPS uses small language models to code free-text industry information to NAICS 2022 codes. It uses a two-step process that first converts (embeds) the text information to numbers and then classifies using a dense classification neural network. Employer name was excluded as an input feature, based on concerns over privacy and the difficulty of obtaining training data. CLIPS provides a score for each of the 689 NAICS 5-digit codes. The industry codes with the highest scores are then included in a list from which the study participant selects the best fit. We validated CLIPS using 1,586 jobs coded to NAICS 2022 by an expert.Results Overall, the industry code with the highest score from our preliminary version of CLIPS had a 47.4% agreement with the expert-assigned codes in our validation data set. The CLIPS score predicted agreement with the expert-assigned code. In addition, the expert-assigned code was in the top 3 CLIPS-suggested codes for 66% of the jobs, in the top 6 for 75%, and in the top 10 for 80%.Conclusion CLIPS’ ability to identify the expert-assigned code in its highest scoring codes makes it suitable for assisting participants in self-coding their industry. We hope to expand the training data to further improve its performance.Funding This work is funded by the Intramural Research Program of the US National Cancer Institute, Division of Cancer Epidemiology and Genetics.