Author
Listed:
- Ardon Kotey
- Allan Almeida
- Hariaksh Pandya
- Arya Raut
- Rayaan Juvale
- Vedant Jamthe
- Tejan Gupta
- Hemaprakash Raghu
- Naman Gupta
- Lalith Samanthapuri
Abstract
Within the field of legal AI, named entity recognition, also known as NER, is an essential step that must be completed before moving on to subsequent processing stages. In this paper, we present the creation of a dataset for the purpose of training natural language understanding models in the legal domain. The dataset is produced by locating and establishing a complete set of legal entities, which goes beyond traditionally employed entities such as person, organization, and location. These are examples of commonly used entities. Annotators are now provided with the means to effectively tag a wide variety of legal documents thanks to these additional entities. The authors tried out several different text annotation tools before settling on the one that proved to be the most effective for this study. The completed annotations are saved in the JavaScript Object Notation (JSON) format, which makes the data more readable and makes it easier to manipulate the data. The dataset that was produced as a result includes approximately thirty documents and five thousand sentences. Following that, these data are use in order to train a pre-trained SpaCy pipeline for accurate legal named entity prediction. There is a possibility that the accuracy of legal named entity recognition can be improved by performing additional fine-tuning on pre-trained models using legal texts.
Suggested Citation
Ardon Kotey & Allan Almeida & Hariaksh Pandya & Arya Raut & Rayaan Juvale & Vedant Jamthe & Tejan Gupta & Hemaprakash Raghu & Naman Gupta & Lalith Samanthapuri, 2023.
"NER Based Law Entity Privacy Protection,"
International Journal of Scientific Research in Computer Science, Engineering and Information Technology, International Journal of Scientific Research in Computer Science, Engineering and Information Technology, vol. 9(6), pages 322-335, December.
Handle:
RePEc:jbh:ijsrcs:v9:y2023:i6:id:hcseit2390665
DOI: 10.32628/CSEIT2390665
Note: Article URL: https://ijsrcseit.com/CSEIT2390665
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:jbh:ijsrcs:v9:y2023:i6:id:hcseit2390665. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Pankaj Sharma (USA) (email available below). General contact details of provider: https://ijsrcseit.com/home .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.