Author
Listed:
- Nagabhushanam Bheemisetty
(Independent Researcher)
Abstract
This paper presents an innovative five-layered architecture for finding equitable representation across the various fragmented sections of large datasets. The architecture employs a set of customizable criteria to group together similar datasets into well-balanced unbiased sample sets; while retaining complete and accurate metadata for all file boundary elements. Important features of this architecture include: a multi-level criteria engine which is capable of processing accepted and rejected streams from a variety of sources; a stratified sampling process which effectively eliminates the presence of partition skew; and a quality assurance mechanism that provides an efficient means to capture and report performance-related metrics. Benchmark partitioning results indicate a significant shift toward overall balance ratio improvements; with improvements of approximately 1.75 to 1.12 ratio, as well as an overall reduction in volume anomaly of 87% (from ±25% down to ±3.2%), and total absence of schema drift occurred at a cost of less than 5% of the overall dataset processing costs to the author(s). By implementing this method, the processing of large datasets has been made much more reliable than previously possible, while simultaneously minimising bias and maximising the ability to provide efficient audit trial capabilities. Future enhancements to the architecture will focus on providing distributed execution capabilities in conjunction with streaming integration into enterprise data systems.
Suggested Citation
Nagabhushanam Bheemisetty, 2026.
"Handling Criteria-Driven Filtering &Sampling Across Distributed Data Partitions,"
International Journal of Latest Technology in Engineering, Management & Applied Science, RSIS International, vol. 15(2), pages 1384-1393, February.
Handle:
RePEc:bjf:ijltem:v:15:y:2026:i:2:a:2133
DOI: 10.51583/IJLTEMAS.2026.15020000122
Download full text from publisher
Corrections
All material on this site has been provided by the respective publishers and authors. You can help correct errors and omissions. When requesting a correction, please mention this item's handle: RePEc:bjf:ijltem:v:15:y:2026:i:2:a:2133. See general information about how to correct material in RePEc.
If you have authored this item and are not yet registered with RePEc, we encourage you to do it here. This allows to link your profile to this item. It also allows you to accept potential citations to this item that we are uncertain about.
We have no bibliographic references for this item. You can help adding them by using this form .
If you know of missing items citing this one, you can help us creating those links by adding the relevant references in the same way as above, for each refering item. If you are a registered author of this item, you may also want to check the "citations" tab in your RePEc Author Service profile, as there may be some citations waiting for confirmation.
For technical questions regarding this item, or to correct its authors, title, abstract, bibliographic or download information, contact: Dr. Pawan Verma (email available below). General contact details of provider: https://www.ijltemas.in/ .
Please note that corrections may take a couple of weeks to filter through
the various RePEc services.