Article Archives
Article Categories
Articles
Can you rely on your structured data to “teach” your AI?
Can you rely on your structured data to "teach" your AI?
By Cheryl Ahrens Young, IGP, CSM, CSPO, CTT+, ERMm, ECMp
[email protected]
The tenets of records management are to provide a reliable, authentic record with auditable integrity. The assumption that unstructured records (a contract, for example) are less reliable than structured data (databases, for example) needs to be investigated more thoroughly. In the past few years, I have been involved in a number of projects where the underlying assumption that the data held in line of business applications, such as Oracle and SAP, were accurate and complete. This assumption proved to be costly, in terms of both time and money, as projects were delayed and additional resources needed to identify and remove real duplicates or rename non-duplicated information and to standardize naming conventions so that the project could move forward, often after the go-live date.
There are so many adages that apply – “pay me now or pay me later”, “assume = ass u (and) me”, “garbage in/garbage out”, it’s almost funny (almost). If you have lived through a project where an underlying assumption proved false, you understand the pain. In many cases, the test environment uses a database of made-up data, so as to not compromise security requirements. In any project, though, live data is used in the production environment. My experience has been that the data needs to be thoroughly examined, and tested, before the go-live date in order to celebrate a successful implementation.
How could the involvement of a records management professional at the inception of creating a new database or rollout of a new project avoided this pain? We’re trained to create classification plans that make business sense, understand the implications of abbreviations and acronyms, and, know that a record will lose its value over time and will need to be discarded. Dealing with the here and now, in migration or business process automation projects, we also understand the importance of cleaning and purging information during the development stages, prior to testing, so that the testing is a reflection of the usefulness of the application and not the integrity of the data.
During the Summer Conference this year, the importance of the profession was emphasized by the speakers in topic ranging from Tyrene Bada’s RIM 101 to Andrew SanAgustin’s Where is the RIM/IG Profession headed and a great panel discussion on RIM in the Real World. A common theme: ROT exists! Records retention schedules are necessary to defensibly dispose of that ROT. Getting rid of the ROT improves the overall results in both AI and human production.
Another lesson learned: AI prompts and any resulting transactions are records! Update your retention schedule to include these!
Is Your Organization AI Ready? (Part II)
Is Your Organization AI Ready?
By Cheryl Ahrens Young, IGP, CSM, CSPO, CTT+, ERMm, ECMp
[email protected]
An AI is only as good as the information in its LLM.
AI developers assume that the large language models built from existing databases and systems’ of record metadata are accurate, factual, reliable and have data integrity. My experience in business process automation with learning servers has taught me that assumptions lead to “ass: u & me”.
You will often see a disclaimer in content created by commercial “free” AI that the information provided may not be accurate.
What is the root cause for the inaccuracies in an AI generated document? For AI hallucinations?
ROT – Redundant, Obsolete and Trivial records fed to the LLM. Added to that, lack of integrity in the information – abbreviations, mis-spellings, homonyms that mean entirely different things (accept vs except), and acronyms that are industry specific (CRM could be Certified Records Manager or Customer Relations Management).
How do you clean up the ROT and improve the integrity of the LLM?
A legally defensible, robust records and information management program which addresses not only retention but integrity, accuracy and accountability. This all starts with a current records retention schedule that is then applied to the systems of record, shared drives and any other source of information that builds the LLM. In addition, processes and procedures outline how to name records and enter consistent data into the systems of record.
Who knows about records retention schedules and naming conventions?
Records and Information Managers!
The question is – how do we get a seat at the AI table?
Where I’ve had success in getting the technical team to listen is inviting them to a lunch and learn meeting and talking about RIM and knowledge management. If your company follows Agile or Scrum project methodologies, ask to be part of the team to develop the customer stories. Your story might revolve around an AI identifying duplicate and near duplicate copies across shared drives and systems of record to reduce the ROT to then provide the best possible information for a iterative process improvement of the information fed to the LLM.