Awarded contract

Enriching judgments and tribunal decisions, entities and hyperlinks for Linked Data, integration

Details

Supplier(s)
MDRX LLP
Value
GBP 74,920
Published
20 September 2022

Tender description

Summary of the work Integration between the enrichment service/NLP pipeline and the Find Caselaw service. Data enrichment to turn the references in judgments to other cases and legislation into hyperlinks. Enrich the documents, by identifying citations and references to other named entities. Expected Contract Length 2 months (with option to extend by a further 2 months) Latest start date Monday 22 August 2022 Budget Range Up to £75,000 (ex VAT) or £90,000 (inc. VAT) Why the Work is Being Done The National Archives manages legislation.gov.uk and is developing a service ‘Find Caselaw’ (https://caselaw.nationalarchives.gov.uk) to provide public access to Judgments and Tribunal Decisions. These documents contain a variety of textual references to other sources, including other cases, legislation and other official documents. Our users would like to be able to move seamlessly between judgments, legislation and other official documents such as guidance documents. We want to turn those textual references into hyperlinks for users and we also want to identify various named entities in the documents to improve our intellectual control over the collection. Once the documents are enriched, we want to extract the additional information into a Linked Data knowledge graph to support searching and browsing features in the new Find Case law service and to enable data analysis of the collection. Problem to Be Solved Our Find Case Law service converts judgments to Legal Document Mark-up Language XML (Akoma Ntoso) and stores them in a MarkLogic database. We have developed an Enrichment service which uses automated Natural Language Processing capability to process the text of judgments to automatically find and then hyperlink neutral citations and legislation. We need to integrate the enrichment service with the Find Case Law service so that the data can be enriched and fed back into Marklogic. We would like to continue to incrementally extend our capability to enrich judgments data with new entities. Immediate priorities: 1. Full implementation of data enrichment pipelines for new data, including testing, documentation, handover training. 2. Finalise an ontology describing significant entities and citations in judgments. 3. Generation of citation RDF to be included in enrichment pipeline and bulk citation enrichment for any existing data. Next priorities (optional if time allows): 4. Create a proposal for an integration with a third-party service (VLex) subject to agreement of licensing. 5. Incorporation of entity mark-up and accompanying RDF extraction to data enrichment pipeline. 6. Bulk generation of RDF (citation and entities) for new datasets. 7. Investigate and create a proposal for enrichment of data in legislation.gov.uk. Who Are the Users The main users of the whole service are likely to be legal professionals, law students, and academics. We expect users to start their user journey from a web search, using Google or Bing say. By linking the documents together, we can improve the search engine’s ability to rank the documents, and ultimately help improve the search results. Once the user has arrived at the service, we know from user research that they value hyperlinks between documents, as it saves time and aids research. However, the wrong link, or a broken link, is frustrating, creates confusion and undermines confidence. Users have sophisticated research questions. For example, how they would like to know how a specific provision in legislation has been interpreted by the courts or which later judgments build on a particular precedent. Our service can’t directly answer these questions, but by enriching the judgments and decisions data, we can begin to support links between different data sources and a more sophisticated user interface for searching and browsing, so that users can more easily research questions like this for themselves. Early Market Engagement Our experience with Find CaseLaw (caselaw.nationalarchives.gov.uk) and with Legislation (legislation.gov.uk) gives us confidence around the feasibility of this work. Work Already Done We have developed a parser that turns Judgments into Legal Document Mark-up Language and stores the documents in a Marklogic database which will also provide search for end-users. We have developed a data enrichment service and pipeline, with capability to process Judgments documents to identify and mark-up references to case law and legislation. Although a pipeline has been developed it hasn’t been fully integrated into the service. We have applied enrichment to a small back catalogue of 50,000 Judgments. We know that there will be further data sets of a similar type that will require enrichment to the same standard. Existing Team The supplier’s team will deliver the work. The National Archives team will include a Data Scientist, a Service Owner, a Product Manager, a Delivery Manager and a User Researcher. Current Phase Alpha Skills & Experience • Experience of natural language processing and enrichment of texts in XML • Experience of modelling linked data • Experience of generating linked data from NLP pipelines • Experience of integrating services with MarkLogic/AWS Nice to Haves • Experience of working with judgments or tribunal decision documents • Experience of the Legal Document Mark-up language • Experience of documenting technical solutions so they can be maintained by others Work Location Mostly remote but some meetings as necessary onsite at The National Archives, Kew, Surrey TW9 4AD. Working Arrangments The supplier will work in accordance with Agile methodologies to scope, plan, and deliver the work incrementally, with daily stand-ups, active communication, and will conduct regular ‘show and tell’ sessions to demonstrate progress. Online meetings will take place via Microsoft Teams with Slack available for quick communication. The National Archives’ staff will be available during UK core hours (10am-4pm) each working day. The supplier will supply their own equipment and technology but will be given access to our organisational tracking app and Slack resources as appropriate. Security Clearance Baseline clearance will be required (BPSS) No. of Suppliers to Evaluate 5 Proposal Criteria • Evidence of delivering natural language processing solutions • Evidence of delivering linked data solutions • Evidence of familiarity the GDS Service Standard • Team structure, including the relevance of the team members' skills and experience Cultural Fit Criteria • Work in an open and transparent way, sharing work in progress and involving others as you go • Explain what methods you propose to use to engage; communicate, constructively challenge and work effectively with our team and other suppliers • Describe how you propose to support positive working relationships throughout the life of the contract Payment Approach Capped time and materials Assessment Method • Work history • Presentation Evaluation Weighting Technical competence 60% Cultural fit 10% Price 30% Questions from Suppliers 1. Is this linked to the previous opportunity “Enriching court judgments and legislation documents, adding hyperlinks and creating Linked Data” that was published in November 2021? Yes, this is a continuation of the work in that previous outcome. 2. Is there a current incumbent delivering this work? There is no incumbent as the previous contract has finished. 3. Who was the incumbent who did the discovery? Details of the previous contract related to this work (including award details) can be found here https://www.contractsfinder.service.gov.uk/Notice/909cef41-a507-4b83-af53-6048c544d897Prior to that there was some minimal Discovery work done in house. 4. Does your existing NLP capability only work on XML or other types of structured files? Judgments received by TNA are converted to XML, conforming to the LegalDocML schema (https://www.oasis-open.org/committees/tc_home.php?wg_abbrev=legaldocml ) for storage and publication on the https://caselaw.nationalarchives.gov.uk/ website. This XML is publicly available through the website e.g. https://caselaw.nationalarchives.gov.uk/ewca/civ/2022/970/data.xml . This is the input format for the enrichment pipeline. 5. Are you using a specific third party text analytics tool which has NLP capability or you have developed your own in-house NLP capability? The current solution uses the spaCy (https://spacy.io/ ) NLP library. You can view the code created to date in our public code repository https://github.com/nationalarchives/ds-caselaw-data-enrichment-service . 6. What are your challenges of extending the existing capability to enrich judgments data with new entities? Does any or all of the following challenges apply: This question appears incomplete. 7. Entity names are normally embedded within large free-flowing text and your existing NLP capability isn’t able to automatically extract and normalise entities within the free-flowing text? This is due to challenges in entity resolution (Name variations in spelling, word order, abbreviated terms etc.)? There is a huge amount of variation in the structure and subject matter of judgments as well as types of entities referenced and the way that they appear in the text. Work is required to identify the possibilities and the best strategies for resolution. Accuracy is extremely important, particularly in relation to potentially sensitive information, e.g. party names, and work will be required to ensure entities are only marked-up where there is adequate confidence in correctness. 8. You need support from people with relevant graph database design experience to bulk generate RDF in the way that is compatible with MarkLogic Triple Store? Some work has been completed to create a draft ontology for citations and entities related to judgments. This modelling will need to be completed together with work to create a workable graph model together with routines to generate this data from the enriched XML content and store it in MarkLogic. 9. Is it your plan to integrate the whole Case Law application and services in AWS? The data enrichment pipeline, Case Law website and associated document publishing workflows all run in AWS.

Timeline

  1. Completed: Award published20 September 2022
    Current notice
  2. Completed: Award date20 September 2022

About the buyer

The National Archives is a public sector buyer in United Kingdom publishing tenders and awards on Stotles. Explore their procurement activity and find more opportunities like this one.

AI insights

  • Is there a preferred supplier?
  • What are the buyers pain points?
  • What has the buyer previously procured?
  • What are the key requirements?
Sign-up to enrich

Decision makers

Connect with the people behind this procurement.

Contact nameJob titlePhone numberWork email
Head of Procurement+44 •••• ••••••
Commercial Director+44 •••• ••••••
Procurement Manager+44 •••• ••••••
Category Lead+44 •••• ••••••
Senior Buyer+44 •••• ••••••
Contracts Manager+44 •••• ••••••

Related topics

Topics related to Enriching judgments and tribunal decisions, entities and hyperlinks for Linked Data, integration, ranked by notice volume.

View all topics
TopicCountValue
  1. 1,745
    £636.1bn
  2. 4,088
    £269.5bn
  3. 4,672
    £63.7bn
  4. 4,815
    £66.2bn

Win more public sector contracts

Track every UK and Ireland tender in one place — set up alerts, find decision-makers, and never miss an opportunity.