Big Data: Data pipeline
Details
- Buyer
- British Museum
- Published
- 15 February 2017
- Submission
- 1 March 2017
- Source
- uk:digital_marketplace
Tender description
Summary of the work The British Museum's Big Data team seeks professional services for a beta phase project of their Data Pipeline. This will see the integration of multiple data sources from Weather to Social Media into the pipeline to support analysis and business recommendations. Expected Contract Length 12 months Latest start date 17/04/17 Why the Work is Being Done As part of the Museum's Digital Strategy the Big Data team deliver analysis and recommendations to departments on ways to understand our visitors to; improve the visitor experience through digital products, look for opportunities to increase revenue, support staff by providing clear, accurate and interactive dashboards. We require a data pipeline to consolidate our information, speed up our processes and visualize our findings to be complete within 9 - 12 months. Problem to Be Solved Combine Museum and external data sources into a Data pipeline. This includes, Audio Guide, Visitor Counting, Website, Social, Online Shop and external reference data like London tourist information and weather. Who Are the Users As a Data Scientist I need to analyse multiple data sources at once so that i can quickly analyse and visualize my findings to make recommendations to the business Early Market Engagement N/A Work Already Done From November 2016 to January 2017 the Big Data team ran an ‘Alpha’ Stage. The pilot successfully built a minimum viable product of the Big Data pipeline bringing three data sources; visitor counting, audio guide visit data and Wi-Fi data together in a centralised data warehouse and facilitated at speed analysis. The Azure Products include; Microsoft Azure SQL Server, Microsoft Azure Blob Storage, Microsoft Power BI, Microsoft Azure Machine Learning Studio. The data was brought together within Microsoft Azure SQL Server using a star schema. All personal identifiable data was anonymized by a GUID. Existing Team The main contacts are two departments in the Museum. The Digital and Publishing team (main users) and the Information Services Team. • Senior Product Manager: Big Data (Digital and Publishing), • Data Scientist (Digital and Publishing), • Junior Project Manager (Digital and Publishing), • Digital Data Analyst (Digital and Publishing), • Head of Business Solutions Delivery (Information Services), • Database Administrator (Information Services), • Analyst Programmer (Information Services), Current Phase Beta Skills & Experience • 2 or more years’ experience working with clients on Microsoft Azure Software products • Specific experience of working with the products for this project which are; Azure SQL Server , Azure Blob Storage, Microsoft Power BI, Machine Learning • Demonstrable experience of ETL, data modelling and data warehousing • Demonstrable experience of Transact SQL (stored procedure queries) • Demonstrable knowledge of data acquisition including, APIs, FTPs and other collection solutions e.g. BCP utility (bulk copy) • Demonstrable knowledge of data schemas (esp. star schemas) • Demonstrable experience of general processes to import, clean, anonymize and insert in SQL server • Excellent stakeholder engagement, to liaise with multiple Museum teams and to understand existing Museum systems and data • Demonstrate how to support and maintain the pipeline for the beta phase Nice to Haves • General stored proc layout experience • Big data (Clickstream/Social Media) experience • Experience of building a data warehouse for the purpose of a single customer view or closed loop marketing Work Location The British Museum, Great Russell Street, London WC1B 3DG Working Arrangments Onsite and offsite working. We envisage that the phases of work will not require extensive time onsite, perhaps 3 days per data source. Should you choose to include travel expenses those should be included in the cost proposal Additional T&Cs The chosen supplier will be expected to sign a non-disclosure agreement due to the sensitive nature of some Museum data. However the Museum is open to discussion as a use case and would act as a reference for the supplier. No. of Suppliers to Evaluate 3 Proposal Criteria • Methodology or approach to delivery and support • Technical capability • Timeframe proposal • Risks, assumptions and challenges identified • Added value experience in Big Data solutions • Estimated timeframes for the work • Value for money • Team structure Cultural Fit Criteria • Share knowledge and skills • Be transparent and collaborative when making decisions • Take responsibility for their work • Can work with clients with low technical expertise • Work in partnership with the Museum Payment Approach Capped time and materials Evaluation Weighting Technical competence 50% Cultural fit 15% Price 35% Questions from Suppliers 1. Could you please advise whether any 3rd party organisation(s) contributed to the Discovery and Alpha stages? The Alpha Stage was assisted by Microsoft directly. The "Evangelist team" advised us on set up and capability as we explored the products and designed the data schema ourselves. The resource for this pilot came from internal staff at the Museum (Digital and Publishing and Information Services departments). For a beta phase, with more data sources, the Museum requires support from professional services. 2. For each of the data sources listed as in scope of the data pipeline i.e. Audio guide, Visitor Counting, Website, Social, Online Shop and reference data e.g. Tourist and Weather can you please supply the following information: a. The location of the data i.e. on premise in Museum IT infrastructure, cloud environment b. The format of the data e.g. Database table, flat file, message bus, API etc. c. The volume of the dataset in row count and data volume, d. For API’s the type of API e.g. SOAP, REST etc. and for external public data a URL to the dataset. Please visit the attached link for further details on the data sources for this Beta Phase. We are still deciding on the public data sets which are called 'Reference Data' in the link below. https://worksmarttrial-my.sharepoint.com/personal/sashby_worksmarttrial_thebritishmuseum_info/_layouts/15/guestaccess.aspx?docid=01c3ee7ab316e43aca7e262734ef23b8f&authkey=AaVXzpVWfWpE8QM6WcmZRg4&expiration=2017-03-03T12%3a22%3a16.000Z 3. What on-going support will be required after the 9-12 month project phase? Is an on-going managed service required? The objective of the beta phase is to show value to the Museum of a data pipeline where the analysis team has access to multiple data sources, delivers at speed analysis and reports/dashboards. A production phase of the pipeline would be dependent on showing this value and we would aim to understand the ongoing support model of a production environment as part of the beta phase. 4. Will any of the Alpha phase solution be re-used or will we be starting from beginning? We will be starting from the beginning with a new environment, but recreating elements of the alpha as the first phase of this project. 5. How many tables are in the SQL database and are they in a dimensional star schema structure? The Alpha phase had 10 tables in the SQL, designed in a scalable star schema. It had 3 ‘Fact’ tables and 7 ‘Dimension’ tables. These have been documented and would be recreated in phase one of the project. 6. What is the current data volume in SQL? The Alpha phase has been closed down. The indicated volumes of data for the beta phase can be found in the document from this link: https://worksmarttrial-my.sharepoint.com/personal/sashby_worksmarttrial_thebritishmuseum_info/_layouts/15/guestaccess.aspx?docid=01c3ee7ab316e43aca7e262734ef23b8f&authkey=AaVXzpVWfWpE8QM6WcmZRg4&expiration=2017-03-03T12%3a22%3a16.000Z 7. Can we see the schema and documentation of the 3 fact and 7 dimension tables of the Alpha phase? The schema can be found at the following link https://worksmarttrial-my.sharepoint.com/personal/sashby_worksmarttrial_thebritishmuseum_info/_layouts/15/guestaccess.aspx?docid=0d003383c5877423bb35ea7676cd68025&authkey=Af-fuXCxTEX0VPLm0F7oV00&expiration=2017-03-03T14%3a44%3a34.000Z 8. Is it an Azure public cloud environment or do you have connectivity between the internal Museum IT network and cloud environment through VPN etc.? The pilot pipeline environment was built within an Azure public cloud. The Museum is planning provision of a secure connection between the Museum’s on premise network and the Azure virtual network via VPN, with AD authentication via Azure Active Directory, but a delivery date cannot be guaranteed for this work at present.It should therefore be anticipated that there is a chance the Beta phase of the pipeline project may begin in an isolated public cloud environment, with a view to transferring to a centralised Azure environment at a later date. 9. If direct network link is not in place to allow automated data refresh and ETL/ELT will a VPN be required or will another method e.g. SFTP file transfer push from on premise be used? The ETL layer between external systems and on premise systems will be a mixture of connections (API to FTP). We will scope what method will be the most suitable as part of the Beta phase, bearing in mind the progress of the work outlined in Question 8. Currently the transformed data is stored within an Azure SQL Database but this may change depending on requirements and data volume. 10. Are their availability and/or disaster recovery requirements that should be considered? Currently there are no anticipated requirements beyond those provided by the relevant Azure services, but these will require testing, for example, point in time restore of Azure SQL Database, and the ability to rerun data extractions on identified failure. However, these requirements will be validated during the Beta phase. 11. Azure Machine Learning is mentioned as having been used during the alpha phase. What will Azure Machine Learning be used for during the beta phase? In the first instance we are aiming to use K means clustering for segmentation of online visitors and linear regression testing for prediction models 12. How is social media being integrated into the pipeline? What requirements do you have around social media and how is this information being used? We are looking to integrate via APIs from Hootsuite and Facebook. This would need a hack day to investigate the details further. We are looking for professional services who could assist us with this. We have a number of business questions on social media. As this is the fifth phase we are aiming to combine the social information with other information e.g. sales and website search. 13. Are there any specific requirements around searching the data? We have a number of business questions and analytic studies that our data scientist and data analyst will perform on the data. 14. Is all data in English or are there multi-lingual requirements? There are no multi-lingual requirements during this phase. 15. Are you essentially trying to build up a single customer view? Understanding our visitors before, during and after a visit is an objective of the team. We will be looking to find out if the same visitors are using different products to build up a customer journey, which is why customer keys are important. However our focus is to understand visitor behaviour to make recommendations on ways we can make better decisions on:o Product development and content for visitor experience o Opportunities to increase revenue o Communicating effectively with new and loyal visitors The Museum does not have a centralized CRM system. 16. Is there any reason you don't plan to use Azure SQL Data Warehouse / are you considering Azure SQL Data Warehouse? Our understanding is that the Azure SQL Data Warehouse does not currently store data in the UK south region, which is a requirement for us. The product is also more expensive than Azure SQL Server which allows us to structure the data as if it were a data warehouse. 17. For the ETL, do you have a tool in mind? e.g. Azure Data Factory We trialed Data Factory during the Alpha Stage, however the product was not suitable at this time. The ETL layer needs to be as flexible as possible. For the most part this is SQL store procedures and power shell scripts to integrate with legacy systems. 18. "The excel ""Data Sources for Datapipeline Beta Project_Gcloud"" mentions the data sources and their corresponding phases. There are 5 phases in total. Please specify if all the phases are in scope of the Beta stage. If not, please specify the phases in scope for Beta stage. The five phases are in scope for the Beta phase. 19. Could you please also provide the number of tables within each of the data sources in scope for Beta? We are seeking professional services to work in partnership with us to finalize the details of what fields we will take from each data source to create tables in the data warehouse. 20. Is the data being analysed in real-time or is it just historical data? How up to date does the data need to be? For the beta phase the data source that may need to be daily is the Audio guide for a product dashboard. The rest will be historical and weekly batches. 21. How long is the data required to be kept for? For the length of the Beta Phase. 22. Are there any reporting requirements for the data? Yes. We are anticipating 5 Museum dashboards using Power Bi and all analysis will need some kind or reporting/visualization to communicate the value to the Museum. 23. Could you please specify if an analytics solution is required to prepare hypotheses based on the source dataset and if visualisations refer to preparing the the findings of this analytics solution in a visually convenient format for the business. Our Data Scientist and Data Analyst will be conducting the analysis using their technical and programming skills (R, Python and SQL) and the support of the Microsoft products listed in the brief. 24. Could you please specify if the Existing internal big data team for digital and publishing will be involved in full capacity in the beta stage. If yes, then please confirm our assumption that the supplier will need to specify the efforts required from the ""existing team"" as well as the efforts that the proposed supplier team will expend, which look like they will be limited to data engineers, visualisation experts and perhaps a business analyst. The Data Scientist and Project Manager will be at full capacity for the beta. The Data Analyst and Snr Product Manager will be at around 50%. Collectively the team has skills in analysis, general process to import- clean- annomymise and insert, programming languages, visualizations, information design and project management. We are looking for technical/ professional services to support the integration of the data sources to support the analysis. Please include all assumption in the proposal. 25. Please can you clarify some details of the feed from the website. The table says 3TB updated weekly, but another row says it should be updated in near real time. Is the 3TB the total storage or the amount transferred every night? Apologies for the confusion here. A year's worth of raw website data is estimated to be 3TB and we would look to move this across on a nightly basis. 26. Please specify if a fully costed solution and complete proposal is expected at this stage (due for submission on 1st of March 2017). The requirements seem to indicate that only responses for ""essential as well as nice-to-have skills and experience"" are expected at this stage and the ""proposal"" stage will follow, soon after 1st March 2017. We are seeking a proposal that will help us understand the following key points:The methodology / approach to delivery.Costs of professional services e.g. day or hourly rate. Where there are significant assumptions on any costs, recommendations on a level of contingency to manage risk.Details of technical staff assigned to the project. Is there ability to provide maintenance support for the ETL layer.Outline any key challenges or risks you may envisage with the project from prior experience. Provide demonstrable experience, references, or case studies in Azure software delivery. Any additional value to the project.Time frame. 27. What are the machine learning requirements? What are you hoping to achieve through the use of machine learning? In the first instance we are aiming to use K means clustering for segmentation of online visitors and linear regression testing for prediction models 28. Can you share with us a solution architecture diagram of the Alpha build including the Azure services and connectivity to internal and external data sources? We are unable to provide the solution diagrams at this stage. If you have any specific question I may be able to answer them in follow up. 29. Have you given any thought with how to deal with unstructured data (as opposed to structure data in the star schema) We have not yet looked at the unstructured data from Social Media. Suggestions or recommendations on how to do this would be very welcome in a proposal. 30. Have you considered an ETL product? Have you considered SSIS as an ETL product? We trialed Microsoft Azure Data Factory during the Alpha Stage, however the product was not suitable at this time. The ETL layer needs to be as flexible as possible. For the most part this is SQL store procedures and power shell scripts to integrate with legacy systems. Python script has also been tested. 31. Who is ‘paying’ for the service? How will you prove ROI? Funding for this initiative is from the Digital and Publishing department at the British Museum. The Big Data team provide recommendations to the business on ways they may be able to increase revenue and then support testing these through analyzing data. There is not direct ROI target for the team. We support the delivery of KPIs for other teams. 32. The requirements mention an NDA to be signed by the supplier. Could you please confirm if the supplier resources need to have a particular security clearance as well. The Cyber Essentials accreditation would also be required. 33. Is there a likelihood that the end of the beta phase will be the end of the product? The aim of the Beta is to show the value of a data pipeline to the Museum. By delivering analysis, dashboards, predictive models and by understanding enough about how this would work long term, we can recommend a production phase for the Museum's approval. 34. Has there been any primary research done on the website? Yes. The Museum conducts surveys on motivations and task completion. We also use Adobe and Google web analytics to track visitor behaviour
Timeline
- Completed: Tender published15 February 2017Current notice
- Completed: Submission date1 March 2017
About the buyer
British Museum is a public sector buyer in United Kingdom publishing tenders and awards on Stotles. Explore their procurement activity and find more opportunities like this one.
Decision makers
Connect with the people behind this procurement.
| Contact name | Job title | Phone number | Work email |
|---|---|---|---|
| Head of Procurement | +44 •••• •••••• | ••••••••@british-museum.gov | |
| Commercial Director | +44 •••• •••••• | ••••••••@british-museum.gov | |
| Procurement Manager | +44 •••• •••••• | ••••••••@british-museum.gov | |
| Category Lead | +44 •••• •••••• | ••••••••@british-museum.gov | |
| Senior Buyer | +44 •••• •••••• | ••••••••@british-museum.gov | |
| Contracts Manager | +44 •••• •••••• | ••••••••@british-museum.gov |
Related topics
Topics related to Big Data: Data pipeline, ranked by notice volume.
- 1,484£14.9bn
Related buyers
Buyers similar to British Museum.
- 2,703£1.7bn
- 909£19.4bn
- 696£8.8bn
- 480£15.4bn
- 436£3.8bn
- 353£20.6m
- 306£371.3m
- 263£233.9m
- 254£580.8m
- 241£10.1bn
Win more public sector contracts
Track every UK and Ireland tender in one place — set up alerts, find decision-makers, and never miss an opportunity.
