Title Exploring variability in the wide field imaging datasets. Abstract The current generation of wide-field imaging cameras is producing a small trickle of data in comparison to the eventual flood expected from planned wide-field imaging, high cadence facilities like LSST and PS1. Current approaches to imaging analysis in search of variability rely on extensive use of operators to reject the 'false positives'. This approach to data analysis does not scale well and is likely not sustainable in the future. Development of new algorithms for detection of source variability via image and catalogue processing will be needed to make full exploitation of these coming data. We propose to develop and implement new variability detection and measuring algorithms that will allow projects searching for variability to better exploit existing and future datasets. We anticipate that this processing will require extensive computational resources, such as those to be made available via the CANFAR infrastructure project. Our strategy is a two step implementation: Our initial project will be to deploy existing algorithms to analyse variability in imaging acquired by - the CFHT NGVS projects - in deep Subaru-CFHT imaging (which will be searched for very high redshift SNe) - in archival imaging from the SNLS project - the ILMT project (in 1-2 years) These algorithms have well tested implementations for solar system and supernova detection and the initial effort will be to understand their deployment within a grid infrastructure. Subsequently we will explore the similarities and differences between the detection systems in an attempt to describe and develop a unified approach to variability analysis. The major outcome from this data-specialists activities will be to produce a generic 'variability' detection capability. Applicability: One of the key methods of 'compressing the data rate' is the creation of image stacks. Such stacks are intentionally designed to reduce the variability of sources, so as to reduce the noise present in the observations and thus achieve deeper images. Variability analysis will require simultaneous access to the entire imaging dataset for a given area of with the expectation (algorithm is TBD) of repeated examination of the images. Thus, given the large volume of data that will provided by the NGVS project and the large archival collection from the CFHT MEGAPrime projects, a fairly large data storage and processing capability is anticipated. The intention of the NGVS project is to provide initial sampling of sky that can used as the discovery sequence for a variability project, there is, however, no follow-up capability built into NGVS. The NGVS project is a multiyear project that will provide imaging of 100 square degrees in 5 filters. For the purposes of this proposal each 'night' of NGVS images will be processed as a 'block' of data. We anticipate that the data will be acquired as images of 6 fields imaged multiple times in a given night, these 1 night sequences of data are referred to a 'block' of observations. Each block will result have approximately 30 images. To fully exploit the NGVS dataset will require triggered followup at external facilities. For detected Kuiper belt objects and stellar transients, external followup will require rapid (results available to allow next nights observations to be planned, ~ 5 hr turn around) turn-around between data acquisition and detection. To achieve this rapid data analysis turn around will require access to significant computational capabilities in 'burst' mode: this rapid response will only be needed on days when new observations are made, ~25\% of nights during dark-lunar phase and 0% of nights during bright time (ie. roughly 15% of the time). Using GRID resources will allow this project to have 'burst' access to the significant resources needed while not requiring a these resources to be idle during the 75% of the time that the project is not running. Resource Requirements: Initial Pipeline Processing: At this time the CFEPS and SNLS projects have operational 'moving object pipeline' (MOP) and 'flux variability' (FV) pipelines. These pipelines will act as our starting point for the variability project and can be used to baseline our resource needs. Current the MOP system is run in the CADC compute cluster environment, to achieve a same-day response we deploy parallel process on a 36 cpu grid (one instance per MEGAPrime ccd). Each ccd requires ~1.5 hours processing, thus the 36 node deployment easily achieves the required data rate of providing the returned processing within ~5 hours of arrival in the CADC archive. The Flux Variability pipeline has been run a variety of small clusters. The amount of processing required is lower by a factor of ~10x for finding flux variable objects; but a significant amount of visual inspection of the images is required in the final stages, and the turnaround is similar to the MOP pipeline. Storage: Each MOP and FV processed block produces roughly 100 G-bytes of processed data that must be examined by an experienced operator. This 'in-term' storage will need to be stable for roughly 1 month time-scales (to allow re-evaluation after attempt at initial followup). Once the analysis has been concluded the long-term storage requirement will be roughly 1 M-byte per block for source catalogues. A set of web pages that provide linkages between the source catalogue and the initial detection images will be required. Memory: 500 Mbytes of memory per processor for analysis of images. Phase 2: Our second stage development effort will be to develop new imaging subtraction and variability measurements. Based on experience with existing image subtraction routines and our initial goal of more comprehensive detection with fewer false candidates, we anticipate that this phase will be roughly twice as CPU intensive as the current approach. Our phase-2 processing will thus require ~3 hours processing per night (given 36 processors availability and an anticipated data rate of 10 fields per night that NGVS is operating). Thus, or processing requirement will likely grow to ~100 CPUs available in a burst mode to process individual sections of CCDs and provide a detection catalogue. Storage requirements will be similar to the Initial Pipeline described above. Memory: ~ 500 Mbytes or memory per processor Process Management: The processing will be triggered by automatic monitoring of the CADC archive, triggering when available images match processing requirements.