GautiMaan Posted October 18, 2013 Posted October 18, 2013 Random Guy and Dm,you guys might as well exchange phone numbers over PM:giggle:
randomGuy Posted October 19, 2013 Posted October 19, 2013 DM, your understanding is incorrect my friend. FPIR of .06% on 1.2billion unique records means that for 99.94% people , there would be no problems at all, Their biometrics will match uniquely with their record. For the rest 0.06% people, the biometrics will match with someone else's biometrics too. Please think over it.
The Outsider Posted October 19, 2013 Posted October 19, 2013 Important thing to note here is that the number of unique records to be compared with' date=' have increased 2500 times (100million/40k). compare this to the increase in FPIR increased by ~22 times. [/quote'] I don't get this. FPIR scales roughly as the population size, N. How does population size increase by 2,500 times and FPIR only by 20 times? If FPIR is 2.5*10^(-5) at 40,000 one would expect FPIR at 100,000,000 to be: 1-(1-(2.5*10^(-5))^2500 = 0.0606 = 6.06%.
randomGuy Posted October 19, 2013 Posted October 19, 2013 I don't get this. FPIR scales roughly as the population size, N. How does population size increase by 2,500 times and FPIR only by 20 times? If FPIR is 2.5*10^(-5) at 40,000 one would expect FPIR at 100,000,000 to be: 1-(1-(2.5*10^(-5))^2500 = 0.0606 = 6.06%. The only possible explanation that I could think is in post#85. Think of set {a,b,c,...} as your set of unique records of 40K people or of 100 million people. The FPIR of 0.06% on 1.2 billion gives you 7.2 lac duplicates(false positives). That would have been acceptable and something that the UIDAI would have anticipated at the beginning. FPIR of ~6% on 1.2 billion would have given 7.2 crore duplicates(false positives) and the UIDAI would have rejected such a project at the outset coz such FPIR is obviously unacceptable.
The Outsider Posted October 19, 2013 Posted October 19, 2013 The only possible explanation that I could think is in post#85. Think of set {a,b,c,...} as your set of unique records of 40K people or of 100 million people. The FPIR of 0.06% on 1.2 billion gives you 7.2 lac duplicates(false positives). That would have been acceptable and something that the UIDAI would have anticipated at the beginning. FPIR of ~6% on 1.2 billion would have given 7.2 crore duplicates(false positives) and the UIDAI would have rejected such a project at the outset coz such FPIR is obviously unacceptable. The 6% figure is for 100,000,000. Will be around 70% for 1,200,000,000. There is no ambiguity here. These are mathematically prove able figures. Either the 40,000 number is wrong or the 100,000,000 number is wrong. They are not reconcilable.
randomGuy Posted October 19, 2013 Posted October 19, 2013 The 6% figure is for 100' date='000,000. Will be around 70% for 1,200,000,000. There is no ambiguity here. These are mathematically prove able figures. Either the 40,000 number is wrong or the 100,000,000 number is wrong. They are not reconcilable.[/quote'] DM's link (of .0025% fpir) says -
Crookbond Posted October 19, 2013 Posted October 19, 2013 DM, your understanding is incorrect my friend. FPIR of .06% on 1.2billion unique records means that for 99.94% people , there would be no problems at all, Their biometrics will match uniquely with their record. For the rest 0.06% people, the biometrics will match with someone else's biometrics too. Please think over it. No RG! You're making a fundamental error on the FPIR. It is based on "comparisons" not "unique records". In fact, my calculations are backed in another article too - How many false positives should India expect? In the Results section of their report UIDAI define FPIR, the false positive identification rate, and they say “we will look at the point where the FPIR (i.e. the possibility that a person is mistaken to be a different person) is 0.0025 %â€. At that point, UIDAI would get 2½ false positives on average for every 100,000 comparisons. http://dematerialisedid.com/BCSL/Drown.html Note comparisons and not unique records.
randomGuy Posted October 19, 2013 Posted October 19, 2013 No RG! You're making a fundamental error on the FPIR. It is based on "comparisons" not "unique records". In fact, my calculations are backed in another article too - http://dematerialisedid.com/BCSL/Drown.html Note comparisons and not unique records. No yaar, if that was the case (your definition's case), the project wouldn't even have begin(at fpir of 0.0025%) at the first place. Also, your link - FPIR: False Positive Identification Rate: This is the likelihood that a person‟s biometrics is seen as a duplicate (i.e., the biometric deduplication software identifies his biometrics as matching with that of a different person), even though it is not a duplicate in reality. http://uidai.gov.in/images/FrontPageUpdates/uid_enrolment_poc_report.pdf
randomGuy Posted October 19, 2013 Posted October 19, 2013 Also, if your definition was correct, FPIR wouldn't have changed 0.0025% to 0.057% on larger population. .057% is what I found after googling(they say the figures are given by UIDAI) - http://www.planetbiometrics.com/creo_files/upload/article-files/India_boldly_takes_biometrics_where_no_country_has_gone_before.pdf http://www.biometrics.org/bc2012/presentations/UIDAI/Technology_Tampa_v1040.pdf
Crookbond Posted October 19, 2013 Posted October 19, 2013 No yaar, if that was the case (your definition's case), the project wouldn't even have begin(at fpir of 0.0025%) at the first place. Also, your link - http://uidai.gov.in/images/FrontPageUpdates/uid_enrolment_poc_report.pdf Huh? The UIDAI is not GOD and they can make mistakes. Just because the project got ahead doesn't mean the FPIR definition would be different. The article I linked above argues why it was a problem to get ahead with the project with that FPIR.
Crookbond Posted October 19, 2013 Posted October 19, 2013 Also, if your definition was correct, FPIR wouldn't have changed 0.0025% to 0.057% on larger population. .057% is what I found after googling(they say the figures are given by UIDAI) - http://www.planetbiometrics.com/creo_files/upload/article-files/India_boldly_takes_biometrics_where_no_country_has_gone_before.pdf http://www.biometrics.org/bc2012/presentations/UIDAI/Technology_Tampa_v1040.pdf Yaar - FPIR grows linearly with size of the dataset [0]. You may also want to read up on who Anil Jain is. [0] [ame=http://www.amazon.com/Handbook-of-Biometrics-ebook/dp/B001A4BG8A/ref=sr_1_1?ie=UTF8&qid=1382195589&sr=8-1&keywords=Handbook+of+Biometrics]Amazon.com: Handbook of Biometrics eBook: Anil K. Jain, Patrick Flynn, Arun A. Ross: Kindle Store[/ame]
Crookbond Posted October 19, 2013 Posted October 19, 2013 I don't get this. FPIR scales roughly as the population size' date= N. How does population size increase by 2,500 times and FPIR only by 20 times? If FPIR is 2.5*10^(-5) at 40,000 one would expect FPIR at 100,000,000 to be: 1-(1-(2.5*10^(-5))^2500 = 0.0606 = 6.06%. You are correct! There's just one explanation - the FNIR rate goes worse. FNIR doesn't change as the dataset size while FPIR does. In biometric systems, the FPIR is kept constant the FNIR is varied i.e. you can observe the same FPIR on a larger dataset you have to increase the FNIR. The relationship of trade off between the FPIR and FNIR is log-log.
randomGuy Posted October 19, 2013 Posted October 19, 2013 Yaar - FPIR grows linearly with size of the dataset [0]. You may also want to read up on who Anil Jain is. [0] Amazon.com: Handbook of Biometrics eBook: Anil K. Jain, Patrick Flynn, Arun A. Ross: Kindle Store This is what I am saying, it IS changing and it is supposed to change. If UIDAI believed that it would get 2½ false positives on average for every 100,000 comparisons on 1.2 billion, they wouldn't have started the project. It means the system will match 30000 records for each person. My friend, your definition can not be correct for this reason.
The Outsider Posted October 19, 2013 Posted October 19, 2013 DM's link (of .0025% fpir) says - Yeah, I read that. But that is not reconcilable with the 100,000,000 FPIR which was quoted. One of them is likely wrong. Please read pages 321-322 from the following link on google books: http://books.google.com/books?id=XPC9ucFbddsC&printsec=frontcover&source=gbs_ge_summary_r&cad=0#v=onepage&q&f=false Quoting some relevant part: FPR(N) = 1 - (1-FPR(1))^N This relationship clearly shows that the FPR is not a constant as a function of sample size. The FPR you expect for 40,000 will scale roughly linearly with N if you increase the sample to 100,000,000 and then to 1,200,000,000. Quoting again: The false positive rate increases drastically with database size because each additional entry in the database provides another opportunity to randomly achieve a high score. A matcher operating at a point where its false positive verification rate is 1% maybe satisfactory in a verification application, but even in a small scale identification application, the error rate will become unacceptable. He goes on to give some examples as well. Also quoting: In order to keep false positive rate within reasonable bounds when operating on large population sizes, a matcher must be operating in a mode for which its false positive rate for verification is in the range of 10^(-9) - 10^(-6). Given, the huge size of the population the number would have to be 10^(-9) or even lower. The quoted numbers are nowhere close to that.
randomGuy Posted October 19, 2013 Posted October 19, 2013 Yeah, I read that. But that is not reconcilable with the 100,000,000 FPIR which was quoted. One of them is likely wrong. Please read pages 321-322 from the following link on google books: http://books.google.com/books?id=XPC9ucFbddsC&printsec=frontcover&source=gbs_ge_summary_r&cad=0#v=onepage&q&f=false Quoting some relevant part: This relationship clearly shows that the FPR is not a constant as a function of sample size. The FPR you expect for 40,000 will scale roughly linearly with N if you increase the sample to 100,000,000 and then to 1,200,000,000. Quoting again: He goes on to give some examples as well. Also quoting: Given, the huge size of the population the number would have to be 10^(-9) or even lower. The quoted numbers are nowhere close to that.I understand what you are saying. Let's leave thepoint of fpir of 0.0025% and 0.057% not reconciling for just a little while. Infact Yours and mine understanding is matching. See post 84, I've calculated fpir for 80k to be ~0.0050% based on fpir of .0025% on 40k. It also comes out ~.0050% for 80k records from the formula you gave. For now, pls tell DM that his understanding/ definition of fpir is incorrect.
Crookbond Posted October 19, 2013 Posted October 19, 2013 This is what I am saying, it IS changing and it is supposed to change. If UIDAI believed that it would get 2½ false positives on average for every 100,000 comparisons on 1.2 billion, they wouldn't have started the project. It means the system will match 30000 records for each person. My friend, your definition can not be correct for this reason. Like I said earlier, UIDAI is not GOD and your assertion lies on the assumption that UIDAI acts rationally. There's no way to know UIDAI's intentions or neither do I care. It's analogous to say - MMS is not responsible for the spectrum scam because if he knew the scam were to take place, he would never give consent.
The Outsider Posted October 19, 2013 Posted October 19, 2013 I understand what you are saying. Let's leave thepoint of fpir of 0.0025% and 0.057% not reconciling for just a little while. Infact Yours and mine understanding is matching. See post 84, I've calculated fpir for 80k to be ~0.0050% based on fpir of .0025% on 40k. It also comes out ~.0050% for 80k records from the formula you gave. For now, pls tell DM that his understanding/ definition of fpir is incorrect. In post #84 you seem to be saying 0.0050% for 80,000 records is wrong and the correct number should be around 0.0026%. In the above post you seem to be saying that the 0.0050% for 80,000 records is correct. If you are indeed saying that the 0.0050% for 80,000 records is correct and by extension 70% for 1,200,000,000 is also correct, then we are in agreement and the scheme is useless based on analysis of the available data. That seems to be DM's inference as well.
Crookbond Posted October 19, 2013 Posted October 19, 2013 In post #84 you seem to be saying 0.0050% for 80' date='000 records is wrong and the correct number should be around 0.0026%. In the above post you seem to be saying that the 0.0050% for 80,000 records is correct. If you are indeed saying that the 0.0050% for 80,000 records is correct and by extension 70% for 1,200,000,000 is also correct, then we are in agreement and the scheme is useless based on analysis of the available data. That seems to be DM's inference as well.[/quote'] I would like your take on Post #87 assertion by RG. His interpretation of FPIR is as follows - I categorically disagree. FPIR just means that there would e around 0.06% viz. 784,800 biometric records for which the system would fail. This doesn't mean that these many people wouldn't enroll. PS - In all of the above calculations, there are MANY assumptions including quality of biometric data captured, independent and identically distributed biometric data, unlimited computational power, algorithm scale, real-time data stream, JIT solutions etc.
The Outsider Posted October 19, 2013 Posted October 19, 2013 I would like your take on Post #87 assertion by RG. His interpretation of FPIR is as follows - I categorically disagree. FPIR just means that there would e around 0.06% viz. 784,800 biometric records for which the system would fail. This doesn't mean that these many people wouldn't enroll. PS - In all of the above calculations, there are MANY assumptions including quality of biometric data captured, independent and identically distributed biometric data, unlimited computational power, algorithm scale, real-time data stream, JIT solutions etc. My understanding is that stating FPR without making a reference to the sample size is incomplete and any conclusions drawn from it would be incorrect. Reason being FPR is a function of N. Stating FPR without specifying if it is a FPR of 1, 40k, 100 million, or 1.2 billion doesn't tell me anything. A FPR of 0.06% for 1.2 billion people is going to be great, but a FPR of 0.06% for 1 person going to be used on a population of 1.2 billion is absolutely shoddy. For this discussion, a FPR of 0.025% on a 40k sample is great but extrapolation to 1.2 billion people of the same FPR will yield a figure of 70% which is shoddy.
randomGuy Posted October 20, 2013 Posted October 20, 2013 In post #84 you seem to be saying 0.0050% for 80' date='000 records is wrong and the correct number should be around 0.0026%. In the above post you seem to be saying that the 0.0050% for 80,000 records is correct. If you are indeed saying that the 0.0050% for 80,000 records is correct and by extension 70% for 1,200,000,000 is also correct, then we are in agreement and the scheme is useless based on analysis of the available data. That seems to be DM's inference as well.[/quote'] In that post, I am trying to reconcile .0025% on 40k with .057% on 100 millions somehow . So trying to say how fpir on 80k(which we are able to derive intuitively) may be say0.0026% and not .005%. You are saying the fpirs on 40k n 100m won't reconcil. Chalo fine, but reconciliation comes later, DM's definition of fpir is incorrect, first we should correct that. Yeah sure, agreed. Now tell him what fpir means, like what does fpir of .06% on 1.2 billion mean?
Recommended Posts